These 4 questions are the antidote to AI hype
The next time you encounter AI claims stop and ask yourself four questions.
Ohmygod I’m still writing about AI on my political philosophy blog
I wasn’t expecting to write a post like this so late into the AI bubble, which I wrote about in March, but after recent AI news stories I felt obligated to post this. Hopefully it helps someone make a little more sense of the world. It’ll be the first part in an ongoing series titled “Take the antidote.”
As I’ve said multiple times before, the media has played a substantial role in inflating the AI bubble, mainly by parroting marketing claims of AI companies and fomenting a sense of dread and inevitability that recently has started to backfire. Still, Big Tech persists in invoking the specter of AI creating a permanent underclass, that politicians like Bernie Sanders have been nudged into taking seriously (I recently ranted about this). However, the evolving economics of AI have started to weaken near-term labor displacement claims. Even Sam Altman felt compelled to walk back immediate job displacement claims, earlier this year.
But AI hype takes other forms, often with more of an existential and urgent tone. If we’re not experts in AI or in the domains AI touches, how can we still read AI stories critically?
Here are four questions that are a great starting point for clarity. You can think of these as bars that a claim or story must clear in order to be coherent rather than just vague hype.
1. What do you mean by “AI?”
AI, or artificial intelligence, is a field of very loosely related techniques and technologies, so the first thing to do when you hear the letters A and I is to ask: “What do you mean?”

Today’s AI boom is being driven by an algorithm called a Transformer. Most people refer to technologies downstream of the Transformer, like language models, as generative AI (GenAI). Although not every GenAI technology uses Transformers under the hood.
I used to be comfortable using GenAI as a term, however, I’ve recently begun moving away from it because it is almost as ambiguous as AI. I think for most people the term is fine—I genuinely don’t expect the average person to follow the field and know the difference between Mamba and a Transformer,[1]1 for example. However, the more specific we can be, the more clarity we can have about a specific claim or topic.
One way to do this is by asking about the modality of a specific system. Is the story about a LLM-based chatbot, like ChatGPT or Claude? Is it a system with a harness, like Claude Code? A narrow-domain model like a chess game AI or scientific model? These distinctions matter. While Claude Code is Claude under the hood, it’s fundamentally a different system as the coding harness gives Claude structure and feedback it does not get in the chat interface. As I said in my dark pattern post, human-designed scaffolding plays a huge part in giving LLMs their capabilities, and different scaffolding changes what those abilities are.
When we know what type of AI a specific story or claim is about, we can begin to contextualize the type of environments and requirements needed for a specific capability claim to be valid.
For the rest of the post, I will walk you through debunking claims that AI boosters made earlier this year about the end of software development due to coding agents. Consider services like Claude Code and OpenAI’s Codex as the answer to this first question given this example.
2. What is the capability claim being made?
The goal of the second question is to identify what capabilities are being attributed to the relevant AI system. Hype-y statements and news stories rely on you reacting to specific capability claims that are either explicitly stated or implied.
Your first job is untangling when an explicit versus implicit claim is being made. In many cases, stories relying on implicit claims are speculation about something that hasn’t happened yet or may never happen. This typically takes the form of extrapolating how well an AI system’s abilities generalize beyond the context provided by the claim.
To give you an example, earlier this year technology stocks took a beating in what was dubbed the SaaSpocalypse (Software as a Service apocalypse). SaaS companies refer to tech companies selling subscription software, like how Google Drive sells monthly storage. The thesis for investors was simple: LLMs have seemingly gotten better at coding, so why would anyone pay for software when they could make and deploy their own with AI?
While LLMs have always been able to code, in the last year or so, coding harnesses have become extremely popular. This changed the type and duration of coding tasks LLMs could perform. Whereas you might use a chatbot to produce a one-off code snippet, a LLM with a coding harness becomes a self-prompting agent that can manage a much larger project with less guidance. We are still debating how consistently reliable these systems are at longer horizon coding tasks, but given limited scope of what chatbots could do even a year ago, this is a notable difference.
While coding harnesses have changed how LLMs behaved, where did the evidence for the SaaSpocalypse come from? Aside from anecdotes, in many cases coding or task benchmarks provided by AI companies or AI safety non-profits like METR (Model Evaluation & Threat Research) were referenced to justify this idea.
AI company benchmarks can be difficult to assess because sometimes basic questions, like whether a benchmark’s questions were in an AI’s existing training set, can’t always be answered. This type of contamination limits the utility of the benchmark to tasks a coding LLM has already seen. And unless the benchmark evaluator specifies, you won’t know if this issue has been addressed.
Alternatively data, like the AI time horizon tracking produced by METR, are more transparent in that there is a published methodology to evaluate. Relying on this, however, comes with its own pitfalls. METR’s time horizon data has had serious criticism made against it, and even METR’s own researchers have noted the divergence that their data has had from real-world conditions. So while METR’s time horizon data has been used to argue that bigger AI models are more capable of taking on longer and more complex tasks, extrapolating much from the data set is challenging. Fundamentally, relying on self-published benchmarks or tests with small sample sizes is not enough information to validate whether AIs can replace both the labor pool and product of an entire industry.
3. What infrastructure is needed for the claim as presented to be true?
All AI systems, including large language models, require scaffolding and humans in the loop to be useful. The dark pattern or illusion that convinces the public to give into hype is the fact that in modern AI systems human discretion and labor are hidden from users. In a past post, I compared this to commodity fetishism, an idea that reflects how the products we interact with have their supply chains obscured from us.
Let’s go back to the SaaSpocalypse and coding job loss narrative, even though we’ve established it was partly built by over-extrapolating benchmarks. In order for SaaSpocalypse fears to be true, what type of infrastructure is necessary?
Systems like Claude Code or OpenAI’s Codex, two popular coding agents, rely on coding harnesses. These are standalone programs that must be built and maintained by humans. When Claude Code accidentally leaked to the public, we got a list of the assumptions its developers made when building the system and what systems were in place to maintain Claude Code. Keep in mind work to maintain the harness is in addition to the human labor that goes into training an LLM and maintaining the underlying subsystems it uses at runtime.
We’re not done as these are just the systems on the AI company’s side. Customers of AI companies who want to deploy a service like Claude Code require experts who know how to integrate it within their corporate environment. This challenge can be steep enough that OpenAI and Anthropic employ so-called Forward Deployed Engineers[2]2 whose job it is to set up and maintain agentic AI for their corporate customers. However, even once an agent is up and running, humans continually remain in the loop to steer and monitor the system, lest it delete critical components of the environment it’s working in.
4. Who benefits from the frame as currently presented?
The final question asks, given the current framing of a story, who ultimately benefits from the portrayal. There are many cases where hype, even when it makes AI companies look irresponsible, serves to bolster what “AI” appears to be capable of. This question can be directed both at the source of a claim, like say an AI lab, and the entity delivering the claim too, like a news outlet. The purpose of this question is not to use the analysis to completely disregard the claim, but to get you to consider what other sides may contribute to a specific story or claim were they included in the discussion.
Do you need to know the answer to all four questions?
No. In fact, while I chose an example I felt comfortable with, in many cases it isn’t possible for you to have answers. The purpose of these questions isn’t for you to respond to hype, but to get you to take a deep breath and be aware that there are other angles to AI reporting.
Don’t fall for the hype
I’ve been sitting on this post since last October, right before I published my popular LLM “intelligence” is a dark pattern post. We’re now even deeper in the bubble and the stories coming out of AI companies are only going to get weirder. Stay safe and don’t freak out when you see a wild AI story. When in doubt, phone a friend, or read Misaligned Markets and feel free to send me your questions!
If you liked this blog post support my tea habit by tipping me!
- Certainly not going to pretend like I do. Here's a short primer I recently watched, though.
- It's insane how normalized military language has been normalized in tech.