Dev iconDevSep 6, 2026 ~1 min source read

The Harness Is Not Intelligence: What Is Actually Improving in AI Agents?

A few months ago, I wrote about a feeling I still have today: AI models, and especially coding agents, no longer give me the same sense of huge leaps that they used to. Newer models usually make fewer mistakes, follow instructions better, and sometimes solve problems that older versions could not.

The Harness Is Not Intelligence: What Is Actually Improving in AI Agents?

Share this story

Send the public story page.

Useful takeaways from this story.

A few months ago, I wrote about a feeling I still have today: AI models, and especially coding agents, no longer give me the same sense of huge leaps that they used to.

Newer models usually make fewer mistakes, follow instructions better, and sometimes solve problems that older versions could not.

It is becoming harder for me to feel those improvements as a real jump in capability.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

A few months ago, I wrote about a feeling I still have today: AI models, and especially coding agents, no longer give me the same sense of huge leaps that they used to. Newer models usually make fewer mistakes, follow instructions better, and sometimes solve problems that older versions could not. It is becoming harder for me to feel those improvements as a real jump in capability.

How it works

  • It even suggested that if someone is still complaining about current models, they probably do not know how to use or control them properly.
  • Recently, I saw a post arguing, more or less, that at this point models are no longer better or worse than each other, but simply have different behaviors, and that what is actually good or bad is the...
  • The harness matters a lot First, we need to separate things that are often thrown into the same bucket.
  • Around the model there is an entire infrastructure: tools, context management, system prompts, planning, retries, and many other things.
  • After spending time building my own coding agent, I am even more convinced that this layer matters enormously.

What to take from it

So yes: simply saying "this model is bad because it performed badly inside X agent" can be unfair. You can put the exact same model inside two different products and get completely different experiences.

Details worth keeping

GPT, Claude, Gemini, GLM, or any other LLM is only one part of the system. One agent may manage context better than another. There is a huge distance between that and saying...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app