Ombulabs iconOmbulabsSep 23, 2026 ~1 min source read

The Two Cheapest Agent Eval Signals Both Fail

Once a team stops trusting leaderboards, the next move is usually to build an evaluation of its own. We've argued before that benchmark scores are a poor predictor of production agent performance, and the natural follow-up question is what to measure instead.

The Two Cheapest Agent Eval Signals Both Fail

Share this story

Send the public story page.

Useful takeaways from this story.

Once a team stops trusting leaderboards, the next move is usually to build an evaluation of its own.

The first is a panel of LLM judges: ask several models whether an output is factually supported, and take the majority vote.

We've argued before that benchmark scores are a poor predictor of production agent performance, and the natural follow-up question is what to measure instead.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

Once a team stops trusting leaderboards, the next move is usually to build an evaluation of its own. We've argued before that benchmark scores are a poor predictor of production agent performance, and the natural follow-up question is what to measure instead. Two answers come up almost every time, because they are the two cheapest evals to stand up.

How it works

  • The second is step-level grading: walk the agent's trajectory one step at a time, score each step, and blame the first one that looks wrong.
  • Neither needs labeled data, and both can be running by the end of the week.
  • The first is a panel of LLM judges: ask several models whether an output is factually supported, and take the majority vote.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app