Dev iconDevSep 5, 2026 ~1 min source read

JuryTrace: make agent-judge failures inspectable

When an agent regression judge says "accept," the interesting question is often not the score. It is whether another judge agrees, whether either judge showed its work, and what happens when they do not.

JuryTrace: make agent-judge failures inspectable

Share this story

Send the public story page.

Useful takeaways from this story.

When an agent regression judge says "accept," the interesting question is often not the score.

Its five synthetic traces cover a clean agreement, a shared rejection, a disagreement, instruction-like text treated as data, and a fixture false acceptance.

JuryTrace is a small Python tool for that boundary: it runs two configured judges over typed JSONL trajectories, validates their evidence, and sends unresolved cases to a durable human queue.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

When an agent regression judge says "accept," the interesting question is often not the score. It is whether another judge agrees, whether either judge showed its work, and what happens when they do not. JuryTrace is a small Python tool for that boundary: it runs two configured judges over typed JSONL trajectories, validates their evidence, and sends unresolved cases to a durable human queue.

How it works

  • It is a workflow and evidence exercise, not a claim that two judges produce objective truth.
  • Its five synthetic traces cover a clean agreement, a shared rejection, a disagreement, instruction-like text treated as data, and a fixture false acceptance.
  • What actually runs The orchestration is a real asynchronous LangGraph workflow.
  • StateGraph moves through hard_gates, dispatch_judges, score_and_compare, then either persist_clean or persist_reviews.
  • A valid result contains accept or reject, a score, a rationale, and one or more evidence spans.

What to take from it

The receipt records four agreements, one review item, and one false acceptance against a supplied reject golden label. The dispatch node creates one coroutine per trajectory and judge and awaits them concurrently with asyncio.gather. The graph also enforces exactly two judges, distinct judge IDs, distinct configurations, maximum traces, and maximum calls.

Details worth keeping

The committed reference run uses deterministic offline fixture judges. Each call has its own asyncio.timeout boundary. The provider boundary returns a typed proposed verdict.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app