Dzone iconDzoneSep 18, 2026

When Your Benchmark Leaks the Answer

A detector I built was scoring 0.067 recall on temporal errors, meaning it caught about one in fifteen of the wrong dates it was supposed to find. The benchmark was the problem, and not in a way that showed up anywhere in the code.

When Your Benchmark Leaks the Answer

Share this story

Send the public story page.

Useful takeaways from this story.

A detector I built was scoring 0.067 recall on temporal errors, meaning it caught about one in fifteen of the wrong dates it was supposed to find.

The benchmark was the problem, and not in a way that showed up anywhere in the code.

I assumed the extraction was broken and went looking for the bug.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

A detector I built was scoring 0.067 recall on temporal errors, meaning it caught about one in fifteen of the wrong dates it was supposed to find. The benchmark was the problem, and not in a way that showed up anywhere in the code. I assumed the extraction was broken and went looking for the bug.

Details worth keeping

I assumed the extraction was broken and went looking for the bug. The contexts had been written in the wrong voice.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app