When Your Benchmark Leaks the Answer
A detector I built was scoring 0.067 recall on temporal errors, meaning it caught about one in fifteen of the wrong dates it was supposed to find. The benchmark was the problem, and not in a way that showed up anywhere in the code.
