69 Tests. All Passing. Zero Bugs Caught.
Together they caught zero of the eleven bugs I had deliberately planted in that module. A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten.

Together they caught zero of the eleven bugs I had deliberately planted in that module. A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten.

Together they caught zero of the eleven bugs I had deliberately planted in that module.
Three approaches, same model, same token ceiling, twelve widely used Python libraries including cachetools, toolz, tenacity and boltons.
A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten.
The page is ready to read now. The fuller skim-friendly version will appear here automatically.
Together they caught zero of the eleven bugs I had deliberately planted in that module. A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten. It changes source code in small ways, flips a comparison, alters a constant, deletes a raise, then runs the existing test suite and records which changes the suite fails to notice.
approach caught one prompt, "write more tests" (556 tests) 9 of 53 one test per call, no targeting (53 tests) 2 of 53 targeted at the specific fault, with a pass/fail gate 44 of 53 I hand-audited the nine it missed. Seven are provably unkillable, six of those being type annotations inside TYPE_CHECKING bloc...
A change nothing catches is a fault your tests cannot detect. Then I pointed an agent at the mutations the tests missed.
Open the app view to save this story, compare related coverage, and continue from the same source.