Towards Data Science iconTowards Data ScienceSep 24, 2026 ~8 min source read

When the Correct Answer Is Nothing: How Reliability Layers Can Produce Confidently Wrong Outputs

Required fields, similarity caches, and binary thresholds remove the option to abstain. That removal can let pipelines return plausible but incorrect values for cases where the true answer is ‘no value’ or ‘different question’.

Share this story

Send the public story page.

Useful takeaways from this story.

I recognised it immediately, because I had just published the same bug wearing different clothes.

What the cache measurement actually showed That embedding models handle negation badly is documented.

The same failure, a different component I maintain a small on-premise retrieval system as an open-source reference implementation: documents that can't leave the perimeter, no GPU, an answer costing tens of...

The useful part

I recognised it immediately, because I had just published the same bug wearing different clothes. The same failure, a different component I maintain a small on-premise retrieval system as an open-source reference implementation: documents that can't leave the perimeter, no GPU, an answer costing tens of seconds on CPU. Repeated questions about the same regulations are the obvious case for a cache, so I added a semantic one.

How it works

  • Reliability mechanisms work by removing degrees of freedom Both of these components were added to make a system more reliable, and both succeeded at what they were added for.
  • The user sees a fluent, sourced, confident answer to a question they didn't ask.
  • What the cache measurement actually showed That embedding models handle negation badly is documented.
  • The same nine negation pairs scored by bge-m3, against two different control populations.
  • The 0.3556 doesn't reach significance on an exact permutation test (p=0.4376), so the honest reading of the top panel is that the score carries no usable information — not that it's reliably inverted.

What to take from it

It passes every gate the pipeline has, because every gate was built to check shape and the problem is provenance. I want to note two things about that second pattern, not to score points but because they're the kind of gap that survives review — including mine. If two questions mean the same thing, serve the stored answer and skip generation entirely.

Example or evidence

  • The ordering is load-bearing, and the published code doesn't have it.
  • It's about what the cache does when the correct answer is none of the above.
  • Constraint violations are loud by construction: a malformed response fails a parser, a type error fails a check, a missing key throws.
  • A required field that gets filled leaves a record you can inspect afterwards — Nweke found his by putting source messages next to extracted rows.

Details worth keeping

He had turned on Structured Outputs for a pipeline that parsed payment confirmation messages into transaction records, and his reconciliation job started flagging a small, steady stream of mismatches — a couple of percent of a week's volume. The schema declared transaction_date as required. The only thing wrong was that a value existed where no value should have.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app