Towards Data Science iconTowards Data ScienceSep 24, 2026 ~7 min source read

Beyond RAGs: How to build AI that proves its own claims

Retrieval provides candidate evidence, but it is not proof. The article outlines a practical design for systems that convert model statements into inspectable claims, enforce a publication gate, and maintain traceable evidence bundles.

Share this story

Send the public story page.

Useful takeaways from this story.

Treat retrieved documents as candidates, not proof: every material claim needs a verifiable evidence link and a recorded support decision.

Replace document-level control with a claim ledger: record atomic claims, exact text spans, verbatim evidence spans, source snapshots, and metadata for traceability.

Enforce a non-negotiable publication gate: material claims without acceptable support must be revised, abstained, labeled as inference, or escalated to review.

The useful part

It can conflict with another source, be untrusted, or even maliciously injected. It's a wonderful technology and extremely useful – however, I'm arguing against treating retrieval as an oracle of truth. The original RAG paper showed why retrieval is valuable for knowledge intensive generation, and it also identified provenance as an open problem.

How it works

  • For any workflow where a reader may act on the narrative (and especially if that action may carry real stakes), I would rather ship a visible 'insufficient evidence' state than an elegant sentence with a...
  • However, when dealing with non-deterministic AI systems in environments where truthfulness is paramount, this human-in-the-loop design is necessary.
  • ALCE — short for Automatic Evaluation of Long-form Answers with Citations — is a research benchmark for question-answering systems that generate citations.
  • FRONT is a research approach that finds supporting quotes before generating the final answer.
  • However, we still need to define what support means in our own system and test against our own source distribution.

What to take from it

They determine what our precision, recall, coverage and abstention metrics will mean. I would start with a versioned, held-out test case: the request, the evidence snapshot, the expected claims, the support policy, permitted abstention behaviour and a risk tier. The post-RAG engineering objective thus has to be stronger than "RAG with citations." The real objective is an evidence-grounded narrative system.

Example or evidence

  • The ledger contains a list of all atomic claims that the system intends to publish, together with the evidence and controls attached to each proposition.
  • A truth-building architecture Here's the architecture to make this actually happen.
  • It can be incomplete: the answer contains a material claim with no adequate evidence.
  • These are research approaches, not a ready-made production threshold.

Details worth keeping

Retrieval gives a model access to candidates for evidence. It doesn't prove that a retrieved passage supports a claim it just made. A retrieved passage can be stale or incomplete.

Related coverage

  • Medium: Have you ever asked a Large Language Model (LLM) a specific question about your internal company policies, only for it to hallucinate or… Continue reading on Medium »
  • Dzone: Retrieval-augmented generation solved a real problem: it grounded LLM outputs in facts the model was never trained on.
  • Towards Data Science: RAG retrieves. Agents act. I built both separately, connected them explicitly, and ran the same nine tasks through all three systems.
  • Medium: A demo that nails every question in the interview room, then embarrasses itself in production a month later. Here's why that happens — and… Continue reading on Medium »

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app