Dev iconDevSep 21, 2026 ~1 min source read

Your LLM Telemetry Table Does Not Have One Denominator

I had an LLM telemetry report with the usual reassuring furniture: medians, interquartile ranges, sample sizes, and bootstrap intervals. Mixed-model threads could contribute to one table but fail the purity rule for another.

Your LLM Telemetry Table Does Not Have One Denominator

Share this story

Send the public story page.

Useful takeaways from this story.

I had an LLM telemetry report with the usual reassuring furniture: medians, interquartile ranges, sample sizes, and bootstrap intervals.

Mixed-model threads could contribute to one table but fail the purity rule for another.

It was actually several different studies sharing a table.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

I had an LLM telemetry report with the usual reassuring furniture: medians, interquartile ranges, sample sizes, and bootstrap intervals. Mixed-model threads could contribute to one table but fail the purity rule for another. Historical routing policy was often unknown, and task family was not observed at all.

How it works

  • Tool-error rate, re-edit rate, validation runs, recovery sequences, and output tokens therefore described epoch-attributed portions of work.
  • It was treating every n beside a model label as if it counted the same kind of thing.
  • For readers auditing an agent harness, hexisteme/hard-gate-hooks contains two MIT-licensed Stop-hook examples, their tests, and a read-only scanner.
  • They are adjacent implementation examples, not the telemetry instrument described here.
  • A model column is not an analysis unit For the core metrics, attribution happened inside a thread.

What to take from it

A multi-model thread could produce separate epoch rows because each assistant turn was assigned to the model epoch that produced it. The dangerous mistake was no longer simply calling an association causal. It was computed once per thread, only for main threads with model purity at or above 0.9, and censored threads were excluded.

Example or evidence

Details worth keeping

It was actually several different studies sharing a table. The core process metrics were attributed to model epochs inside threads. The completion proxy existed only at thread level.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app