An AI agent can finish a run cleanly and still be wrong. It may call a tool multiple times, produce a plausible but incorrect answer, or use different internal steps on repeated runs. Health checks that report "up" and dashboards built around requests-per-second miss these failure modes because errors aren't thrown and request counts don't reflect token usage or multi-step chains.
# What to capture in logs
# Distributed tracing and trace waterfalls
Distributed traces turn individual log events into an ordered tree of spans. Instrument model calls, retrievals, and tool invocations as spans. A trace waterfall shows nesting and timing: you can see which model call consumed most time or tokens, whether a tool was called twice with slightly different arguments, and how side effects flowed through the run. OpenTelemetry-style tracing automatically propagates trace IDs to logs, which is key for joining traces and logs into a single forensic record.
# Token metrics and cost visibility
Token counts matter more than request counts. Measure tokens consumed per model call and aggregate tokens by run. Track token costs as metrics so you spot runs that were unusually expensive or slow because they consumed far more tokens than typical. Correlate token metrics with trace spans to find where the cost concentrated.
# Practical debugging workflow
- 1Instrumentation: Add structured logging and tracing to model calls, tools, and retrievals.
- 2Reconstruct: Open the trace waterfall for a suspect run, follow spans, and read the arguments and responses tied to each span.
- 3Diagnose: Look for duplicated tool calls, divergent retrieval results, or a model response that ignored an earlier tool result.
- 4Remediate: Fix tool argument handling, add checks on tool results, or adjust prompt/context logic.
Recording seeds or other reproducibility metadata can help when you want to rerun a model call with the same settings, but a single rerun rarely proves the distribution of behavior across production runs.
# Privacy, compliance, and redaction
Prompts and retrievals often include personal or confidential data. Naive logging of full prompt text creates compliance risks. Instead, log structured metadata and redact or hash sensitive fields. Where full content is required for debugging, use access controls and retention policies that limit exposure.
# Concrete instrumentation goals
- Tie every log line to a trace ID.
- Capture model parameters (model name, temperature), token counts, and finish reason.
- Instrument tool calls and retrievals as spans so they appear in the same waterfall.
- Emit metrics for token consumption, latency per span, and tool error rates.
# How this changes incident response
When a customer reports a wrong answer, the team should open the trace for that run and see the exact sequence: which retrievals happened, which tools returned which results, and which model outputs were used to form the final answer. That replaces guesswork and re-running prompts with a forensic reconstruction of the original run.
# Bottom line