Machinelearningmastery iconMachinelearningmasteryOct 1, 2026 ~7 min source read

AI Agent Observability: How to Log, Trace, and Debug Agent Runs

Practical overview of why traditional monitoring fails for agentic systems and how structured logging, distributed tracing, and debugging workflows let you reconstruct what an agent actually did.

AI Agent Observability: Logging, Tracing, and Debugging Explained

Share this story

Send the public story page.

Useful takeaways from this story.

Capture every model call, tool execution, token usage, and reasoning step as structured data and tie each event to a run-specific trace ID.

Use OpenTelemetry-style distributed tracing and trace waterfalls to see nested model and tool calls, token costs, and timing for forensic reconstruction.

An AI agent can finish a run cleanly and still be wrong. It may call a tool multiple times, produce a plausible but incorrect answer, or use different internal steps on repeated runs. Health checks that report "up" and dashboards built around requests-per-second miss these failure modes because errors aren't thrown and request counts don't reflect token usage or multi-step chains.

# What to capture in logs

# Distributed tracing and trace waterfalls

Distributed traces turn individual log events into an ordered tree of spans. Instrument model calls, retrievals, and tool invocations as spans. A trace waterfall shows nesting and timing: you can see which model call consumed most time or tokens, whether a tool was called twice with slightly different arguments, and how side effects flowed through the run. OpenTelemetry-style tracing automatically propagates trace IDs to logs, which is key for joining traces and logs into a single forensic record.

# Token metrics and cost visibility

Token counts matter more than request counts. Measure tokens consumed per model call and aggregate tokens by run. Track token costs as metrics so you spot runs that were unusually expensive or slow because they consumed far more tokens than typical. Correlate token metrics with trace spans to find where the cost concentrated.

# Practical debugging workflow

  1. Instrumentation: Add structured logging and tracing to model calls, tools, and retrievals.
  2. Reconstruct: Open the trace waterfall for a suspect run, follow spans, and read the arguments and responses tied to each span.
  3. Diagnose: Look for duplicated tool calls, divergent retrieval results, or a model response that ignored an earlier tool result.
  4. Remediate: Fix tool argument handling, add checks on tool results, or adjust prompt/context logic.

Recording seeds or other reproducibility metadata can help when you want to rerun a model call with the same settings, but a single rerun rarely proves the distribution of behavior across production runs.

# Privacy, compliance, and redaction

Prompts and retrievals often include personal or confidential data. Naive logging of full prompt text creates compliance risks. Instead, log structured metadata and redact or hash sensitive fields. Where full content is required for debugging, use access controls and retention policies that limit exposure.

# Concrete instrumentation goals

  • Tie every log line to a trace ID.
  • Capture model parameters (model name, temperature), token counts, and finish reason.
  • Instrument tool calls and retrievals as spans so they appear in the same waterfall.
  • Emit metrics for token consumption, latency per span, and tool error rates.

# How this changes incident response

When a customer reports a wrong answer, the team should open the trace for that run and see the exact sequence: which retrievals happened, which tools returned which results, and which model outputs were used to form the final answer. That replaces guesswork and re-running prompts with a forensic reconstruction of the original run.

# Bottom line

More context around this story.

Graham Dumpleton: Zero-code tracing with wrapture
Grahamdumpleton iconGrahamdumpletonSep 8, 2026

Graham Dumpleton: Zero-code tracing with wrapture

The previous post traced the shop with three bindings and a sink, all applied from the program's own entry point. That is fine when the program is yours. It is less fine when the application is one you inherited and would rather not touch, when someone else owns the deployment, or when you simply do not want observatio

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app