# The problem in one example
# Why source diffs are limited for agents Traditional deterministic programs map source changes closely to runtime changes. Agent systems add moving parts: prompts, model versions and sampling, tool descriptions and schemas, retrieval content, external services, memory, conversation state, and orchestration/fallback policies. That makes identical code behave differently across runs, or tiny source edits produce large behavioral shifts.
Code review remains necessary, but it needs an extra layer: a behavior diff that answers questions reviewers care about: what path changed, where runs diverged, which steps or errors differ, and whether the implementation is sound.
# Capture two comparable runs before diffing A useful behavioral comparison requires a deliberate baseline and candidate:
- Baseline: same synthetic input, pinned or recorded configuration, known-good trace.
- Candidate: same synthetic input, intended code/prompt/model change, newly captured trace.
Control what you can and record what you cannot. If many variables change at once (input, retrieval corpus, model, tools), the diff may be accurate but hard to interpret.
# Read structural differences as evidence, not verdicts Behavioral diffs report added or removed steps, status differences, timing deltas, and warnings. For example, a diff might show that a plan step exists in the baseline but not the candidate, and a failing-step appears only in the candidate. That evidence doesn't tell you why. Possible causes include skipping planning, a step rename, changed instrumentation boundaries, a different branching choice, or an early termination.
Treat the behavioral diff as a starting point for causal investigation. Pair the observed divergence with the source diff and inspect the execution tree around the divergence to form a hypothesis.
# Focus diffs on the question you're asking AgentInspect supports scoped checks so you can narrow the comparison:
- Structure check to focus on added/removed steps and ignore status/timing noise.
- Timing check with a configurable duration threshold for performance-oriented reviews.
Begin with the broad human-readable diff, then narrow to test a specific hypothesis. JSON output, ignore-duration, and verbose modes support automation and CI integration.
# Timing differences need controlled interpretation A single run showing 120ms versus 70ms does not justify a general performance claim. Latency varies with network, cache state, provider load, token volume, and concurrency. A timing diff becomes meaningful when calls are stubbed deterministically, the difference is large relative to noise, representative runs repeat the pattern, and the structural diff explains the extra work. Choose duration thresholds based on the environment, not to force a test to pass.
AgentInspect's diff compares persisted traces. It does not replay or re-run agents. That determinism is useful for review because the comparison is repeatable for those two artifacts, but it also means representative capture and fixture control are essential. Treat the workflow as three actions: execute and capture evidence, compare traces to describe observed differences, then judge whether the change is acceptable.
# Practical review workflow
- 1Pin or stub inputs, models, and tools where possible. 2. Capture a known-good baseline trace. 3. Apply the candidate change and capture a new trace. 4. Run a behavior diff to locate the first divergence and structural differences. 5. Use focused checks (structure or timing) to test hypotheses. 6. Combine source diffs and execution evidence to reach a review decision.
This approach makes reviews more actionable by giving reviewers concrete runtime evidence tied back to the source changes.