Testmuai iconTestmuaiSep 5, 2026 ~7 min source read

Salesforce’s Four Phases for Testing AI Agents, Explained

John Liang of Salesforce describes how their Service Cloud quality engineering team moved from output-only checks to multi-layered, agent-to-agent and scale-aware testing across four phases.

Salesforce's Four Phases of Agentic Testing [Testμ 2026]

Share this story

Send the public story page.

Useful takeaways from this story.

Test every internal step of an agent’s decision process—utterance, topic classification, plan, tool selection, and grounding—rather than relying on the final output alone.

Simulate high-volume voice traffic and evaluate a sampled subset (about 50–60 calls) in depth rather than making thousands of real calls.

# Overview Salesforce's John Liang lays out a practical progression for testing agentic systems. The core insight: traditional deterministic testing (known input, known output) still matters but is insufficient by itself. Agentic systems allow many valid paths and varied outcomes, so testing must expand to validate internal steps, simulate real-world conditions, and scale appropriately.

# The four phases, in brief Phase 1: Validate every step. Move validation inside the agent. Confirm what the customer actually said (utterance), verify topic classification, check the plan the agent chose, confirm the correct tools and actions were used, and ensure grounding to real data.

Phase 2: Close the lab-to-field gap. Laboratory scenarios miss real customer behaviors and industry-specific patterns. An airline example showed customers asking multiple questions in one message—something lab tests hadn't modeled—so internal testing missed failures that appeared in production.

Phase 3: Use agents to test agents, especially for voice. Real callers bring accents, background noise, interruptions, mixed requests, and emotional tone. Programmatic generation of that variety is impractical, so Salesforce used an AI agent to simulate realistic callers and run multi-turn conversations.

# Practical points about judgments and metrics

  • LLM-based judges are useful across validation layers but cannot replace deterministic checks. Assertions, trace validation, and metric retrieval remain necessary.

# Voice testing at scale Salesforce does not place thousands of real phone calls. They simulate the volume with virtual calls and then select roughly 50–60 of the simulated calls for deep evaluation. Those sampled calls are scored on accuracy, conciseness, completeness, tone, and voice quality.

# Why single-turn testing fails A single successful turn does not guarantee a resolved customer journey. A later turn can fail, and the overall outcome can be unresolved even if earlier steps looked correct. Tests must cover multi-turn interactions and state continuity across turns.

# Unexpected weaknesses discovered Even when agents were correctly grounded and prompted, they struggled to pick the single most relevant document among near-duplicate knowledge articles. Retrieval among similar documents proved a recurring failure mode.

# What changed for quality engineering Liang frames the shift as adapting a long-established engineering discipline to an environment with multiple valid paths and broader outcome distributions. The consequence of failure is customer trust loss rather than merely a functional bug.

# Short checklist for teams

  • Add step-level validation: utterance → topic → plan → tool → grounding.
  • Keep deterministic tests and add LLM-based judgment layers.
  • Simulate high volume and evaluate a representative sample deeply.
  • Package test tooling so customers or product teams can run them consistently.

# Bottom line Testing agents requires broader, layered validation that inspects internal decisions, mimics messy real-world interactions, and scales through simulation plus targeted human-grade evaluation. The strategy keeps deterministic testing while adding agentic and judgment layers to avoid deployed surprises.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app