Salesforce’s Four Phases for Testing AI Agents, Explained
John Liang of Salesforce describes how their Service Cloud quality engineering team moved from output-only checks to multi-layered, agent-to-agent and scale-aware testing across four phases.
![Salesforce's Four Phases of Agentic Testing [Testμ 2026]](/api/proxy/image?url=https%3A%2F%2Fassets.testmuai.com%2Fresources%2Fimages%2Fmeta%2Fsalesforce-agentic-testing-phases.webp)
John Liang of Salesforce describes how their Service Cloud quality engineering team moved from output-only checks to multi-layered, agent-to-agent and scale-aware testing across four phases.
![Salesforce's Four Phases of Agentic Testing [Testμ 2026]](/api/proxy/image?url=https%3A%2F%2Fassets.testmuai.com%2Fresources%2Fimages%2Fmeta%2Fsalesforce-agentic-testing-phases.webp)
Test every internal step of an agent’s decision process—utterance, topic classification, plan, tool selection, and grounding—rather than relying on the final output alone.
Simulate high-volume voice traffic and evaluate a sampled subset (about 50–60 calls) in depth rather than making thousands of real calls.
# Overview Salesforce's John Liang lays out a practical progression for testing agentic systems. The core insight: traditional deterministic testing (known input, known output) still matters but is insufficient by itself. Agentic systems allow many valid paths and varied outcomes, so testing must expand to validate internal steps, simulate real-world conditions, and scale appropriately.
# The four phases, in brief Phase 1: Validate every step. Move validation inside the agent. Confirm what the customer actually said (utterance), verify topic classification, check the plan the agent chose, confirm the correct tools and actions were used, and ensure grounding to real data.
Phase 2: Close the lab-to-field gap. Laboratory scenarios miss real customer behaviors and industry-specific patterns. An airline example showed customers asking multiple questions in one message—something lab tests hadn't modeled—so internal testing missed failures that appeared in production.
Phase 3: Use agents to test agents, especially for voice. Real callers bring accents, background noise, interruptions, mixed requests, and emotional tone. Programmatic generation of that variety is impractical, so Salesforce used an AI agent to simulate realistic callers and run multi-turn conversations.
# Practical points about judgments and metrics
# Voice testing at scale Salesforce does not place thousands of real phone calls. They simulate the volume with virtual calls and then select roughly 50–60 of the simulated calls for deep evaluation. Those sampled calls are scored on accuracy, conciseness, completeness, tone, and voice quality.
# Why single-turn testing fails A single successful turn does not guarantee a resolved customer journey. A later turn can fail, and the overall outcome can be unresolved even if earlier steps looked correct. Tests must cover multi-turn interactions and state continuity across turns.
# Unexpected weaknesses discovered Even when agents were correctly grounded and prompted, they struggled to pick the single most relevant document among near-duplicate knowledge articles. Retrieval among similar documents proved a recurring failure mode.
# What changed for quality engineering Liang frames the shift as adapting a long-established engineering discipline to an environment with multiple valid paths and broader outcome distributions. The consequence of failure is customer trust loss rather than merely a functional bug.
# Short checklist for teams
# Bottom line Testing agents requires broader, layered validation that inspects internal decisions, mimics messy real-world interactions, and scales through simulation plus targeted human-grade evaluation. The strategy keeps deterministic testing while adding agentic and judgment layers to avoid deployed surprises.

The agentic testing life cycle points the QE loop at the agent itself. See the six phases, what a verdict proves, and the gap between tested and verified.
![Redefining Quality Leadership in an Agentic World [Testμ 2026]](/api/proxy/image?url=https%3A%2F%2Fassets.testmuai.com%2Fresources%2Fimages%2Fmeta%2Fredefining-quality-leadership-agentic-world.webp)
Sobhitha Neelanath of Salesforce on why a 94.2% pass rate missed a $400,000 refund loop, and the trust equation quality leaders should measure instead.
![Agentic AI and the Next Decade of Quality Engineering [Testμ 2026]](/api/proxy/image?url=https%3A%2F%2Fassets.testmuai.com%2Fresources%2Fimages%2Fmeta%2Fai-intern-quality-engineering.webp)
Mallika Fernandes of Accenture on the five stages of AI reliance, why code got cheap, and the Concorde lesson waiting for teams that ignore token economics.
![What a QE Org Will Look Like in 2027 [Testμ 2026]](/api/proxy/image?url=https%3A%2F%2Fassets.testmuai.com%2Fresources%2Fimages%2Fmeta%2Fquality-engineering-org-2027.webp)
A Testμ 2026 panel with CITY Furniture, Vail Resorts and TestMu AI on pooled QA teams, change failure rate, and why everyone becomes a QA architect.
![Testing LLM Responses You Cannot Predict [Testμ 2026]](/api/proxy/image?url=https%3A%2F%2Fassets.testmuai.com%2Fresources%2Fimages%2Fmeta%2Ftesting-llm-responses.webp)
Gil Zilberfeld of TestinGil on golden data sets, scorecards and sanity tests, and why fixing a bug inside a prompt is only the start of fixing it.

At Dreamforce 2026, Salesforce presented its most complete vision yet for the “agentic enterprise.” It shifted the core focus away from individual apps (such as those focused on sales, services, or marketing) and unified them through a single architecture built around data, logic and AI agents. This vision includes: (1
Loading more related stories...
Open the app view to save this story, compare related coverage, and continue from the same source.