# What this is about
This brief explains the agentic testing life cycle described by TestMu AI. It reorganizes quality engineering (QE) for software that makes decisions and acts—agents that take natural-language instructions, call tools, change the world, and then report what they did.
# Why change testing for agents
# The QE loop re-derived for agents
# Six phases of the loop
- Discover: read the agent codebase and folder structure to learn what the agent can do. TestMu AI found that about 20 real agent repositories rarely included a usable agent manifest, so discovery treats the repo as the input.
- Profile: define environment, constraints, and assumptions to run the agent safely.
- Run: invoke the agent against generated scenarios and capture its actions and side effects.
- Report: emit a machine-readable JSON report with per-criterion verdicts and an assurance gap metric.
Phases are composable. Asking for a later step triggers planning of earlier steps so you can see cost and scope before you run expensive actions.
# The criterion as the unit of judgement
A criterion is a single, gradable claim (for example: "refund issued and recorded in ledger"). Each criterion is evaluated with three items: what was expected, what happened, and a quoted piece of evidence. The run also records criteria that could not be verified at all.
# The assurance gap and pass rates
# How CI gating changes
A finished test run exits 0 whether scenarios passed or failed. CI and release pipelines gate on the per-criterion verdicts in the JSON report. Exit 1 indicates that a run never started due to infrastructure problems, not that the agent failed functionally.
# Practical implication for teams
If there is no declared contract to test against, treat the codebase and filesystem as the input. Design tests to collect external evidence of actions and express failures at the criterion level. Track the assurance gap and use it as a design signal to improve observability and recording of agent activity.
# Where this shipping is today
TestMu AI provides this loop as Agent Assurance. The platform covers conversational agents in general availability, while autonomous execution tooling is at earlier stages of release.