Testmuai iconTestmuaiSep 10, 2026 ~8 min source read

Agentic Testing Life Cycle: Re-centering QE on Autonomous Agents

A practical summary of a quality-engineering loop designed for agents that act on the world: six phases, how verdicts are formed from evidence, and the assurance gap that separates tested from verified.

Agentic Testing Life Cycle: The QE Loop for AI Agents

Share this story

Send the public story page.

Useful takeaways from this story.

Treat the agent itself as the test target: read its codebase and derive scenarios rather than assuming a contract exists.

The QE loop has six phases—Discover, Generate, Profile, Run, Judge, Report—and runs can be composed so later steps plan earlier ones.

# What this is about

This brief explains the agentic testing life cycle described by TestMu AI. It reorganizes quality engineering (QE) for software that makes decisions and acts—agents that take natural-language instructions, call tools, change the world, and then report what they did.

# Why change testing for agents

# The QE loop re-derived for agents

# Six phases of the loop

  • Discover: read the agent codebase and folder structure to learn what the agent can do. TestMu AI found that about 20 real agent repositories rarely included a usable agent manifest, so discovery treats the repo as the input.
  • Profile: define environment, constraints, and assumptions to run the agent safely.
  • Run: invoke the agent against generated scenarios and capture its actions and side effects.
  • Report: emit a machine-readable JSON report with per-criterion verdicts and an assurance gap metric.

Phases are composable. Asking for a later step triggers planning of earlier steps so you can see cost and scope before you run expensive actions.

# The criterion as the unit of judgement

A criterion is a single, gradable claim (for example: "refund issued and recorded in ledger"). Each criterion is evaluated with three items: what was expected, what happened, and a quoted piece of evidence. The run also records criteria that could not be verified at all.

# The assurance gap and pass rates

# How CI gating changes

A finished test run exits 0 whether scenarios passed or failed. CI and release pipelines gate on the per-criterion verdicts in the JSON report. Exit 1 indicates that a run never started due to infrastructure problems, not that the agent failed functionally.

# Practical implication for teams

If there is no declared contract to test against, treat the codebase and filesystem as the input. Design tests to collect external evidence of actions and express failures at the criterion level. Track the assurance gap and use it as a design signal to improve observability and recording of agent activity.

# Where this shipping is today

TestMu AI provides this loop as Agent Assurance. The platform covers conversational agents in general availability, while autonomous execution tooling is at earlier stages of release.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app