# Summary Production telemetry often holds enough clues to explain a bug but not enough structure to run as an executable test. The engineering move is to transform that evidence into a deterministic regression test that reproduces the failure on the recorded revision and becomes passing evidence after a fix. Recent reproducer systems follow a staged approach that combines repository retrieval with runtime interaction instead of a single prompt-response test generation step.
Reduce telemetry before model or agent access. Request bodies, auth headers, customer identifiers, and database values are usually unnecessary for control-flow reproduction. Use collector processors to remove, redact, and transform attributes so sensitive or noisy fields never reach the agent.
# Reduce the failure to an executable context Raw traces are too broad. The agent input should be a compact slice containing:
- failing application frames and call path
- request shape and relevant inputs
- downstream interactions and the responses that triggered the defect
- nearby or related tests that show local conventions and setup
Exclude unrelated controllers, persistence code, and full payloads. The agent needs the smallest slice that still supplies the runtime cause and repository anchors. Example systems like ReProAgent structure the process into localization, root-cause analysis, test planning, and test generation rather than treating test creation as code completion.
# Produce an explicit test contract Make the agent's goal explicit: generate a single deterministic JUnit (or equivalent) regression test, do not change production code, and do not assert the observed bug as correct behavior. The last constraint prevents producing tests that would pass trivially on the buggy implementation by simply asserting the exception the system emitted.
# Generate against repository evidence
# Validate with execution feedback (fail-to-pass)
# Practical constraints and safeguards
# Bottom line Transforming incident evidence into regression tests requires a small, precise failure context, repository-anchored generation, and execution feedback that enforces fail-to-pass behavior. When those elements are combined, agents can supply deterministic tests that become living artifacts of the fix rather than another incident summary.