# Clear recovery definitions When an AI agent causes an unacceptable external effect, saying you "rolled back" can hide important differences. Replacing a changed value with its previous setting fixes the current configuration but does not automatically restore anything that was lost or undo any side effects that already occurred. Use precise recovery labels and choose actions that meet an explicit recovery objective.
# The recovery taxonomy Use these operational categories instead of a generic "rollback":
- Reversal: directly undo the prior mutation when it's safe and nothing else depends on it. Example: removing a newly added firewall rule that no other change relies on.
- Compensation: perform a new, consequential action that corrects or mitigates the original effect when literal reversal is unsafe or impossible.
- Forward recovery: accept the original transition and move the system to a different acceptable state rather than restoring the old one.
- Containment: stop further harm without claiming to repair prior effects (for example, disable a compromised identity).
- Remediation: address downstream consequences after the technical state is corrected (notify recipients, invalidate exposed data).
- Irreversible: acknowledge the effect cannot be meaningfully undone using the system (delivered messages, permanently deleted data).
Each category establishes a different operational outcome. Pick the one that matches what the organization actually needs.
# Why compensation is not a free undo A compensating action is itself a privileged, consequential operation. Treat it like any other power change. Reversing or deleting things can break legitimate concurrent work. For example:
- Restoring an old snapshot can discard valid transactions that arrived after the snapshot.
- Revoking a role to undo an erroneous grant can also revoke access legitimately added afterward.
- Reapplying a previous configuration can remove recently approved changes.
Because compensation can cause new harm, it must be governed, constrained, and documented independently of the original agent action.
# What to bind to a compensating action When authorizing a compensation, include concrete items so the action is auditable and safe:
- The original action and the observed effect being addressed
- The approved recovery objective or target state
- The identity authorized to execute compensation and required preconditions
- The exact compensation operation and parameters
- Evidence requirements and idempotency expectations
- Timeout, retry behaviour, and escalation if compensation fails
- Any residual effects the compensation cannot repair
# Start with the recovery objective Before choosing a command, define the state the organization needs. In the backup-retention example, the objective might include: reestablish approved retention, determine what recovery points were lost, recreate required backups where possible, and preserve compliance evidence. Changing the retention number alone does not complete those objectives.
# Practical checklist for incident recovery
- Reconcile: prove whether the agent's external effect actually occurred.
- Classify: pick reversal, compensation, forward recovery, containment, remediation, or irreversible.
- Define objective: list target-state preconditions, evidence, and acceptable residuals.
- Authorize: bind explicit execution envelope and approvals for the compensation.
- Execute with safeguards: prefer idempotent operations, log evidence, and avoid blind inverse calls.
- Review aftermath: record what was lost, what remains exposed, and what further remediation is required.
# Short conclusion Recovery after an agent-caused incident is a decision about acceptable outcomes, not a single API call. Use precise language, require explicit authorization for compensating actions, and choose the recovery mechanism that actually produces the organization's required state instead of assuming a rollback recreates the past.