This tutorial builds a controlled evaluation workflow in Python that tests whether an AI reviewer preserves and applies a declared policy across presentation variations. The goal is to detect context-sensitivity bugs: cases where the same substantive evidence leads to different decisions because of prompt wording, evidence order, or execution conditions.
Start by writing the requirement as a deterministic rule with explicit outcomes. The article's worked example is a two-node power-domain requirement with three cases and expected decisions:
- shared: both nodes in domain A -> expected decision: hold
- independent: domains A and B -> expected decision: approve
- unknown: one node domain unknown -> expected decision: insufficient_evidence
Treat unknown as an explicit value (Python None / JSON null) so graders don't invent interpretations. Include all outcomes so a grader that always approves or always holds will fail some cases.
Keep the answer key outside the application
Create variants that change only presentation elements: equivalent prompts, different orders of evidence, and alternate instruction phrasing. Do not change the task, the evidence values, or the policy while varying presentation. Serialize None as JSON null when sending prompts so the application sees the explicit unknown.
Before connecting a live model, run deterministic fixtures locally. Include a passing fixture (the expected behavior) and a deliberately defective fixture to ensure the harness detects regressions and known failures. Use only the standard library in the supplied reference implementation so it can run without external dependencies.
The harness defines an adapter interface for provider-specific integrations. When ready to test live systems, connect an approved application in read-only mode. Keep deployment credentials and mutation-capable tools out of the evaluation environment.
Preserve failures and analyze by condition
Record all failed attempts and configuration used for each trial. Analyze results by condition (prompt variant, evidence order, policy case) instead of relying only on aggregate scores. Use the suite as a regression check that flags when behavior changes under controlled variations, not as an authority that certifies model reliability.
The example harness focuses on a single, advisory decision and synthetic fixtures. It is a reference implementation for local demonstration and regression testing, not a comprehensive architecture review, retrieval benchmark, or production assurance product.
A defensible context-sensitivity harness requires: explicit policy definitions, externalized answer keys, controlled presentation variants, independent graders, preserved failures, and upfront validation. Treated as a regression tool, the harness makes prompt changes testable rather than anecdotal.