Digitalthoughtdisruption iconDigitalthoughtdisruptionSep 22, 2026 ~6 min source read

How to Build a Context-Sensitivity Test Harness for AI in Python

A practical walkthrough that shows how to build a provider-neutral Python harness that varies presentation while holding task, evidence, and policy constant, so you can detect when context changes alter decisions.

Share this story

Send the public story page.

Useful takeaways from this story.

Define the policy and expected outcomes before creating prompt or evidence variants so graders can’t silently reinterpret unknowns.

Generate controlled prompt and evidence variants (including evidence-order changes) and preserve failing attempts and configuration for postmortem analysis.

Validate the harness locally with passing and deliberately defective fixtures before connecting a live model or application through a read-only adapter.

This tutorial builds a controlled evaluation workflow in Python that tests whether an AI reviewer preserves and applies a declared policy across presentation variations. The goal is to detect context-sensitivity bugs: cases where the same substantive evidence leads to different decisions because of prompt wording, evidence order, or execution conditions.

Start by writing the requirement as a deterministic rule with explicit outcomes. The article's worked example is a two-node power-domain requirement with three cases and expected decisions:

  • shared: both nodes in domain A -> expected decision: hold
  • independent: domains A and B -> expected decision: approve
  • unknown: one node domain unknown -> expected decision: insufficient_evidence

Treat unknown as an explicit value (Python None / JSON null) so graders don't invent interpretations. Include all outcomes so a grader that always approves or always holds will fail some cases.

Keep the answer key outside the application

Create variants that change only presentation elements: equivalent prompts, different orders of evidence, and alternate instruction phrasing. Do not change the task, the evidence values, or the policy while varying presentation. Serialize None as JSON null when sending prompts so the application sees the explicit unknown.

Before connecting a live model, run deterministic fixtures locally. Include a passing fixture (the expected behavior) and a deliberately defective fixture to ensure the harness detects regressions and known failures. Use only the standard library in the supplied reference implementation so it can run without external dependencies.

The harness defines an adapter interface for provider-specific integrations. When ready to test live systems, connect an approved application in read-only mode. Keep deployment credentials and mutation-capable tools out of the evaluation environment.

Preserve failures and analyze by condition

Record all failed attempts and configuration used for each trial. Analyze results by condition (prompt variant, evidence order, policy case) instead of relying only on aggregate scores. Use the suite as a regression check that flags when behavior changes under controlled variations, not as an authority that certifies model reliability.

The example harness focuses on a single, advisory decision and synthetic fixtures. It is a reference implementation for local demonstration and regression testing, not a comprehensive architecture review, retrieval benchmark, or production assurance product.

A defensible context-sensitivity harness requires: explicit policy definitions, externalized answer keys, controlled presentation variants, independent graders, preserved failures, and upfront validation. Treated as a regression tool, the harness makes prompt changes testable rather than anecdotal.

More context around this story.

69 Tests. All Passing. Zero Bugs Caught.
Dev iconDevSep 17, 2026

69 Tests. All Passing. Zero Bugs Caught.

An AI model wrote 69 tests for a Python module. Every one passed. Together they caught zero of the eleven bugs I had deliberately planted in that module. A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten. That contrast is the whole project. What I built A mutation-t

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app