Ministryoftesting iconMinistryoftestingSep 23, 2026 ~2 min source read

Guardrails and evals people can trust

A practical summary of Maneesh Maddala’s guidance: build testable guardrails from real failures, make enforcement predictable, and design evals that match cost to evidence and measure repeatability.

Guardrails and evals people can trust

Share this story

Send the public story page.

Useful takeaways from this story.

Run cheap deterministic checks frequently and reserve judgment-based LLM grading for when evidence requires it.

Include adversarial attempts in evals and measure repeated runs—use pass@k for occasional success and pass^k to require consistent success.

# What this is about

# Practical distinctions A guardrail acts during execution to block or alter outputs that violate rules. An eval runs after the fact to measure whether the Skill meets its requirements. The guidance stresses a clear line between the two: pick a guardrail when you need runtime protection, pick an eval when you need measurement or diagnosis.

# How to write guardrails

  • Make it testable. As Maddala puts it, "A rule without a test is a suggestion. Pair every guardrail with an eval." That means express the guardrail so you can write an automated check against it.
  • Make enforcement predictable. If the same violation sometimes blocks and sometimes passes, people will stop trusting the rule. Predictable outcomes maintain developer confidence and enable reliable workflows.

# How to write evals

  • Match cost to evidence. Run deterministic checks on every commit. Use an LLM judge only when the answer genuinely needs human-like judgment. This preserves budget and keeps fast feedback cheap.
  • Include at least one attack attempt. Test at least one prompt injection or adversarial request to ensure guardrails hold against realistic misuse.

# Operational recommendations Three tiers of checking were mentioned as a useful model: cheap deterministic checks for every change, more expensive probabilistic or LLM-based checks selectively, and manual review only when automation can't decide. This prevents running the costly graders everywhere while still giving coverage where it matters.

# Concrete checklist to start

  • Convert guardrail statements into executable checks.
  • Instrument every guardrail with a paired eval to detect regressions.
  • Add at least one adversarial test case to eval suites.
  • Decide which checks run on every commit (deterministic) versus which run on a schedule or gate (LLM-judged).
  • Track pass@k and pass^k metrics separately and require pass^k for any guardrail intended to protect users at runtime.

# Final point Treat guardrails and evals like the test suites you already write for software. Apply the same discipline: express rules so they can be tested, keep enforcement consistent, and balance cost by running the right kind of check at the right cadence.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app