Guardrails and evals people can trust
A practical summary of Maneesh Maddala’s guidance: build testable guardrails from real failures, make enforcement predictable, and design evals that match cost to evidence and measure repeatability.
A practical summary of Maneesh Maddala’s guidance: build testable guardrails from real failures, make enforcement predictable, and design evals that match cost to evidence and measure repeatability.
Run cheap deterministic checks frequently and reserve judgment-based LLM grading for when evidence requires it.
Include adversarial attempts in evals and measure repeated runs—use pass@k for occasional success and pass^k to require consistent success.
# What this is about
# Practical distinctions A guardrail acts during execution to block or alter outputs that violate rules. An eval runs after the fact to measure whether the Skill meets its requirements. The guidance stresses a clear line between the two: pick a guardrail when you need runtime protection, pick an eval when you need measurement or diagnosis.
# How to write guardrails
# How to write evals
# Operational recommendations Three tiers of checking were mentioned as a useful model: cheap deterministic checks for every change, more expensive probabilistic or LLM-based checks selectively, and manual review only when automation can't decide. This prevents running the costly graders everywhere while still giving coverage where it matters.
# Concrete checklist to start
# Final point Treat guardrails and evals like the test suites you already write for software. Apply the same discipline: express rules so they can be tested, keep enforcement consistent, and balance cost by running the right kind of check at the right cadence.
![Build Trustworthy AI Agents Powered by Evals [Testμ 2026]](/api/proxy/image?url=https%3A%2F%2Fassets.testmuai.com%2Fresources%2Fimages%2Fmeta%2Fbuild-trustworthy-ai-agents.webp)
Rushabh Mehta of Meta on agent evals: idempotency keys, checkpointing, memory TTLs, the three grader types, and why GAIA 2 shows temporal awareness still fails.

AI guardrails explained: input, output and tool-call rails, how Guardrails AI, NeMo Guardrails and Llama Guard differ, and how to test that each rail holds.
Loading more related stories...
Open the app view to save this story, compare related coverage, and continue from the same source.