Why skills matter and what they change
Skills encode domain-specific procedures—instructions, tool bindings, reference knowledge, workflows, and guardrails—so agents can load only the needed procedure at runtime. That keeps core agent instructions smaller, makes procedures portable across compatible agent harnesses, and lets teams update domain logic without retraining the underlying model.
Why standard output metrics aren't enough
A fluent, convincing response does not prove the agent picked the right skill or actually followed the skill's prescribed steps. Two failure modes are common:
- The agent invokes an inappropriate skill for the task (wrong routing).
- The agent invokes the correct skill but skips or partially follows the skill's steps (execution failure).
Both failures can produce plausible answers that pass ordinary output-quality checks but violate required procedures such as compliance checks, workflows, or business rules.
Strands Evals SDK and Amazon Bedrock AgentCore Evaluations introduce three skill-focused evaluators:
- Skill Selection Accuracy: A binary check that determines whether each invoked skill was appropriate for the task and whether the agent invoked the correct skill. This uses a prompt template and rubric stored with the evaluator.
- Skill Instruction Following: A graded check that measures how fully the agent followed the prescribed steps of a skill. It returns a five-level rating grounded in evidence for each required step, based on a rubric and prompt template.
- SkillInvoked (Strands-only): A deterministic, non-model check that confirms whether a named skill was loaded successfully.
- 1Record the agent run. The agent's decision and execution are captured either as a Strands Evals trajectory or an OpenTelemetry trace in your observability layer.
- 1Choose evaluators. Apply Skill Selection Accuracy to judge routing decisions and Skill Instruction Following to inspect step-by-step execution. Use SkillInvoked in Strands for deterministic checks.
Interpreting evaluator results and choosing fixes
- If Skill Selection Accuracy flags the invoked skill as inappropriate, the problem is routing. Fixes focus on the agent's selection logic or catalog-level guidance: change the skill catalog, adjust selection prompts, or refine decision heuristics.
- If Skill Instruction Following reports skipped or incomplete steps, the problem is execution. Fixes target skill content, guardrails, or how tools are bound: harden the step checks in the skill, add validation rules, or strengthen tool bindings and examples.
A hypothetical HR assistant has separate skills for PTO planning and employee benefits. If an employee asks about dental and vision benefits and the agent loads the benefits skill, routing is correct. If the agent loads the PTO skill instead, Skill Selection Accuracy isolates that routing error. If the agent correctly loads PTO but skips checking rollover rules when required, Skill Instruction Following identifies the skipped step. The two outcomes lead to different remediation paths.
Why per-skill evaluation matters in practice
Where these evaluations fit in a workflow
Use Strands Evals or AgentCore Evaluations as part of test suites, pre-production checks, or CI pipelines. Combine deterministic SkillInvoked checks for routing hygiene with model-based rubrics for selection and instruction following. The record-based approach (trajectories or traces) means you can evaluate recorded runs retrospectively and block regressions before they reach users.