Amazon iconAmazonSep 22, 2026 ~6 min source read

How to evaluate skill-equipped agents using Strands Evals and Amazon Bedrock AgentCore

Skills package domain procedures for agents, but fluent output can hide routing and execution errors. Strands Evals and Amazon Bedrock AgentCore Evaluations add skill-focused checks to measure whether the agent chose the right skill and followed its steps.

Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore

Share this story

Send the public story page.

Useful takeaways from this story.

Strands Evals and AgentCore Evaluations provide three evaluators: Skill Selection Accuracy, Skill Instruction Following, and SkillInvoked (deterministic check on Strands).

Use recorded trajectories (Strands) or OpenTelemetry traces (AgentCore) and the AgentCore CLI to run per-skill evaluations and choose fixes based on whether the issue is routing or execution.

Why skills matter and what they change

Skills encode domain-specific procedures—instructions, tool bindings, reference knowledge, workflows, and guardrails—so agents can load only the needed procedure at runtime. That keeps core agent instructions smaller, makes procedures portable across compatible agent harnesses, and lets teams update domain logic without retraining the underlying model.

Why standard output metrics aren't enough

A fluent, convincing response does not prove the agent picked the right skill or actually followed the skill's prescribed steps. Two failure modes are common:

  • The agent invokes an inappropriate skill for the task (wrong routing).
  • The agent invokes the correct skill but skips or partially follows the skill's steps (execution failure).

Both failures can produce plausible answers that pass ordinary output-quality checks but violate required procedures such as compliance checks, workflows, or business rules.

Strands Evals SDK and Amazon Bedrock AgentCore Evaluations introduce three skill-focused evaluators:

  • Skill Selection Accuracy: A binary check that determines whether each invoked skill was appropriate for the task and whether the agent invoked the correct skill. This uses a prompt template and rubric stored with the evaluator.
  • Skill Instruction Following: A graded check that measures how fully the agent followed the prescribed steps of a skill. It returns a five-level rating grounded in evidence for each required step, based on a rubric and prompt template.
  • SkillInvoked (Strands-only): A deterministic, non-model check that confirms whether a named skill was loaded successfully.
  1. Record the agent run. The agent's decision and execution are captured either as a Strands Evals trajectory or an OpenTelemetry trace in your observability layer.
  1. Choose evaluators. Apply Skill Selection Accuracy to judge routing decisions and Skill Instruction Following to inspect step-by-step execution. Use SkillInvoked in Strands for deterministic checks.

Interpreting evaluator results and choosing fixes

  • If Skill Selection Accuracy flags the invoked skill as inappropriate, the problem is routing. Fixes focus on the agent's selection logic or catalog-level guidance: change the skill catalog, adjust selection prompts, or refine decision heuristics.
  • If Skill Instruction Following reports skipped or incomplete steps, the problem is execution. Fixes target skill content, guardrails, or how tools are bound: harden the step checks in the skill, add validation rules, or strengthen tool bindings and examples.

A hypothetical HR assistant has separate skills for PTO planning and employee benefits. If an employee asks about dental and vision benefits and the agent loads the benefits skill, routing is correct. If the agent loads the PTO skill instead, Skill Selection Accuracy isolates that routing error. If the agent correctly loads PTO but skips checking rollover rules when required, Skill Instruction Following identifies the skipped step. The two outcomes lead to different remediation paths.

Why per-skill evaluation matters in practice

Where these evaluations fit in a workflow

Use Strands Evals or AgentCore Evaluations as part of test suites, pre-production checks, or CI pipelines. Combine deterministic SkillInvoked checks for routing hygiene with model-based rubrics for selection and instruction following. The record-based approach (trajectories or traces) means you can evaluate recorded runs retrospectively and block regressions before they reach users.

More context around this story.

Migrate agentic workloads to Amazon Bedrock AgentCore
Amazon iconAmazonSep 3, 2026

Migrate agentic workloads to Amazon Bedrock AgentCore

An agent that works in a notebook is not an agent in production. This post walks through migrating a LangGraph customer support agent to Amazon Bedrock AgentCore in two stages: onto Runtime, Gateway, and Memory, then to model-driven planning on Strands Agents, retiring operational burdens along the way.

Improving HCLS AI reasoning with open-source agent skills
Amazon iconAmazonSep 16, 2026

Improving HCLS AI reasoning with open-source agent skills

AI agents on foundation models often misapply healthcare and life sciences decision frameworks, citing the right guideline but applying it incorrectly. This post shares 38 open-source agent skills across 11 HCLS domains that close this gap, with installation steps, three worked use cases, and a 410-prompt evaluation sh

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app