AWS has documented a reference implementation for evaluating a Strands-based supply-chain multi-agent system with Amazon Bedrock AgentCore Evaluations. The design gives teams a way to assess tool use, operational correctness and explanation quality alongside a response’s helpfulness, which matters for planners relying on agent recommendations.

The post is an implementation guide, not a product launch or benchmark. It uses the fictitious AnyCompany Retail to demonstrate a system in which an orchestrator delegates requests to four specialist agents: optimization, distribution, routing and analytics. The example is designed for inventory allocation, distribution adjustments, routing scenarios and supply-chain diagnostics.

Reference architecture

In AWS’s design, the orchestrator receives a planner request and invokes specialist agents exposed as tools. The optimization agent uses MCP tools backed by mock API Gateway REST interfaces; the distribution agent calls recommendation APIs; the routing agent calls logistics APIs; and the analytics agent answers diagnostics questions. The agents run on AgentCore runtime with memory and observability enabled, while Amazon Bedrock foundation models run the agent loop.

The distinction is important because an agentic workflow can fail even when its prose sounds plausible. AWS frames correctness as including the selection of tools, workflow execution, compliance with business constraints and grounding in data. The guide also separates evaluation after execution from Amazon Bedrock Guardrails, which it describes as execution-time controls such as content filtering, denied-topic detection and grounding validation.

Three evaluation layers

The first layer uses built-in evaluators without additional setup. Helpfulness is the common baseline, supplemented by Tool Selection Accuracy for the orchestrator, Response Relevance for optimization and distribution, Instruction Following for routing, and Faithfulness for analytics.

The second layer adds custom checks for operational validity. AWS proposes plan-coherence and tool-trajectory evaluators for the orchestrator; constraint satisfaction and KPI attainment for optimization; groundedness and risk-impact checks for distribution; route feasibility and service-level checks for routing; and SQL correctness, data grounding and unsupported-claim checks for analytics.

For the optimization example, the constraint evaluator checks remaining budget, warehouse capacity and inventory coverage. Recommended inventory must meet forecast demand while remaining no higher than 2× demand. This makes the example more than a test of response wording: it evaluates whether a stocking recommendation can meet stated business rules.

Explainability as a separate test

AWS makes explainability a third, independent layer rather than treating it as a side effect of accuracy. Its six cross-cutting evaluators test decision-rationale quality, evidence attribution, constraint reasoning, trade-off explanation, tool-use explainability and assumption disclosure.

That separation allows teams to identify two different failure modes. A recommendation might satisfy a custom business evaluator yet fail to state the data, constraints or trade-offs behind it. Conversely, it might offer a persuasive explanation while violating constraints. AWS says those outcomes call for different remediation: improving how an agent communicates in the first case and correcting its decision logic in the second.

The optimization walkthrough illustrates the split. It first runs a custom constraint evaluator with built-in Helpfulness and Response Relevance, then applies Decision Rationale Quality and Constraint Reasoning to the same session. AWS’s illustrative rationale cites a 1,200-unit demand forecast, a 1.25× safety factor and a 1,500-unit recommendation; a separate example describes a budget allowing 1,800 units while warehouse capacity limits the choice to 1,600 units. These are examples of the evidence and trade-offs the evaluators are intended to inspect, not reported system results.

Testing and deployment path

The supplied test client runs 20 sample queries—five for each sub-agent—across four multi-turn sessions. It can instead target individual categories: an optimization-only run uses one session of five turns, while a routing-and-analytics run uses two sessions of five turns each. The client prints session IDs, which are then used to run evaluations against the resulting traces.

Custom evaluators are registered through the supplied evaluator API, and an evaluation run accepts at least one custom or built-in evaluator ID. The API returns HTTP 202 immediately; AWS says the evaluation runs asynchronously and saves Markdown results to S3. Before launching evaluations, the guide advises waiting 3–5 minutes for traces to propagate to CloudWatch.

The implementation requires the AWS CLI, AWS SAM CLI v1.100.0 or later, Docker v20.x or later, Node.js v18.x or later, and Python v3.11 or later. Its packaged dependencies include strands-agents, strands-agents-tools, requests, bedrock-agentcore and boto3. Deployment uses Terraform after configuring a VPC ID and runtime subnet availability zones; AWS advises running terraform destroy afterward to avoid recurring charges.

Development and production modes

AWS positions on-demand evaluation for development benchmarking, regression testing and CI/CD gates. Its online mode is intended for production monitoring and alerts: the same custom evaluators can be referenced through an OnlineEvaluationConfig with evaluator ARNs, optional session filters and a sampling rate. AWS gives 1–10% of production traces as an example range; AgentCore reads observability traces, scores them, and streams results to CloudWatch dashboards and alarms.

The guide provides an architectural pattern rather than comparative performance evidence. Its retailer is fictional, and the optimization interfaces are mock REST APIs, so the post demonstrates how to structure evaluation rather than proving outcomes for a live supply-chain deployment. Source: AWS Machine Learning Blog.

Definition. The framework evaluates an agent system’s helpfulness and task execution, operational correctness against custom rules, and the quality of its explanations as separate concerns.

Evaluation layerWhat it evaluates
Built-in evaluationHelpfulness and agent-specific measures such as tool selection, relevance, instruction following and faithfulness.
Custom operational evaluationBusiness-rule validity, including constraints, KPIs, feasibility, SQL correctness and grounding.
Explainability evaluationDecision rationale, evidence attribution, constraint reasoning, trade-offs, tool use and assumptions.

Key takeaways

  • Built-in evaluators provide a baseline for helpfulness, tool selection, relevance, instruction following and faithfulness.
  • Custom evaluators test operational requirements such as constraints, KPI attainment, route feasibility, SQL correctness and data grounding.
  • Six cross-cutting explainability evaluators assess rationale, evidence, constraints, trade-offs, tool use and disclosed assumptions.
  • The example tests whether inventory recommendations meet forecast demand without exceeding twice that demand.
  • On-demand evaluations support development and CI/CD use cases, while online evaluation supports production monitoring.
  • The guide presents an architectural pattern using fictional data and mock APIs, not comparative performance evidence.

FAQ

What are the three evaluation layers in AWS’s framework?

The layers cover built-in response-quality evaluation, custom operational-validity checks and independent explainability evaluation.

Why does AWS evaluate explainability separately?

A recommendation can be operationally correct but poorly explained, or persuasive in wording while violating business constraints; the failures require different remediation.

What does the optimization constraint evaluator check?

It checks remaining budget, warehouse capacity and inventory coverage, including that recommended inventory meets forecast demand and does not exceed twice demand.

How does AWS position on-demand and online evaluations?

AWS positions on-demand evaluation for development benchmarking, regression testing and CI/CD gates, and online evaluation for production monitoring and alerts.

Sources