AWS has published a reference design for evaluating multi-agent systems on more than fluent output. The example uses Amazon Bedrock AgentCore to test whether a supply-chain assistant selects appropriate tools, respects business constraints and explains the evidence behind its recommendations.

The fictional workflow has an orchestrator delegate work to optimization, distribution, routing and analytics agents. Built-in evaluators score broad qualities such as helpfulness, task completion and instruction following. Custom evaluators then check domain-specific properties including route feasibility, inventory grounding, SQL correctness, constraint satisfaction and whether trade-offs are explained.

AgentCore Evaluations supports an on-demand mode for development benchmarks, regression tests and CI/CD gates. The same custom evaluators can also run online against a sampled share of production traces, with scores sent to CloudWatch dashboards and alarms. AWS distinguishes this after-the-fact measurement from Bedrock Guardrails, which enforce content and grounding controls while a workflow runs.

The implementation is a reference architecture rather than evidence that the checks guarantee correct supply-chain decisions. Custom evaluators can encode incomplete assumptions or share weaknesses with the model they grade. Teams still need representative test cases, deterministic validation where possible, human review of consequential decisions and monitoring for failures the chosen metrics do not capture.