AWS has introduced Agent-EvalKit, an open-source toolkit for evaluating AI agents systematically.
The toolkit walks through six evaluation phases and integrates with AI coding assistants including Claude Code, Kiro CLI, and Kilo Code. Its example uses a travel research agent built with Strands Agents SDK and Amazon Bedrock.
The release fits a broader shift: as agents move into real workflows, teams need repeatable evaluation infrastructure rather than ad hoc demos and vibes-based testing.