Evaluate AI agents systematically with Agent-EvalKit
AWS released Agent-EvalKit, an open-source toolkit for systematically evaluating AI agents' full execution paths.
“An agent might deliver a well-structured, actionable response while hallucinating, fabricating facts because its tools returned empty results.”
AWS launched Agent-EvalKit, an Apache 2.0 open-source toolkit that integrates with AI coding assistants (Claude Code, Kiro CLI, Kilo Code) to evaluate agents by tracing their full execution path rather than just final outputs. It addresses the gap that output-level testing misses—hallucinations over empty tool results and broken tool-call sequences—and ends in code-level improvement recommendations. This matters because robust agent evaluation infrastructure is becoming a key bottleneck as teams move agents into production.