14 tracked signals on agent-evaluation.
SimulationMaxxing: How Nubank ships agents 20× faster with simulations — Shreya Rajpal, Snowglobe
AI Engineer · Jul 29, 2026
Nubank ships AI customer support agents 20x faster by generating eval data in simulation rather than waiting for production data.
“If you generate your eval data in sim instead of waiting on production data you can ship agents 20x faster and we'll give you evidence for that.”
How We Built an Agent That Improves Itself — Zubin Aysola, Weights & Biases
AI Engineer · Sep 26, 2026
Weights & Biases released Agent Arya, a self-improving research agent using evaluation-driven self-reinforcement.
“tests, assessments, agents, and how you configure them are covariant”
Introducing: LangSmith Tuned Evaluators
LangChain · Aug 18, 2026
LangChain launches tuned evaluators that beat frontier models on agent evals at lower cost
“our post-trained perceived error evaluator outperformed all frontier closed and open models, while still remaining the most cost-effective”
Evaluate AI agents systematically with Agent-EvalKit
AWS Machine Learning Blog · Jun 11, 2026
AWS released Agent-EvalKit, an open-source toolkit for systematically evaluating AI agents' full execution paths.
“An agent might deliver a well-structured, actionable response while hallucinating, fabricating facts because its tools returned empty results.”
The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI
AI Engineer · Jun 04, 2026
Our ability to measure AI agents in practice is falling behind their actual capabilities.
“our ability to actually measure these agents in practice that is falling behind of where the capabilities actually are”
The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse
AI Engineer · Oct 06, 2026
An OSS reference stack for self-improving agents combines online monitoring with offline evaluation cycles.
“I mean the transition from "I don't write prompts" to "I have cycles."”
From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI
AI Engineer · Oct 05, 2026
Arize AI presents a 101 framework for evaluating AI agents from tracing to LLM-as-judge meta-evaluation
How Do You Actually Evaluate an AI Agent?
LangChain · Aug 26, 2026
Agent evals must run continuously using LLM-as-judge since non-deterministic outputs make exact-match testing impossible.
“The teams that really get this right treat eval sort of as a muscle. You build it early, you're running it constantly because I think the alternative is like finding out your agent tried recommending the wrong product three weeks ago from a customer complaint.”
From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI
AI Engineer · Jul 25, 2026
Snorkel AI argues every company needs a custom benchmark built from production agent traces.
“It's not a static benchmark. It's a constantly populated data set from your production traces.”
Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
AI Engineer · Jul 24, 2026
Harbor framework treats all agent runs as RL rollouts, requiring empirical evaluation like ML models.
“Agentic coding is a form of machine learning. Generated code is best treated as a blackbox artifact whose behavior and generalization should be managed via empirical evaluation like with any ML model.”
Evaluating AI Agents: A production blueprint with Strands and AgentCore
AWS Machine Learning Blog · Jul 23, 2026
AWS and Motorway cut AI agent error rates from 1-in-8 to 1-in-50 with a production eval pipeline.
“The agent gives a confident-sounding response, but how do you prove it works reliably with real money on the line?”
EVA-Bench Data 2.0: 3 Domains, 121 Tools, 213 Scenarios
Hugging Face · Hugging Face Blog · Jun 04, 2026
Hugging Face expands EVA-Bench to 3 domains, 121 tools, and 213 scenarios for agent evaluation.
Evaluating Deep Agents using LangSmith on AWS
AWS Machine Learning Blog · May 28, 2026
LangSmith on AWS provides five evaluation patterns for deep agents in production
“Validating AI agent behavior before production is one of the hardest problems in applied AI.”
What is LangSmith?
LangChain · Sep 22, 2026
LangSmith is LangChain's framework-agnostic platform for tracing, testing, deploying, and monitoring LLM agents.
“Agents are a black box, making them difficult to observe and debug.”