From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI
Snorkel AI argues every company needs a custom benchmark built from production agent traces.
“It's not a static benchmark. It's a constantly populated data set from your production traces.”
Rustem Feyzkhanov of Snorkel AI argues that agent evaluation must evolve beyond analyzing production traces to building offline simulation environments that replay those traces as repeatable experiments. The core insight is that public benchmarks are too domain-generic and A/B testing in production lacks reproducibility, so companies should continuously populate private benchmarks from their own traces. Snorkel AI runs millions of such simulations monthly and frames benchmark construction as an engineering discipline.