The Hallway Track
Engineering Insights

Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust

AI Engineer · Aug 20, 2026 · Engineering Insights

AI agents evolve rapidly but evaluation frameworks fail to keep pace with model changes

“building a demo with AI is really easy but making it production quality is really hard”

Braintrust field CTO argues that AI applications are now in a 'replatforming' era — models have made step-function improvements in tool use, long context, and code execution that render old application assumptions obsolete. The core problem: evals designed around a system's original model constraints become misleading when a new model drops in, creating invisible quality regressions. The talk sets up Braintrust's observability platform as the solution, though the transcript cuts off before the prescriptive content.

evals observability agent-reliability model-upgrades production-ai

Watch / read the original source →