Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust
AI agents evolve rapidly but evaluation frameworks fail to keep pace with model changes
“building a demo with AI is really easy but making it production quality is really hard”
Braintrust field CTO argues that AI applications are now in a 'replatforming' era — models have made step-function improvements in tool use, long context, and code execution that render old application assumptions obsolete. The core problem: evals designed around a system's original model constraints become misleading when a new model drops in, creating invisible quality regressions. The talk sets up Braintrust's observability platform as the solution, though the transcript cuts off before the prescriptive content.