How Lyft Builds Evals That Actually Matter in Production | Interrupt 26
Lyft scaled to seven+ production AI agents at 35% resolution by building an offline-eval quality gate before shipping.
“You don't want to use your users as test data.”
Lyft's data science lead Nick details how the AI Assist customer-care agent team built an offline evaluation system—inspired by Tau Bench, using a LangGraph agent simulator, LLM-as-user role-play, config-driven YAML scenarios, mocked MCP outputs, and task-specific LLM judges—as a quality gate before production. This eval discipline let them ship seven+ agents and raise resolution rates from 10% to 35%, offering a concrete, replicable playbook for evaluating agents at scale rather than treating users as test data.