The Hallway Track
Engineering Insights

How Lyft Builds Evals That Actually Matter in Production | Interrupt 26

LangChain · Jun 15, 2026 · Engineering Insights

Lyft scaled to seven+ production AI agents at 35% resolution by building an offline-eval quality gate before shipping.

“You don't want to use your users as test data.”

Lyft's data science lead Nick details how the AI Assist customer-care agent team built an offline evaluation system—inspired by Tau Bench, using a LangGraph agent simulator, LLM-as-user role-play, config-driven YAML scenarios, mocked MCP outputs, and task-specific LLM judges—as a quality gate before production. This eval discipline let them ship seven+ agents and raise resolution rates from 10% to 35%, offering a concrete, replicable playbook for evaluating agents at scale rather than treating users as test data.

evals ai-agents production llm-as-judge lyft langchain

Watch / read the original source →