How to Build a Model Router in the Harness
LangChain achieved 64% cost reduction in their coding agent via model routing with no quality loss
“we were able to see a 64% reduction in median cost per thread with no measurable change in quality”
LangChain PM Sydney demonstrates a 4-step model routing framework that selects cheap vs. expensive LLMs per task, tested on their Open SWE coding agent. The experiment yielded a 64% median cost reduction with no measurable quality degradation, validated via A/B test on live traffic. The pattern — classify task complexity, map to a cost/intelligence Pareto frontier, route once per thread — is broadly applicable to any agent operating at scale.