Introducing: LangSmith Tuned Evaluators
LangChain launches tuned evaluators that beat frontier models on agent evals at lower cost
“our post-trained perceived error evaluator outperformed all frontier closed and open models, while still remaining the most cost-effective”
LangChain launched LangSmith Tuned Evaluators, specialized models managed end-to-end by LangChain that claim to outperform frontier models on specific evaluation tasks at a fraction of the cost. The first evaluator, 'Perceived Error,' detects agent failures in multi-turn conversations by spotting subtle signals like contradictory answers and unresolved outcomes — no custom prompts or inference infrastructure required. This matters because most production AI failures don't surface as system errors and users rarely leave explicit ratings, making scalable automated quality signal a critical gap for agent teams.