How Salesforce Standardizes Agent Evals with LangSmith
Salesforce AgentForce uses LangSmith to standardize evals across teams at scale
“LangSmith helps us to amplify our internal expertise, allowing us to run thousands and thousands of test cases.”
Salesforce's AgentForce vibe coding product adopted LangSmith as a unified evaluation platform after teams were independently building MCP tools with no shared quality bar. LangSmith now standardizes scoring across metrics like instruction following, coherence, factuality, and deployability at scale before code reaches customers. The product has surpassed 100 million lines of accepted code, making consistent eval infrastructure a critical quality gate.