Inside Clay's Eval Stack: 300M Agent Runs, One LangSmith Pipeline
Clay runs 300M agent executions monthly and calls evals non-negotiable at that scale
“eval's became non-negotiable”
Clay describes running 300M monthly agent runs through their Claygent research agent and 100K+ weekly messages to their Sculptor workflow agent, forcing them to build rigorous eval pipelines via LangSmith rather than manually reviewing traces. The talk frames robust evals as the prerequisite for safe agentic development at scale, with teams using Claude, Codex, and Devin to iterate on prompts once evals are in place. This is a useful practitioner data point on what production-scale agentic infrastructure looks like, though it is more a vendor case study than a major industry signal.