RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor
Mercor scaled to $2B revenue run rate in 4 months as primary agentic data vendor to frontier labs
“this giant transition away from the low-skilled crowdsourcing era of behavior cloning data and moving towards the agentic era of data”
Mercor's Brendan Foody describes a fundamental shift in AI training data from low-skill crowdsourcing for behavior cloning to high-skill expert networks building RL environments and frontier evals. The company grew from $1B to $2B ARR in ~4 months serving labs and app-layer companies like Harvey, Cognition, and Ramp. The talk signals that RL environments — rich simulated worlds teaching agents to use real software tools — are the new frontier of post-training infrastructure.