Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs
Bespoke Labs says AI post-training has shifted from knowledge data to RL environments for agentic tasks
“we have moved on from knowing to doing right so that's the idea of agents”
Mahesh Sathiamoorthy of Bespoke Labs (ex-Google DeepMind) outlines their open-source work on post-training data curation, including the Curator tool, Bespoke Stratos reasoning dataset, Open Thoughts project, and contributions to Terminal Bench. The central thesis is that the AI industry has shifted evaluation focus from what models 'know' (STEM/knowledge benchmarks) to what they can 'do' (agent benchmarks), and that data curators should also act as researchers to understand what actually moves model metrics. The talk is practitioner-level engineering signal relevant to teams building or fine-tuning models for agentic workloads.