Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute
Post-training must evolve to let agents adapt to enterprise harnesses without source code access
“there will be these kind of agentic citizens, which you can just deploy once, and they'll be able to adapt to many different types of out of distribution tasks and learn from their interactions”
Raymond Feng of Applied Compute outlines a tiered post-training roadmap: from single-turn Q&A, to training on custom enterprise harnesses without source code access, to a future of 'agentic citizens' that continuously learn from live interactions. The enterprise demand driver is plug-and-play deployability — companies want to fine-tune models to fit existing workflows without handing over infrastructure internals. This frames the next frontier of model customization as harness-agnostic, self-improving agents rather than static fine-tunes.