23 tracked signals on post-training.
[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law
Latent Space Blog · Aug 20, 2026
GLM-5.3 proves post-training RL on long-horizon tasks beats parameter scaling for reasoning
“Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions.”
DeepSeek Just Made Closed AI Look Ridiculous
Two Minute Papers · Aug 19, 2026
DeepSeek 4 Pro achieves near-frontier quality with MIT-licensed open weights, pressuring closed AI labs.
“DeepSeek has MIT licensed open weights. Anyone can run the exact same model at their own price.”
Another DeepSeek Moment Has Arrived
Two Minute Papers · Aug 03, 2026
DeepSeek's updated flash model beats its larger pro version via post-training improvements alone
“Just the post-training step changed?”
Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs
AI Engineer · Jul 24, 2026
Opus 4.8 and Fable regressed on business-task evals after Anthropic removed business skills from post-training
“they removed a part of the post-training recipe that was meant to do business skills”
Frontier post-training recipe review with Finbarr Timbers
Nathan Lambert · Interconnects · Jun 16, 2026
2026 frontier post-training has shifted to Multi-Teacher On-Policy Distillation (MOPD), merging many specialist models into one.
“The shape of a post-training recipe has changed more in the last year than in the prior three.”
How to go from your agent's traces to a fine-tuned model in one workflow
LangChain · Sep 24, 2026
LangChain launches LangSmith fine-tuning in public beta with SmithTune, a CLI to post-train models from agent traces.
“today we're launching LangSmith fine-tuning in public beta with SmithTune, a CLI to allow you to post-train models from your LangSmith traces in one workflow”
I wrote an AI textbook — how long until AI can do it better?
Sam Altman · Interconnects · Aug 12, 2026
LLMs are stagnant at long-form non-fiction writing, a prerequisite for autonomous science
“Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future.”
RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor
Sequoia Capital · Aug 12, 2026
Mercor scaled to $2B revenue run rate in 4 months as primary agentic data vendor to frontier labs
“this giant transition away from the low-skilled crowdsourcing era of behavior cloning data and moving towards the agentic era of data”
How Harvey Built a Research Lab on a Budget | Gabe Pereyra
Sequoia Capital · Aug 11, 2026
Harvey built a domain-specific research lab by leveraging frontier ecosystem rather than competing with it directly.
“it's an unfair game competing with the frontier labs if you're an application layer company.”
Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal
AI Engineer · Aug 10, 2026
Modal proposes decoupling RL rollout workers from trainer clusters to use distributed GPU capacity across datacenters.
“IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time.”
[AINews] not much happened today
Latent Space Blog · Aug 01, 2026
DeepSeek V4-Flash 0731 matches GPT-5.6 performance at 60% lower cost via post-training alone
“Terminal-Bench 82.7, up +25.8 from the April preview's 56.9”
What's Next After RLHF? — Diogo Almeida, TypeSafe AI
AI Engineer · Jul 31, 2026
OpenAI post-training pioneer argues Claude Code and ChatGPT are the same era, not successive ones
“the team I was part of basically invented post-training as a concept”
Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute
AI Engineer · Jul 31, 2026
Post-training must evolve to let agents adapt to enterprise harnesses without source code access
“there will be these kind of agentic citizens, which you can just deploy once, and they'll be able to adapt to many different types of out of distribution tasks and learn from their interactions”
Post-Training Is How You Keep Your Taste | Fireworks CEO Lin Qiao
Sequoia Capital · Aug 12, 2026
Post-training is the primary moat for AI application companies over black-box APIs
“right now, with one person, a few weeks, uh without understanding how to write a single line of code, you can do that”
Hugging Face Journal Club: Direct On-Policy Distillation
Hugging Face · Aug 11, 2026
On-policy distillation uses RL policy shift signals to train student models without full RL cost
“the student might be stronger than the teacher in this case”
5 useful things you'll learn in my new post-training textbook (shipping now!)
Nathan Lambert · Interconnects · Aug 10, 2026
Nathan Lambert publishes definitive RLHF post-training textbook, free online with 12-hour course
“This is the book I wanted to read when I was getting started a few years ago!”
Hugging Face Journal Club: Scaling Laws for Pre-training & RL
Hugging Face · Aug 04, 2026
New research uses chess as a testbed to derive scaling laws for post-training compute allocation
“In pre-training we had the chinchilla scaling laws uh which is roughly like uh uh you would allocate 20x the tokens per parameters of u of your model. But in a we didn't know like how much you would allocate.”
Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs
AI Engineer · Jul 31, 2026
Bespoke Labs says AI post-training has shifted from knowledge data to RL environments for agentic tasks
“we have moved on from knowing to doing right so that's the idea of agents”
Emulated: The Data for Fully Autonomous Software Engineers and Companies — Joseph Wang
AI Engineer · Jul 31, 2026
AI agent capability gaps in infrastructure reasoning stem from missing training data, not model limitations.
“like with everything in NML, uh the gap in models is usually a gap in data”
The Base Model Is Dead — Varun Singh, Arcee AI
AI Engineer · Jul 31, 2026
The traditional base-model paradigm of web-scale pre-training is being displaced by post-training
“RL was mostly just a cherry on top, shaping the flavor of the interactions more than conferring extra knowledge or quality onto the base model itself.”
Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect
AI Engineer · Jul 31, 2026
Prime Intellect is extending RL to real-world tasks lacking verifiable reward signals.
State of the blog, mid-2026
Nathan Lambert · Interconnects · Jun 17, 2026
Nathan Lambert keeps Interconnects intentionally raw and independent, disclosing new advising roles at Arcee AI and Mercor.
“Interconnects is the tip of the spear of all of my missions in AI.”
How to Stop Shipping Low-Quality RL Environments (with Examples)
Latent Space Blog · Jun 05, 2026
Low-quality RL training environments and broken harnesses actively degrade models and ruin training runs.
“researchers don’t want your broken RL environments because they will make our models worse”