Frontier post-training recipe review with Finbarr Timbers
2026 frontier post-training has shifted to Multi-Teacher On-Policy Distillation (MOPD), merging many specialist models into one.
“The shape of a post-training recipe has changed more in the last year than in the prior three.”— Nathan Lambert
Nathan Lambert and Finbarr Timbers review how frontier post-training recipes evolved from InstructGPT's single SFT-RM-RL pipeline to 2026's fragmented specialist-and-merge approach centered on Multi-Teacher On-Policy Distillation (MOPD). MOPD trains many domain-specialist teachers, then distills a single general student via reverse-KL on its own rollouts, a pattern now scaling past 10 teachers in models like DeepSeek V4 and Nemotron 3 Ultra. It matters because it signals how the open frontier is solving the cost and reward-conflict problems of large-scale RL.