The Hallway Track

post-training

23 tracked signals on post-training.

DeepSeek Just Made Closed AI Look Ridiculous

Two Minute Papers · Aug 19, 2026

DeepSeek 4 Pro achieves near-frontier quality with MIT-licensed open weights, pressuring closed AI labs.

“DeepSeek has MIT licensed open weights. Anyone can run the exact same model at their own price.”
Another DeepSeek Moment Has Arrived

Two Minute Papers · Aug 03, 2026

DeepSeek's updated flash model beats its larger pro version via post-training improvements alone

“Just the post-training step changed?”
Frontier post-training recipe review with Finbarr Timbers

Nathan Lambert · Interconnects · Jun 16, 2026

2026 frontier post-training has shifted to Multi-Teacher On-Policy Distillation (MOPD), merging many specialist models into one.

“The shape of a post-training recipe has changed more in the last year than in the prior three.”
How to go from your agent's traces to a fine-tuned model in one workflow

LangChain · Sep 24, 2026

LangChain launches LangSmith fine-tuning in public beta with SmithTune, a CLI to post-train models from agent traces.

“today we're launching LangSmith fine-tuning in public beta with SmithTune, a CLI to allow you to post-train models from your LangSmith traces in one workflow”
I wrote an AI textbook — how long until AI can do it better?

Sam Altman · Interconnects · Aug 12, 2026

LLMs are stagnant at long-form non-fiction writing, a prerequisite for autonomous science

“Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future.”
How Harvey Built a Research Lab on a Budget | Gabe Pereyra

Sequoia Capital · Aug 11, 2026

Harvey built a domain-specific research lab by leveraging frontier ecosystem rather than competing with it directly.

“it's an unfair game competing with the frontier labs if you're an application layer company.”
Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal

AI Engineer · Aug 10, 2026

Modal proposes decoupling RL rollout workers from trainer clusters to use distributed GPU capacity across datacenters.

“IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time.”
[AINews] not much happened today

Latent Space Blog · Aug 01, 2026

DeepSeek V4-Flash 0731 matches GPT-5.6 performance at 60% lower cost via post-training alone

“Terminal-Bench 82.7, up +25.8 from the April preview's 56.9”
Hugging Face Journal Club: Scaling Laws for Pre-training & RL

Hugging Face · Aug 04, 2026

New research uses chess as a testbed to derive scaling laws for post-training compute allocation

“In pre-training we had the chinchilla scaling laws uh which is roughly like uh uh you would allocate 20x the tokens per parameters of u of your model. But in a we didn't know like how much you would allocate.”
The Base Model Is Dead — Varun Singh, Arcee AI

AI Engineer · Jul 31, 2026

The traditional base-model paradigm of web-scale pre-training is being displaced by post-training

“RL was mostly just a cherry on top, shaping the flavor of the interactions more than conferring extra knowledge or quality onto the base model itself.”
State of the blog, mid-2026

Nathan Lambert · Interconnects · Jun 17, 2026

Nathan Lambert keeps Interconnects intentionally raw and independent, disclosing new advising roles at Arcee AI and Mercor.

“Interconnects is the tip of the spear of all of my missions in AI.”