The Hallway Track
Research Findings

Hugging Face Journal Club: Direct On-Policy Distillation

Hugging Face · Aug 11, 2026 · Research Findings

On-policy distillation uses RL policy shift signals to train student models without full RL cost

“the student might be stronger than the teacher in this case”

A Hugging Face Journal Club session reviewed a paper on Direct On-Policy Distillation, which proposes reusing signals from RL training runs to guide knowledge distillation rather than running full RL pipelines. The approach inverts the usual distillation assumption — the student model can be larger and stronger than the RL-specialized teacher. This is a cost-reduction technique for post-training that could make RL-style improvements accessible without expensive full RL runs.

distillation reinforcement-learning post-training efficiency LLM-training

Watch / read the original source →