Hugging Face Journal Club: Direct On-Policy Distillation
On-policy distillation uses RL policy shift signals to train student models without full RL cost
“the student might be stronger than the teacher in this case”
A Hugging Face Journal Club session reviewed a paper on Direct On-Policy Distillation, which proposes reusing signals from RL training runs to guide knowledge distillation rather than running full RL pipelines. The approach inverts the usual distillation assumption — the student model can be larger and stronger than the RL-specialized teacher. This is a cost-reduction technique for post-training that could make RL-style improvements accessible without expensive full RL runs.