The Hallway Track
Research Findings

Hugging Face Journal Club: AsyncOPD and How Stale Can On-Policy Distillation Be?

Hugging Face · Jul 21, 2026 · Research Findings

AsyncOPD enables fully asynchronous on-policy distillation by decoupling student generation from teacher scoring

A Hugging Face journal club session reviews AsyncOPD, a method for asynchronous on-policy distillation that continuously generates student rollouts while teacher scoring and backprop happen in parallel. The core challenge is that fully decoupled async training causes student/teacher log probability divergence, introducing training instability. This is relevant to practitioners optimizing GPU utilization in distillation pipelines, building on prior work like GRPO and libraries like Verl.

on-policy distillation async training knowledge distillation LLM training efficiency reinforcement learning

Watch / read the original source →