Hugging Face Journal Club: AsyncOPD and How Stale Can On-Policy Distillation Be?
AsyncOPD enables fully asynchronous on-policy distillation by decoupling student generation from teacher scoring
A Hugging Face journal club session reviews AsyncOPD, a method for asynchronous on-policy distillation that continuously generates student rollouts while teacher scoring and backprop happen in parallel. The core challenge is that fully decoupled async training causes student/teacher log probability divergence, introducing training instability. This is relevant to practitioners optimizing GPU utilization in distillation pipelines, building on prior work like GRPO and libraries like Verl.