The Hallway Track
Engineering Insights

Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth

AI Engineer · Jul 17, 2026 · Engineering Insights

Unsloth fixed a gradient accumulation bug improving training accuracy by 1-3% across the entire stack.

“fixed a gradient accumulation bug fix um which increased accuracy by 1 to 3% um across the entire training stack”

Daniel Han of Unsloth (one of Hugging Face's top model distributors with 300M+ downloads) outlines the org's broad contributions to the open-source AI stack, including async gradient checkpointing, flex attention, and a gradient accumulation bug fix that boosted training accuracy 1-3%. The talk is a workshop intro covering kernels, RL, and reward hacking in agents. Content is cut off before the substantive technical sections begin, limiting assessable signal.

unsloth open-source training quantization gradient-checkpointing hugging-face

Watch / read the original source →