Smaller, faster, smarter: Distilling models with fine‑tuning | DEM322
Microsoft Foundry distills production agent traces into smaller fine-tuned models that run cheaper and faster.
“the question is not just whether my agent works. It is about whether I can afford to run my agent a 100 million times”
At Microsoft Build, the Foundry fine-tuning team detailed how enterprises can distill large 'teacher' models into smaller 'student' models by fine-tuning on production agent traces, capturing real tool-call sequences and arguments. This matters because agentic AI is token-intensive, and distillation makes intelligence cheap enough to run at enterprise scale as a utility rather than a luxury.