Post-Training and Deploying Open Source Reasoning Models in Foundry | BRK232
Post-training open-source reasoning models with reinforcement learning keeps agent costs flat while quality improves in Foundry.
“deploying a model into production is just the beginning”
A Microsoft Build session demonstrates using post-training (fine-tuning) and reinforcement learning on production agent data to teach open-source reasoning models a company's domain and tools. The pitch is that token-hungry agents drive escalating costs, and learning from production traffic can keep costs flat while holding or improving quality. It matters as practical guidance for scaling agents economically, though it is a technical how-to rather than a major announcement.