Your Voice Agent Doesn't Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft
Voice agents need sub-950ms response; small models beat frontier models on latency
“A frontier model that think for a full second has already lost the room, no matter how good the answer is.”
Microsoft's Ace team demonstrates that frontier reasoning models are architecturally wrong for real-time voice applications because their thinking latency exceeds the ~950ms human tolerance threshold. Their solution offloads all lesson logic, student tracking, and decision-making into a state machine, feeding the small model only what it needs to speak. This 'thin model, fat orchestration' pattern is a transferable architectural principle for any low-latency voice AI product.