Build realtime multimodal agents with LiveKit and Azure | ODSP937
Real-time multimodal voice agents are bottlenecked by media infrastructure, not models, which LiveKit handles via WebRTC on Azure.
“You see models they are the easy part. Now everything around it is the hard part.”
A Microsoft Build session argues that the hard part of real-time voice agents is the media infrastructure (latency, turn detection, interruption, noise cancellation, scaling) rather than the models, and positions LiveKit—the open-source WebRTC media layer ChatGPT chose for its voice mode—as the solution paired with Azure. It matters as a practical blueprint for production voice agents, but is a technical how-to rather than a major industry announcement.