From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind
Google DeepMind launched Gemini 3.1 Flash Live, a full-duplex real-time sound-to-sound conversational model, plus Gemma 4 with on-device audio understanding.
“we recently launched Gemini 3.1 flash life, which is our kind of full duplex, you know, sound to sound real-time conversational model”
Google DeepMind's Thor Schaeff outlined Gemini's audio stack, highlighting Gemma 4 with on-device audio understanding, Gemini 3's nuanced audio comprehension beyond transcription, and the new Gemini 3.1 Flash Live full-duplex real-time conversational model. These releases signal Google pushing multimodal, real-time, and edge-deployable audio AI as core developer capabilities.