DeepMind Just Changed How AI Sees The World
DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.
“Throw that all away. Out. Right now.”
DeepMind revealed the architecture behind Gemma 4's multimodal capabilities: rather than chaining separate vision and audio encoders, it slices images into patches and audio into 40ms chunks and feeds them directly as tokens into a single transformer. This unified approach enables a 12B parameter model to handle vision, audio, and language without the overhead of dedicated encoder networks. With over 300 million downloads, Gemma 4 represents a meaningful efficiency signal for the industry's push toward capable local models.