Bring the power of on-device AI to life with Google AI Edge and Gemma
Gemma 4 4B model outperforms last year's 70B model on most benchmarks
“The 4B model actually outperforms the much larger GMA 370B from last year on most of these metrics, not all, which shows how much things have changed.”
Google announced major capability improvements in on-device LLMs at Google I/O 2026, highlighting that the new Gemma 4 4B model outperforms the 70B model from 2024 on most benchmarks. The presentation demonstrated real-time multimodal inference running entirely on a Pixel 10 without an internet connection. This signals that on-device AI has crossed a practical capability threshold, with implications for offline use cases, data privacy, and reduced cloud API costs.