Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.
“The North Star of our smaller models is intelligence per byte of memory footprint.”
18 tracked signals on gemma.
Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.
“The North Star of our smaller models is intelligence per byte of memory footprint.”
Google released DiffusionGemma, an open-weight Apache 2 diffusion-based text generation model.
“That research has returned in the best possible way: as a new open weight (Apache 2 licensed) Gemma model”
Google DeepMind released Gemma 4, a new family of open models in four sizes.
“In some situations, you want to own the model. You want to be able to run on your own hardware.”
Google DeepMind announces Gemma 4 12B, a unified encoder-free multimodal open model.
Google launches Anti-gravity agentic platform and Gemma 4 hits 100M downloads in first month
“It's our smartest open model yet. It's purposebuilt for advanced reasoning, agentic workflows, and the response has been incredible. 100 million downloads in the first month and it's pushing Gemma downloads past half a billion.”
Google launched Gemini 3.5 Live Translate, a speech-to-speech model interpreting 70+ languages in near real time, now available via API.
“Developers can now build experiences that preserve speaker's tone, pitch, and pacing, resulting in more natural-sounding experiences than voice translation apps of the past.”
Google's on-device Gemma 4 models hit 5 million AI Edge Gallery downloads within one month of launch.
“Within one month of the launch, we see more than 5 million downloads”
Google's Gemma models run entirely on-device on phones, enabling multimodal AI, agent skills, and offline use.
“And what's also important to remember is that this is running entirely on the device. So, it will work offline or in areas of low connectivity.”
Google DeepMind unveils DiffusionGemma, a diffusion-based model generating text 4x faster.
Google launches Gemma 4 open model family with four sizes from 2B to 31B parameters
“Our Gemma model, here evaluated on L M Arena, are scoring as well as model 20X the size.”
Google releases EmbeddingGemma 2 under Apache 2.0, enabling portable open-weight embeddings
“I don't think it makes sense to use a closed, proprietary, hosted-only model.”
Gemma 4's per-layer embeddings boost model performance without increasing compute parameters
“It's a very nice way to increase the performance of the model without actually increasing parameters because these parameters are stored on flash storage.”
Tiny LLMs are essential to scale AI intelligence beyond expensive robots to mass-market devices
“If we want intelligence to get into lots and lots and lots of devices and not just really expensive robots, we are going to need tiny models.”
Google demos Gemma running on-device on Hugging Face's open-source Reachy Mini robot for local AI interaction.
“Faster processing means quicker reactions, which is great for robots navigating the real world.”
Gemma 4 runs 10 parallel local sub-agents at 170+ tokens/sec on one machine.
“This is the new standard for local parallel AI.”
Gemma autonomously explores an unfamiliar database, plans, and runs experiments to optimize Citi Bike placement for revenue.
“These are the places where you have to increase the number of bicycles.”
Gemma 4 E2B runs fully on-device for multimodal inference on robot ducks via Raspberry Pi 5 and Jetson Orin Nano.
“And through multimodal inputs of Gemma 4, they're able to process and understand their environment like never before.”
Google released AIventure, an open-source game teaching vibe coding and agentic workflows with the Gemma 4 open-weights model.
“you have complete flexibility in how the model is served”