The Hallway Track
Engineering Insights

Per-Layer Embeddings (PLE) in Gemma 4 explained

Google Developers (Google I/O) · Sep 25, 2026 · Engineering Insights

Gemma 4's per-layer embeddings boost model performance without increasing compute parameters

“It's a very nice way to increase the performance of the model without actually increasing parameters because these parameters are stored on flash storage.”

Google's Gemma 4 E2B and E4B models use Per-Layer Embeddings (PLE), a technique that gives each token a distinct embedding per transformer layer rather than a single shared representation. The key architectural insight is that these embeddings are stored on flash storage and only a tiny fraction is accessed during inference, meaning the model gains representational power without increasing effective compute parameters. This is a noteworthy efficiency technique relevant to edge/on-device deployment of small language models.

gemma embeddings model-architecture google efficiency on-device-ai

Watch / read the original source →