The Hallway Track

on-device-ai

56 tracked signals on on-device-ai.

WWDC26 Platforms State of the Union Recap

Apple Developer (WWDC) · Jun 08, 2026

Apple rebuilt Apple Intelligence on Google Gemini technology to power its new Foundation Models.

“Working together with Google and leveraging the technologies behind their Gemini family of models, we created the latest Apple Foundation Models for our integrated Apple Intelligence experiences.”
WWDC26: Platforms State of the Union | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple's latest Foundation Models are now built with Google, leveraging Gemini technology to power Apple Intelligence.

“Working together with Google and leveraging the technologies behind their Gemini family of models, we created the latest Apple Foundation models to power our Apple intelligence experiences”
The Billion Dollar AI Advantage Is Disappearing

Two Minute Papers · Oct 05, 2026

Sonnet 5.5 is 5x cheaper than Fable while matching frontier-level quality

“similar intelligence keeps getting possible in smaller and smaller systems”
Introducing Core AI

Apple Developer (WWDC) · Aug 17, 2026

Apple launches Core AI framework for on-device model inference across iOS and macOS.

Dub Dub Daily: Day 3 | WWDC26

Apple Developer (WWDC) · Jun 10, 2026

Apple opens its Foundation Models framework to Private Cloud Compute and third-party models like Gemini and Claude.

“This year, we're offering developers access to our Private Cloud Compute as well. So you'll have access to a very powerful server model, the one that actually powers our Apple Intelligence features, as well as access to third-party models like Gemini and Claude from Anthropic.”
WWDC26: Bring an LLM provider to the Foundation Models framework | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple opens its Foundation Models framework to nearly any LLM via a new open-source Language Model protocol.

“Anthropic and Google will soon extend the Foundation Models framework with Swift packages of their own, making state-of-the-art Claude and Gemini models available to all Swift developers.”
WWDC26: Meet Core AI | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple opens Core AI, its on-device inference framework powering Apple Intelligence, to all developers.

“Core AI is the inference framework powering on-device Apple Intelligence. And now, it's available for you to use, bringing that same power to your app's own intelligence.”
WWDC26: What’s new in the Foundation Models framework | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple is open-sourcing the Foundation Models framework and adding cloud reasoning models plus third-party model access from Anthropic and Google.

“The Foundation Models framework, including many of the brand new APIs that we're announcing today, is going open source!”
Leverage Private Cloud Compute in your app

Apple Developer (WWDC) · Aug 25, 2026

Apple launches zero-cost server LLM for developers via Private Cloud Compute with no API keys

“there are no token costs to you, the developer”
This Small AI Will Change Everything

Two Minute Papers · Aug 24, 2026

Qwen 3.8B runs on consumer laptops while matching frontier model performance

“if we wait a bit, we might get frontier level systems running on our laptops”
Gemini 3.5 Live Translate, Gemini in Xcode, and more! - Google Developer News June 2026

Google Developers (Google I/O) · Jun 25, 2026

Google launched Gemini 3.5 Live Translate, a speech-to-speech model interpreting 70+ languages in near real time, now available via API.

“Developers can now build experiences that preserve speaker's tone, pitch, and pacing, resulting in more natural-sounding experiences than voice translation apps of the past.”
Gemma Playground: AI Edge Gallery

Google Developers (Google I/O) · Jun 18, 2026

Google's Gemma models run entirely on-device on phones, enabling multimodal AI, agent skills, and offline use.

“And what's also important to remember is that this is running entirely on the device. So, it will work offline or in areas of low connectivity.”
Deploy models on-device with Core AI

Apple Developer (WWDC) · Jun 17, 2026

Apple's Core AI lets developers run cutting-edge AI models on-device for private, offline, zero-cost inference.

“Inference happens on device so data stays private, AI features can be readily available and work offline, and there is no per inference cost to you or your users.”
Dub Dub Daily: Day 2 | WWDC26

Apple Developer (WWDC) · Jun 09, 2026

Apple's Foundation Models framework now offers frictionless frontier-model access via Private Cloud Compute, plus agentic coding in Xcode.

“the ability to now just access one of these frontier models in Private Cloud Compute with basically no setup or friction at all”
WWDC26: Improve your prompts by hill-climbing with Evaluations | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple's new Evaluations framework lets developers hill-climb and align model judges to reduce drift in AI features.

“This discrepancy between model and human is known as drift, and it is a problem faced by all developers trying to evaluate intelligent features.”
WWDC26: Meet the Evaluations framework | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple's new Evaluations framework lets developers measure quality of generative-AI features powered by on-device models.

“These models break a contract that is fundamental to software testing.”
WWDC26: Explore distributed inference and training with MLX | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple's MLX now supports distributed LLM inference and training across multiple Macs via RDMA over Thunderbolt 5 and the new JACCL library.

“JACCL is an open-source collective communication library built by Apple. It leverages RDMA over Thunderbolt and gives you collective communication primitives for sending data between machines and combining results across the group”
WWDC26: Build intelligent Siri experiences with App Schemas | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple's iOS 27 Siri, powered by Apple Intelligence, can access app entities, take actions via App Intents, and understand on-screen context.

“This year, Siri becomes more powerful in three key ways. Siri can now access your app's entities, the real meaningful content inside your app.”
WWDC26: Dive into Core AI model authoring and optimization | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple introduces Core AI, a suite for optimizing, converting, and deploying models including LLMs on Apple Silicon.

“It includes a Swift package for running LLMs in your app. But at its core, it's an open-source repository of models that are ready to go, including generative architectures like cutting-edge large language models.”
What's new in the Gemma open model family

Google Developers (Google I/O) · May 22, 2026

Google launches Gemma 4 open model family with four sizes from 2B to 31B parameters

“Our Gemma model, here evaluated on L M Arena, are scoring as well as model 20X the size.”
Per-Layer Embeddings (PLE) in Gemma 4 explained

Google Developers (Google I/O) · Sep 25, 2026

Gemma 4's per-layer embeddings boost model performance without increasing compute parameters

“It's a very nice way to increase the performance of the model without actually increasing parameters because these parameters are stored on flash storage.”
Frontier results, on device - RL Nabors, Arize

AI Engineer · Jun 29, 2026

Local on-device models can replace frontier models like GPT-5 and Claude to cut inference costs, latency, and security risks.

“Every time you reach for foundation models like GPT-5 or Claude, it's costing you, your users, and the environment.”
Ship AI with confidence

Apple Developer (WWDC) · Jun 26, 2026

Apple introduces a new Evaluations Framework at WWDC26 to measure model outputs, including model-judge evaluators.

“if models are judging other models, who's judging the model judges?”
Run Gemma on Reachy Mini, an open source robot

Google Developers (Google I/O) · Jun 26, 2026

Google demos Gemma running on-device on Hugging Face's open-source Reachy Mini robot for local AI interaction.

“Faster processing means quicker reactions, which is great for robots navigating the real world.”
Build local AI agents on Mac with MLX

Apple Developer (WWDC) · Jun 09, 2026

MLX on Apple Silicon runs the entire agentic AI loop locally on a Mac, with M5 neural accelerators making prompt processing up to 4x faster.

“What if your AI agent never needed the internet?”
Gemma Playground: Robot Duck

Google Developers (Google I/O) · Jun 09, 2026

Gemma 4 E2B runs fully on-device for multimodal inference on robot ducks via Raspberry Pi 5 and Jetson Orin Nano.

“And through multimodal inputs of Gemma 4, they're able to process and understand their environment like never before.”
WWDC26: LLM search using Core Spotlight | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple introduces SpotlightSearchTool, letting on-device LLMs search app content via Core Spotlight for grounded responses.

“today, we're introducing SpotlightSearchTool. It's a tool that adopts the tool protocol, to let a language model directly search your app's content in Core Spotlight for contextual response generation.”
WWDC26: What’s new in Shortcuts | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple's Shortcuts now integrates more capable Apple Intelligence models with web access, transcript debugging, and persistent storage.

“The Use Model action lets you tap into the power of large language models, directly within your shortcuts.”
WWDC26: Create robust evaluations for agentic apps | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple's new Evaluations framework in Xcode 27 lets Swift developers test and scale agentic AI app quality.

“The quality of your evaluation results is only as good as the data behind them.”
From prompt to app, build AI powered apps on Windows | DEM345

Microsoft Developer (Build) · Jun 03, 2026

Microsoft demos building offline AI-powered Windows apps using GitHub Copilot, WinUI, and the Windows AI stack.

“whether you're bringing your own model, whether you're looking for models on Hugging Face, or use some of these primitives uh primitives that we provide as part of the Windows SDK, we're trying to make it easy for you to effectively uh run AI models on Windows.”
Deprecation of the ImageCreator class

Apple · Apple Developer News · Jun 11, 2026

Apple is deprecating the ImageCreator class, removing programmatic on-device image generation in OS version 27.

“the ImageCreator class is being discontinued and will no longer work in iOS 27, iPadOS 27, macOS 27, and visionOS 27 or later”
PM vs. Eng

Google Developers (Google I/O) · Jul 23, 2026

Google I/O skit highlights PM-engineer tension behind Android feature development

“I'm a PM and I want a notification summaries so on device AI can automatically summarize 30 plus unread group chat messages into a single haik coup.”