56 tracked signals on on-device-ai.
WWDC26 Platforms State of the Union Recap
Apple Developer (WWDC) · Jun 08, 2026
Apple rebuilt Apple Intelligence on Google Gemini technology to power its new Foundation Models.
“Working together with Google and leveraging the technologies behind their Gemini family of models, we created the latest Apple Foundation Models for our integrated Apple Intelligence experiences.”
WWDC26: Platforms State of the Union | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's latest Foundation Models are now built with Google, leveraging Gemini technology to power Apple Intelligence.
“Working together with Google and leveraging the technologies behind their Gemini family of models, we created the latest Apple Foundation models to power our Apple intelligence experiences”
The Billion Dollar AI Advantage Is Disappearing
Two Minute Papers · Oct 05, 2026
Sonnet 5.5 is 5x cheaper than Fable while matching frontier-level quality
“similar intelligence keeps getting possible in smaller and smaller systems”
Introducing Core AI
Apple Developer (WWDC) · Aug 17, 2026
Apple launches Core AI framework for on-device model inference across iOS and macOS.
Inside Apple Intelligence and Xcode: Special Presentation | WWDC26
Apple Developer (WWDC) · Jun 17, 2026
Apple unveils integrated AI platform at WWDC26 with Xcode 27 agentic coding and new on-device frameworks.
“At Apple, AI isn't a layer you bolt on. It's integrated into our platforms.”
Dub Dub Daily: Day 3 | WWDC26
Apple Developer (WWDC) · Jun 10, 2026
Apple opens its Foundation Models framework to Private Cloud Compute and third-party models like Gemini and Claude.
“This year, we're offering developers access to our Private Cloud Compute as well. So you'll have access to a very powerful server model, the one that actually powers our Apple Intelligence features, as well as access to third-party models like Gemini and Claude from Anthropic.”
WWDC26: Integrate on-device AI models into your app using Core AI | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's new Core AI framework lets developers run advanced AI models entirely on-device in their apps.
“With Core AI, you can build app experiences where user's data never leaves their device.”
WWDC26: Bring an LLM provider to the Foundation Models framework | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple opens its Foundation Models framework to nearly any LLM via a new open-source Language Model protocol.
“Anthropic and Google will soon extend the Foundation Models framework with Swift packages of their own, making state-of-the-art Claude and Gemini models available to all Swift developers.”
WWDC26: Meet Core AI | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple opens Core AI, its on-device inference framework powering Apple Intelligence, to all developers.
“Core AI is the inference framework powering on-device Apple Intelligence. And now, it's available for you to use, bringing that same power to your app's own intelligence.”
WWDC26: What’s new in the Foundation Models framework | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple is open-sourcing the Foundation Models framework and adding cloud reasoning models plus third-party model access from Anthropic and Google.
“The Foundation Models framework, including many of the brand new APIs that we're announcing today, is going open source!”
WWDC26: Build with the new Apple Foundation Model on Private Cloud Compute | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple opens its Private Cloud Compute server LLM to third-party developers via the Foundation Models framework at WWDC26.
“And now, by changing just one line of code, you can switch to the new server model on PCC.”
Leverage Private Cloud Compute in your app
Apple Developer (WWDC) · Aug 25, 2026
Apple launches zero-cost server LLM for developers via Private Cloud Compute with no API keys
“there are no token costs to you, the developer”
This Small AI Will Change Everything
Two Minute Papers · Aug 24, 2026
Qwen 3.8B runs on consumer laptops while matching frontier model performance
“if we wait a bit, we might get frontier level systems running on our laptops”
Memory Harnesses for Long-Running Research Agents — Stefania Druga, Sakana.ai
AI Engineer · Aug 12, 2026
Memory harnesses for local models can solve context rot in long-horizon research agents.
“that makes this issue of dealing with context rot a priority”
Gemini 3.5 Live Translate, Gemini in Xcode, and more! - Google Developer News June 2026
Google Developers (Google I/O) · Jun 25, 2026
Google launched Gemini 3.5 Live Translate, a speech-to-speech model interpreting 70+ languages in near real time, now available via API.
“Developers can now build experiences that preserve speaker's tone, pitch, and pacing, resulting in more natural-sounding experiences than voice translation apps of the past.”
Gemma 4 and the AI Edge Gallery: On-Device AI Gets an Upgrade
Google Developers (Google I/O) · Jun 18, 2026
Google's on-device Gemma 4 models hit 5 million AI Edge Gallery downloads within one month of launch.
“Within one month of the launch, we see more than 5 million downloads”
Gemma Playground: AI Edge Gallery
Google Developers (Google I/O) · Jun 18, 2026
Google's Gemma models run entirely on-device on phones, enabling multimodal AI, agent skills, and offline use.
“And what's also important to remember is that this is running entirely on the device. So, it will work offline or in areas of low connectivity.”
Deploy models on-device with Core AI
Apple Developer (WWDC) · Jun 17, 2026
Apple's Core AI lets developers run cutting-edge AI models on-device for private, offline, zero-cost inference.
“Inference happens on device so data stays private, AI features can be readily available and work offline, and there is no per inference cost to you or your users.”
WWDC26: Machine Learning & AI Group Lab | Apple
Apple Developer (WWDC) · Jun 12, 2026
Apple announced new Core AI and evaluations frameworks, foundation model updates, and open-sourced Core ML models at WWDC26.
WWDC26: Coding Intelligence, Machine Learning & AI Group Lab | Apple
Apple Developer (WWDC) · Jun 10, 2026
Apple's Foundation Models framework adds a language model protocol letting developers plug in third-party inference backends.
“We've got a Google package that just went live. Anthropic released theirs this morning.”
Dub Dub Daily: Day 2 | WWDC26
Apple Developer (WWDC) · Jun 09, 2026
Apple's Foundation Models framework now offers frictionless frontier-model access via Private Cloud Compute, plus agentic coding in Xcode.
“the ability to now just access one of these frontier models in Private Cloud Compute with basically no setup or friction at all”
WWDC26: Build agentic app experiences with the Foundation Models framework | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's WWDC26 Foundation Models framework adds DynamicProfile APIs and an open-source Utilities package for building agentic app experiences.
“You can think of this as swapping hats, or switching agents.”
WWDC26: Build AI-powered scripts with the fm CLI and Python SDK | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple is adding an fm command-line tool and Python SDK to access on-device Foundation Models at WWDC26.
“We're introducing a new command line tool called fm, and a new Foundation Models SDK for Python.”
WWDC26: Improve your prompts by hill-climbing with Evaluations | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's new Evaluations framework lets developers hill-climb and align model judges to reduce drift in AI features.
“This discrepancy between model and human is known as drift, and it is a problem faced by all developers trying to evaluate intelligent features.”
WWDC26: Meet the Evaluations framework | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's new Evaluations framework lets developers measure quality of generative-AI features powered by on-device models.
“These models break a contract that is fundamental to software testing.”
WWDC26: Explore distributed inference and training with MLX | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's MLX now supports distributed LLM inference and training across multiple Macs via RDMA over Thunderbolt 5 and the new JACCL library.
“JACCL is an open-source collective communication library built by Apple. It leverages RDMA over Thunderbolt and gives you collective communication primitives for sending data between machines and combining results across the group”
WWDC26: Build intelligent Siri experiences with App Schemas | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's iOS 27 Siri, powered by Apple Intelligence, can access app entities, take actions via App Intents, and understand on-screen context.
“This year, Siri becomes more powerful in three key ways. Siri can now access your app's entities, the real meaningful content inside your app.”
WWDC26: Dive into Core AI model authoring and optimization | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple introduces Core AI, a suite for optimizing, converting, and deploying models including LLMs on Apple Silicon.
“It includes a Swift package for running LLMs in your app. But at its core, it's an open-source repository of models that are ready to go, including generative architectures like cutting-edge large language models.”
Scale agentic AI from on-device to cloud orchestration | BRKSP92
Microsoft Developer (Build) · Jun 04, 2026
On-device AI inference on Intel Panther Lake NPUs eliminates per-token cloud cost and latency for agentic systems.
“every token has a round trip. It has latency. It is a cost item associated with it.”
What's new in the Gemma open model family
Google Developers (Google I/O) · May 22, 2026
Google launches Gemma 4 open model family with four sizes from 2B to 31B parameters
“Our Gemma model, here evaluated on L M Arena, are scoring as well as model 20X the size.”
Per-Layer Embeddings (PLE) in Gemma 4 explained
Google Developers (Google I/O) · Sep 25, 2026
Gemma 4's per-layer embeddings boost model performance without increasing compute parameters
“It's a very nice way to increase the performance of the model without actually increasing parameters because these parameters are stored on flash storage.”
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Simon Willison · Aug 16, 2026
Qwen 3.8 27B defaults to extreme overthinking, consuming 22K tokens for simple tasks
“This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware.”
Compression at the Edge — NVIDIA, Unsloth, HuggingFace, Ollama
AI Engineer · Aug 07, 2026
Panel from NVIDIA, Unsloth, HuggingFace, and Ollama frames quantization as AI democratization at the edge.
“same cost more intelligence”
Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google
AI Engineer · Jul 25, 2026
Tiny LLMs are essential to scale AI intelligence beyond expensive robots to mass-market devices
“If we want intelligence to get into lots and lots and lots of devices and not just really expensive robots, we are going to need tiny models.”
Frontier results, on device - RL Nabors, Arize
AI Engineer · Jun 29, 2026
Local on-device models can replace frontier models like GPT-5 and Claude to cut inference costs, latency, and security risks.
“Every time you reach for foundation models like GPT-5 or Claude, it's costing you, your users, and the environment.”
Ship AI with confidence
Apple Developer (WWDC) · Jun 26, 2026
Apple introduces a new Evaluations Framework at WWDC26 to measure model outputs, including model-judge evaluators.
“if models are judging other models, who's judging the model judges?”
Run Gemma on Reachy Mini, an open source robot
Google Developers (Google I/O) · Jun 26, 2026
Google demos Gemma running on-device on Hugging Face's open-source Reachy Mini robot for local AI interaction.
“Faster processing means quicker reactions, which is great for robots navigating the real world.”
Build On-Device AI Companions with the NVIDIA ACE Game Agent SDK and Unreal Engine 5 Plugins
NVIDIA Developer Blog · Jun 16, 2026
NVIDIA launched the ACE Game Agent SDK and Unreal Engine 5 plugins for building on-device AI companions.
Build local AI agents on Mac with MLX
Apple Developer (WWDC) · Jun 09, 2026
MLX on Apple Silicon runs the entire agentic AI loop locally on a Mac, with M5 neural accelerators making prompt processing up to 4x faster.
“What if your AI agent never needed the internet?”
Gemma Playground: Robot Duck
Google Developers (Google I/O) · Jun 09, 2026
Gemma 4 E2B runs fully on-device for multimodal inference on robot ducks via Raspberry Pi 5 and Jetson Orin Nano.
“And through multimodal inputs of Gemma 4, they're able to process and understand their environment like never before.”
WWDC26: LLM search using Core Spotlight | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple introduces SpotlightSearchTool, letting on-device LLMs search app content via Core Spotlight for grounded responses.
“today, we're introducing SpotlightSearchTool. It's a tool that adopts the tool protocol, to let a language model directly search your app's content in Core Spotlight for contextual response generation.”
WWDC26: What’s new in Shortcuts | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's Shortcuts now integrates more capable Apple Intelligence models with web access, transcript debugging, and persistent storage.
“The Use Model action lets you tap into the power of large language models, directly within your shortcuts.”
WWDC26: Create robust evaluations for agentic apps | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's new Evaluations framework in Xcode 27 lets Swift developers test and scale agentic AI app quality.
“The quality of your evaluation results is only as good as the data behind them.”
WWDC26: Debug and profile agentic app experiences with Instruments | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's Instruments now profiles Foundation Models apps, giving developers observability into on-device and server LLM agent loops.
“Traditional code is predictable. LLMs are non-deterministic - the same input can produce different outputs.”
WWDC26: What’s new in image understanding | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's WWDC26 adds image inputs to Foundation Models and a new Tap to Segment Vision API.
“this year Foundation Models is supporting image inputs”
Updated Apple Developer Program License Agreement and App Review Guidelines now available
Apple · Apple Developer News · Jun 08, 2026
Apple revised its Developer Program License Agreement to formalize AI/ML terms, including Foundation Models framework and access to Apple models.
Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773
TWIML AI Podcast · Aug 12, 2026
Qualcomm VP argues image generation needs architectural innovation, not just scale
Nativ: Run AI models locally on your Mac
Simon Willison · Jul 21, 2026
Nativ is a new macOS desktop app for running AI models locally via MLX
WWDC26: Discover generated subtitles and subtitle styles | Apple
Apple Developer (WWDC) · Jun 08, 2026
Apple's WWDC26 adds on-device AI-generated subtitles via speech transcription and translation for video playback.
“Apple AI generated subtitles can be created live locally on the device as the media plays.”
From prompt to app, build AI powered apps on Windows | DEM345
Microsoft Developer (Build) · Jun 03, 2026
Microsoft demos building offline AI-powered Windows apps using GitHub Copilot, WinUI, and the Windows AI stack.
“whether you're bringing your own model, whether you're looking for models on Hugging Face, or use some of these primitives uh primitives that we provide as part of the Windows SDK, we're trying to make it easy for you to effectively uh run AI models on Windows.”
Build automated agents using optimized AI Foundry models on Snapdragon | DEMSP380
Microsoft Developer (Build) · Jun 03, 2026
LLMware Model HQ runs scheduled, multi-step SLM agent workflows locally on Snapdragon NPUs using AI Foundry-optimized models.
“what if also without having to prompt you could run automated workflows that are scheduled driven”
Build Personal AI Agents on Windows PCs with New Tools from Microsoft and NVIDIA
NVIDIA Developer Blog · Jun 02, 2026
Microsoft and NVIDIA release new tools to build on-device personal AI agents on Windows PCs.
Deprecation of the ImageCreator class
Apple · Apple Developer News · Jun 11, 2026
Apple is deprecating the ImageCreator class, removing programmatic on-device image generation in OS version 27.
“the ImageCreator class is being discontinued and will no longer work in iOS 27, iPadOS 27, macOS 27, and visionOS 27 or later”
Stop routing docstrings to 70B models with on-device AI on Snapdragon | BRKSP90
Microsoft Developer (Build) · Jun 11, 2026
Qualcomm argues small on-device tasks like docstrings shouldn't be routed to large cloud models.
“if you want to predict the future, you have to go and create it”
Stop routing docstrings to 70B models with on-device AI on Snapdragon | BRKSP90
Microsoft Developer (Build) · Jun 04, 2026
On-device AI on Snapdragon can offload lightweight tasks from large cloud models to cut token costs.
“if you want to predict the future, you have to go and create it”
PM vs. Eng
Google Developers (Google I/O) · Jul 23, 2026
Google I/O skit highlights PM-engineer tension behind Android feature development
“I'm a PM and I want a notification summaries so on device AI can automatically summarize 30 plus unread group chat messages into a single haik coup.”