The Hallway Track

multimodal

58 tracked signals on multimodal.

[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0

Latent Space Blog · May 20, 2026

Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches

“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Gemma 4 in Action: Bringing Frontier AI to the Edge

Google Developers (Google I/O) · Jun 29, 2026

Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.

“The North Star of our smaller models is intelligence per byte of memory footprint.”
Introducing EmbeddingGemma 2: An open model for natively multimodal embeddings

Google Developers (Google I/O) · Oct 06, 2026

Google releases EmbeddingGemma 2, a sub-billion open model unifying text, image, video, and audio embeddings on-device.

“a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in a shared high-dimensional embedding space”
Image understanding with on-device AI

Apple Developer (WWDC) · Aug 11, 2026

Apple's on-device Foundation Models framework gains multimodal image understanding and Vision Framework integration.

“This opens up new categories of experiences you can build with image understanding. It's as simple as attaching an image to your prompt.”
DeepMind Just Changed How AI Sees The World

Two Minute Papers · Aug 07, 2026

DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.

“Throw that all away. Out. Right now.”
HTML Is All Agents Need — James Russo, HeyGen

Andrej Karpathy · AI Engineer · Jul 21, 2026

HeyGen uses HTML as the native canvas for AI agents to generate full video compositions

“Uh when you try to teach a model a new DSL or even your own custom JSON structure, it's forcing it to speak another language.”
Gemma Playground: AI Edge Gallery

Google Developers (Google I/O) · Jun 18, 2026

Google's Gemma models run entirely on-device on phones, enabling multimodal AI, agent skills, and offline use.

“And what's also important to remember is that this is running entirely on the device. So, it will work offline or in areas of low connectivity.”
Top 3 new model launches at Gemini Audio at Night

Google Developers (Google I/O) · Oct 06, 2026

Google DeepMind releases new audio models including TTS, live translation, and Gemini Live upgrades

“Our new text-to-speech models are capable of doing really compelling voice personalization, which means that you can create a voice and guide it to have specific emotions, specific resonance, and even to incorporate things like pauses.”
Agentic approaches to processing long videos with Gemini

Google Developers (Google I/O) · Sep 01, 2026

Gemini's agentic video understanding selectively samples transcripts and frames to cut token costs while boosting accuracy.

“not only does it reduce the number of tokens it really needs, the performance increases because it can zoom in on certain functions and really pay attention to the things in the video that are actually important to the query”
Agentic video understanding in Gemini

Google Developers (Google I/O) · Sep 01, 2026

Gemini uses agentic tool-calling loops to selectively process video segments instead of entire videos

“With agentic video understanding, we don't give the model the entire video. We give the reference to the video and a model can then decide”
Qwen3.8-Flash-Next

Simon Willison · Aug 26, 2026

Qwen releases 125B MoE model with only 6B active parameters as Qwen4 architecture preview

“a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4”
How to build with Gemini 3.5 Transcribe

Google Developers (Google I/O) · Aug 26, 2026

Google launches Gemini 3.5 Transcribe, its first LLM-based transcription model with live streaming support.

“this model is both available for unary on the interactions API but also for life transcription on the life API”
Sketch your idea, let Xcode build it

Apple Developer (WWDC) · Aug 14, 2026

Apple Xcode agents can now interpret hand-drawn sketches and generate matching UI code

“Xcode took my image and used it as inspiration to create this chart.”
PipeNetwork/minimax-h3-mlx

Simon Willison · Aug 04, 2026

MiniMax-H3 omni-modal model now runs locally on Apple Silicon via MLX port

“a rainbow colored skunk leaps over a mossy log in a supermarket”
Pair Nova 2 Lite with Claude for cost-optimized document processing

AWS Machine Learning Blog · Jun 29, 2026

Pairing Amazon Nova 2 Lite with Claude Sonnet 4.6 in a two-model Bedrock pipeline cuts per-page document-digitization cost by about two-thirds.

“This two-model approach costs about two-thirds less per page than a single-model alternative that sends the entire task to one vision-language model.”
Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI

NVIDIA Developer Blog · Jun 16, 2026

NVIDIA launches XR AI, a reusable foundation for building AI agents on AR glasses and XR devices.

“The hardware is ready, but creating AI experiences requires integrating live camera and microphone streams, multimodal AI models, enterprise data, tool use, deployment infrastructure, and device-specific runtimes.”
Gemma Playground: Robot Duck

Google Developers (Google I/O) · Jun 09, 2026

Gemma 4 E2B runs fully on-device for multimodal inference on robot ducks via Raspberry Pi 5 and Jetson Orin Nano.

“And through multimodal inputs of Gemma 4, they're able to process and understand their environment like never before.”
Build realtime multimodal agents with LiveKit and Azure | ODSP937

Microsoft Developer (Build) · Jun 03, 2026

Real-time multimodal voice agents are bottlenecked by media infrastructure, not models, which LiveKit handles via WebRTC on Azure.

“You see models they are the easy part. Now everything around it is the hard part.”
Run Step 3.7 Flash on NVIDIA GPUs with Enterprise-Ready Multimodal AI

NVIDIA Developer Blog · May 29, 2026

StepFun's Step 3.7 Flash 198B multimodal model is now available on NVIDIA-accelerated infrastructure

“AI applications are moving beyond text generation to multimodal systems that can perceive, search, and reason across images, documents, video, and language in real time”
Build agents with Gemini API

Google Developers (Google I/O) · May 21, 2026

Google launches Gemini 3.1 Flash Live as its latest multimodal real-time conversational model

“Google AI Studio is really the best place to get started with all of the Google DeepMind models.”
Build a live translation broadcast app with the Gemini Live API and LiveKit

Google Developers (Google I/O) · Aug 18, 2026

Gemini 3.5 live translation model enables real-time multi-language broadcast via API

“This is using the Gemini API together with Life Kit and Google Cloud Run to deploy an application that can broadcast speech in many, many different languages using Gemini live translate.”
Object detection with Amazon Nova 2 Lite

AWS Machine Learning Blog · Jun 02, 2026

Amazon Nova 2 Lite enables zero-shot object detection via natural language prompts with no model training required.

“This multimodal foundation model detects objects through natural language prompts with no training required.”
Kākāpō Party

Simon Willison · Sep 26, 2026

Claude Opus 5.5 generates pixel art animations from photos via simple prompts

“I had seen some buzz around how good Claude Opus 5.5 was at creating pixel art animations.”