The Hallway Track

Google I/O 2026

gemini search agents multimodal

Dates
2026-05-19 → 2026-05-20
Location
Mountain View, CA
Ecosystem
hyperscaler
Importance
10/10

Related coverage & signals

[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0

Latent Space Blog · May 20, 2026

Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches

“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Build agents with Gemini API

Google Developers (Google I/O) · May 21, 2026

Google launches Gemini 3.1 Flash Live as its latest multimodal real-time conversational model

“Google AI Studio is really the best place to get started with all of the Google DeepMind models.”
Developer Keynote (Google I/O '26) - Audio Described

Google Developers (Google I/O) · May 26, 2026

Google launches Anti-gravity agentic platform and Gemma 4 hits 100M downloads in first month

“It's our smartest open model yet. It's purposebuilt for advanced reasoning, agentic workflows, and the response has been incredible. 100 million downloads in the first month and it's pushing Gemma downloads past half a billion.”
Introducing Gemini 3.7 Flash

Google Developers (Google I/O) · Aug 13, 2026

Google launches Gemini 3.7 Flash, its most capable coding and agent workhorse model

“a model that just feels better to build with, landing where you want to go in fewer shots, less back and forth, and with higher fidelity”
Top 3 new model launches at Gemini Audio at Night

Google Developers (Google I/O) · Oct 06, 2026

Google DeepMind releases new audio models including TTS, live translation, and Gemini Live upgrades

“Our new text-to-speech models are capable of doing really compelling voice personalization, which means that you can create a voice and guide it to have specific emotions, specific resonance, and even to incorporate things like pauses.”
Agentic approaches to processing long videos with Gemini

Google Developers (Google I/O) · Sep 01, 2026

Gemini's agentic video understanding selectively samples transcripts and frames to cut token costs while boosting accuracy.

“not only does it reduce the number of tokens it really needs, the performance increases because it can zoom in on certain functions and really pay attention to the things in the video that are actually important to the query”
Agentic video understanding in Gemini

Google Developers (Google I/O) · Sep 01, 2026

Gemini uses agentic tool-calling loops to selectively process video segments instead of entire videos

“With agentic video understanding, we don't give the model the entire video. We give the reference to the video and a model can then decide”
How to build with Gemini 3.5 Transcribe

Google Developers (Google I/O) · Aug 26, 2026

Google launches Gemini 3.5 Transcribe, its first LLM-based transcription model with live streaming support.

“this model is both available for unary on the interactions API but also for life transcription on the life API”
What is Gemini 3.7 Flash?

Google Developers (Google I/O) · Aug 13, 2026

Google launches Gemini 3.7 Flash as its top coding and agent workhorse model

“our most intelligent workhorse model yet for coding and agents”
Build a live translation broadcast app with the Gemini Live API and LiveKit

Google Developers (Google I/O) · Aug 18, 2026

Gemini 3.5 live translation model enables real-time multi-language broadcast via API

“This is using the Gemini API together with Life Kit and Google Cloud Run to deploy an application that can broadcast speech in many, many different languages using Gemini live translate.”
[AINews] SpaceXAI Grok 4.6 and Grok @Bot

Elon Musk · Latent Space Blog · Aug 13, 2026

SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents

“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
Quoting John Gruber

Simon Willison · Sep 25, 2026

Meta's Muse is the first consumer-accessible agentic AI with persistent Linux VMs per user

“It's the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it's a genuinely open question whether consumers have any understanding what this means.”
Gemini Live audio

Simon Willison · Sep 15, 2026

Google releases Gemini 3.8 Live, a new speech-to-speech model.

🪄 Gemini Live API in action

Google Developers (Google I/O) · Sep 15, 2026

Gemini Life API introduces async function calling for enhanced user interactions.

“This really, really neatly shows the higher reasoning capabilities that are now available within Gemini life.”
What's new in the Gemini Live API

Google Developers (Google I/O) · Sep 15, 2026

Gemini Live API introduces async function calling and proactive audio features.

“We've introduced async function calling for faster execution, proactive audio so we only speak when relevant.”
llm-gemini 0.34

Simon Willison · Sep 02, 2026

Google released the Gemini 3.8 Flash model today with new features.

“Something I appreciate about Gemini Flash is that it's fast, cheap, and competent at things like HTML and JavaScript.”