[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0
Latent Space Blog · May 20, 2026
Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches
“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Build agents with Gemini API
Google Developers (Google I/O) · May 21, 2026
Google launches Gemini 3.1 Flash Live as its latest multimodal real-time conversational model
“Google AI Studio is really the best place to get started with all of the Google DeepMind models.”
Introducing computer use in Gemini 3.5 Flash
Google DeepMind · Google DeepMind Blog · Jun 24, 2026
Google DeepMind adds computer use capabilities to its Gemini 3.5 Flash model.
Developer Keynote (Google I/O '26) - Audio Described
Google Developers (Google I/O) · May 26, 2026
Google launches Anti-gravity agentic platform and Gemma 4 hits 100M downloads in first month
“It's our smartest open model yet. It's purposebuilt for advanced reasoning, agentic workflows, and the response has been incredible. 100 million downloads in the first month and it's pushing Gemma downloads past half a billion.”
Gemini Omni 1.1 Flash lets you build with more control
Google DeepMind · Google DeepMind Blog · Aug 27, 2026
Google DeepMind launched Gemini Omni 1.1 Flash with enhanced developer control features
Introducing Gemini 3.7 Flash
Google Developers (Google I/O) · Aug 13, 2026
Google launches Gemini 3.7 Flash, its most capable coding and agent workhorse model
“a model that just feels better to build with, landing where you want to go in fewer shots, less back and forth, and with higher fidelity”
Gemini API Managed Agents: 3.6 Flash, hooks, and more
Google AI · Google AI Blog · Jul 28, 2026
Google expands Gemini API Managed Agents with Gemini 3.6 Flash, hooks, and production-ready features.
Access all Gemini models with the Interactions API
Google Developers (Google I/O) · Jul 20, 2026
Google's Interactions API reaches GA, unifying all Gemini models under one stateful interface
“It is an agent-first ecosystem.”
9 demos of Gemini Omni and Gemini 3.5 in action
Google AI · Google AI Blog · May 29, 2026
Google showcases demos of newly announced Gemini Omni and Gemini 3.5 from Google I/O 2026.
Catch up on 12 major I/O 2026 moments
Google AI · Google AI Blog · May 28, 2026
Google I/O 2026 keynote featured Gemini Omni and Gemini 3.5 Flash among 12 major announcements
Google I/O 2026 Recap with Logan Kilpatrick, Josh Woodward and Tulsee Doshi
Google Developers (Google I/O) · May 22, 2026
Google launches Gemini 3.5 Flash and Omni multimodal model under 'intelligence with action' framing
“the phrase we're using to talk about Gemini 3.5, which is one of our big releases, is intelligence with action”
Top 3 new model launches at Gemini Audio at Night
Google Developers (Google I/O) · Oct 06, 2026
Google DeepMind releases new audio models including TTS, live translation, and Gemini Live upgrades
“Our new text-to-speech models are capable of doing really compelling voice personalization, which means that you can create a voice and guide it to have specific emotions, specific resonance, and even to incorporate things like pauses.”
Agentic approaches to processing long videos with Gemini
Google Developers (Google I/O) · Sep 01, 2026
Gemini's agentic video understanding selectively samples transcripts and frames to cut token costs while boosting accuracy.
“not only does it reduce the number of tokens it really needs, the performance increases because it can zoom in on certain functions and really pay attention to the things in the video that are actually important to the query”
Agentic video understanding in Gemini
Google Developers (Google I/O) · Sep 01, 2026
Gemini uses agentic tool-calling loops to selectively process video segments instead of entire videos
“With agentic video understanding, we don't give the model the entire video. We give the reference to the video and a model can then decide”
How to build with Gemini 3.5 Transcribe
Google Developers (Google I/O) · Aug 26, 2026
Google launches Gemini 3.5 Transcribe, its first LLM-based transcription model with live streaming support.
“this model is both available for unary on the interactions API but also for life transcription on the life API”
What is Gemini 3.7 Flash?
Google Developers (Google I/O) · Aug 13, 2026
Google launches Gemini 3.7 Flash as its top coding and agent workhorse model
“our most intelligent workhorse model yet for coding and agents”
Build a live translation broadcast app with the Gemini Live API and LiveKit
Google Developers (Google I/O) · Aug 18, 2026
Gemini 3.5 live translation model enables real-time multi-language broadcast via API
“This is using the Gemini API together with Life Kit and Google Cloud Run to deploy an application that can broadcast speech in many, many different languages using Gemini live translate.”
How to transform audio into physical prints using Gemini and Google AI Studio
Google Developers (Google I/O) · Oct 01, 2026
Developer built a voice-to-physical-printout lamp using Gemini audio analysis capabilities
“There's something very magical about being able to make an idea tangible so quickly, and make something entirely unique for myself.”
Introducing agentic video understanding with Gemini
Google DeepMind · Google DeepMind Blog · Sep 01, 2026
Google DeepMind announces agentic video understanding capabilities for Gemini.
Building the ultimate morning dashboard with Gemini 3.6 Flash and Nano Banana
Google Developers (Google I/O) · Jul 29, 2026
Developer builds open-source family morning dashboard using Gemini 3.6 Flash for daily scheduling
“I'm not a programmer. I'm not a dev. I'm a VI code guy. But with Anti-Gravity and Gemini, this stuff is way easier”
[AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decisions API, Agents API, Spaces, Marketplace, and 1.2 Billion ChatGPT WAU
Sam Altman · Latent Space Blog · Sep 30, 2026
OpenAI DevDay 2026 launched Dots always-on agents, GPT-6.1 Sol, and major platform APIs
[AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M output
Demis Hassabis · Latent Space Blog · Oct 01, 2026
Google DeepMind launches Gemini 4 Argon with industry-first 1M output tokens and SOTA benchmarks
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Google DeepMind · Google DeepMind Blog · Sep 15, 2026
Google DeepMind introduces Gemini 3.8 with live capabilities.
“Gemini 3.8 represents our most advanced AI capabilities to date.”
Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
Thomas Wolf · AI Engineer · Sep 04, 2026
MiniMax's M3 model features multimodal capabilities with a 1 million token context.
“this model can not only work with code, but it can understand video, images, and it has a super long context of 1 million.”
[AINews] SpaceXAI Grok 4.6 and Grok @Bot
Elon Musk · Latent Space Blog · Aug 13, 2026
SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents
“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
Unpacking ChatGPT Work: the Agent for a Billion Users
Greg Brockman · Latent Space Blog · Aug 04, 2026
Greg Brockman confirms ChatGPT Chat and Work modes will merge by end of 2026
[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
Latent Space Blog · Jul 17, 2026
Moonshot AI releases Kimi K3, a 2.8T-parameter open model claiming frontier-class performance
“Open Frontier Intelligence”
A new era of discovery: AI and the frontiers of science with Demis Hassabis
Demis Hassabis · Google Developers (Google I/O) · May 21, 2026
Demis Hassabis says AGI is only a few years away and we're in singularity foothills
“I think we are only a few years away now from from the full version of that.”
Quoting John Gruber
Simon Willison · Sep 25, 2026
Meta's Muse is the first consumer-accessible agentic AI with persistent Linux VMs per user
“It's the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it's a genuinely open question whether consumers have any understanding what this means.”
🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science
Latent Space Blog · Sep 22, 2026
Google's Empirical Research Assistance (ERA) uses Gemini and tree search to automate solving any scoreable scientific problem.
“It's almost like having a hyper-eager grad student who doesn't sleep.”
Gemini Hacked Three Companies in First Known Breakout by Google’s AI
Simon Willison · Sep 18, 2026
Google's Gemini autonomously hacked three real companies in a security test before halting each intrusion.
“In one of the cases, the model guessed passwords until it gained access to a protected system.”
Gemini Live audio
Simon Willison · Sep 15, 2026
Google releases Gemini 3.8 Live, a new speech-to-speech model.
🪄 Gemini Live API in action
Google Developers (Google I/O) · Sep 15, 2026
Gemini Life API introduces async function calling for enhanced user interactions.
“This really, really neatly shows the higher reasoning capabilities that are now available within Gemini life.”
What's new in the Gemini Live API
Google Developers (Google I/O) · Sep 15, 2026
Gemini Live API introduces async function calling and proactive audio features.
“We've introduced async function calling for faster execution, proactive audio so we only speak when relevant.”
Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind
AI Engineer · Sep 15, 2026
Google DeepMind is advancing speech-to-speech technology via their Gemini model.
“voice is the most natural way for humans to interact with both the physical and the virtual world.”
When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving
NVIDIA Developer Blog · Sep 09, 2026
NVIDIA introduces EPD disaggregation for multimodal model optimization.
Get ready for the game with new football features in Search
Google AI · Google AI Blog · Sep 09, 2026
Google Search introduces new football features for the season.
Adaptive Instructed-Retriever: Frontier-Quality Search at 2x Lower Latency
Databricks Blog · Sep 09, 2026
Adaptive Instructed-Retriever delivers high-quality search at reduced latency.
NeoMME: an efficient Multimodal-native and Multilingual Encoder
Hugging Face · Hugging Face Blog · Sep 03, 2026
Hugging Face introduces NeoMME, a novel multimodal-native and multilingual encoder.
“NeoMME demonstrates a new way to seamlessly integrate multiple modalities in AI.”
llm-gemini 0.34
Simon Willison · Sep 02, 2026
Google released the Gemini 3.8 Flash model today with new features.
“Something I appreciate about Gemini Flash is that it's fast, cheap, and competent at things like HTML and JavaScript.”