The Hallway Track

CVPR 2026

vision robotics multimodal research

Dates
2026-06-20 → 2026-06-25
Location
Denver, CO
Ecosystem
frontier research
Importance
8/10

Official site →

Related coverage & signals

DeepMind Just Changed How AI Sees The World

Two Minute Papers · Aug 07, 2026

DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.

“Throw that all away. Out. Right now.”
Gemma Playground: Robot Duck

Google Developers (Google I/O) · Jun 09, 2026

Gemma 4 E2B runs fully on-device for multimodal inference on robot ducks via Raspberry Pi 5 and Jetson Orin Nano.

“And through multimodal inputs of Gemma 4, they're able to process and understand their environment like never before.”
[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0

Latent Space Blog · May 20, 2026

Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches

“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Funding grants for new research into AI and teen development

OpenAI · OpenAI Blog · Sep 08, 2026

OpenAI announces a $5 million grant for research on AI's impact on teens.

“Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.”
Introducing Gemini Robotics 2

Google Developers (Google I/O) · Jul 31, 2026

Google DeepMind launches Gemini Robotics 2 with whole-body intelligence, dexterity, and multi-robot collaboration.

“What we're building here is the intelligence layer to power any robot to do a broad range of useful tasks.”
Gemma 4 in Action: Bringing Frontier AI to the Edge

Google Developers (Google I/O) · Jun 29, 2026

Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.

“The North Star of our smaller models is intelligence per byte of memory footprint.”
Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era

Jack Clark · Import AI · Jun 29, 2026

NVIDIA's ENPIRE framework lets coding agents run a closed-loop self-improvement process for real-world robots, hitting 99% on dexterous tasks.

“Frontier coding agents can autonomously develop a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks in the real world, such as PushT, organizing pins into a pin box, and using a cutter to cut a zip tie,”
Introducing EmbeddingGemma 2: An open model for natively multimodal embeddings

Google Developers (Google I/O) · Oct 06, 2026

Google releases EmbeddingGemma 2, a sub-billion open model unifying text, image, video, and audio embeddings on-device.

“a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in a shared high-dimensional embedding space”
From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI

AI Engineer · Sep 23, 2026

Perceptron AI wants to unify VLMs, VLAs, and world models into 'embodied foundational models' for real-time physical-world interaction.

“we want to move away from the distinction between VLM, VLA, world models, whatever you want to call it, to what we call embodied fundamental models”