Black Forest Labs launches FLUX 3 Video, beating Seedance 2.0, Gemini Omni, and Grok Imagine
CVPR 2026
- Dates
- 2026-06-20 → 2026-06-25
- Location
- Denver, CO
- Ecosystem
- frontier research
- Importance
- 8/10
Related coverage & signals
Today's top vision models fail at basic visual reasoning, relying on pattern recognition over spatial understanding.
“there's actually a big gap between how these models handle visual thinking and how humans do it”
DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.
“Throw that all away. Out. Right now.”
Gemma 4 E2B runs fully on-device for multimodal inference on robot ducks via Raspberry Pi 5 and Jetson Orin Nano.
“And through multimodal inputs of Gemma 4, they're able to process and understand their environment like never before.”
AMD acquires World Labs for $8.2B, gaining spatial AI and Atlas sparse reconstruction model
“Recently we released Atlas, a first of its kind omni model architecture that solves a key outstanding problem in spatial intelligence: new camera view prediction.”
AI generates a solution to the Navier–Stokes Millennium Prize Problem.
MiniMax's M3 model features multimodal capabilities with a 1 million token context.
“this model can not only work with code, but it can understand video, images, and it has a super long context of 1 million.”
Sam Altman expects OpenAI to declare AGI achieved internally by December 2026.
“Automated AI Research Intern”
Moonshot AI releases Kimi K3, a 2.8T-parameter open model claiming frontier-class performance
“Open Frontier Intelligence”
Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches
“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Claude Opus 5.5 ships, leads SimpleBench at 88.4% and dominates explainer video creation
DeepMind launches new institute for AGI research and governance.
“@demishassabis and @ShaneLegg launched the DeepMind Institute, a new in-house platform for interdisciplinary research and debate on AGI governance, economics, transparency, and human flourishing.”
Mass Magnetics aims to recycle rare earth magnets for robotics and defense.
“A huge portion of the world's magnets will come from recycled materials.”
NVIDIA introduces EPD disaggregation for multimodal model optimization.
AI tokenomics and understanding token value are crucial as model costs rise.
AlphaGenome Atlas predicts effects of 9 billion DNA variations.
OpenAI announces a $5 million grant for research on AI's impact on teens.
“Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.”
Atlas is a new next-generation model capable of high-quality world simulation.
“Yesterday was a big day. You launched a new cutting-edge model that received a great reception.”
Hugging Face introduces NeoMME, a novel multimodal-native and multilingual encoder.
“NeoMME demonstrates a new way to seamlessly integrate multiple modalities in AI.”
Zhipu's GLM 5.3 Flash anonymously dominated Open Router and costs 40x less than Claude
“at its peak, OX Alpha accounted for nearly a third of Open Router's entire weekly traffic”
Qwen releases 2.4T open-weight model with autonomous 10-day coding and AI research capabilities
Google DeepMind launches Gemini Robotics 2 with whole-body intelligence, dexterity, and multi-robot collaboration.
“What we're building here is the intelligence layer to power any robot to do a broad range of useful tasks.”
NVIDIA declares physical AI has arrived after 15 years, unveiling full robotics stack and Japan mechatronics partnership.
“After 15 years of work, physical AI is here.”
Google DeepMind launches Gemini Robotics ER 2 with multi-robot collaboration and video understanding.
Thinking Machines releases Inkling, a 970B-parameter Apache-licensed open-weights multimodal model.
“the weights are Apache licensed and sitting on hugging face right now”
NVIDIA Cosmos is an open foundation model generating synthetic data for physical AI training.
“For physical AI, compute is data.”
Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.
“The North Star of our smaller models is intelligence per byte of memory footprint.”
NVIDIA's ENPIRE framework lets coding agents run a closed-loop self-improvement process for real-world robots, hitting 99% on dexterous tasks.
“Frontier coding agents can autonomously develop a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks in the real world, such as PushT, organizing pins into a pin box, and using a cutter to cut a zip tie,”
NVIDIA's Jensen Huang frames the AI factory as the largest infrastructure buildout in history and the best enterprise investment of the next decade.
“you can't think if you don't generate words”
Google DeepMind announces Gemma 4 12B, a unified encoder-free multimodal open model.
NVIDIA unveils Cosmos 3, an open omnimodel for physical AI that perceives, generates, and acts.
“Cosmos, the foundation for developers of the age of physical AI.”
NVIDIA launched Cosmos 3, Nemotron 3 Ultra, and previewed the RTX Spark superchip at Computex.
NVIDIA releases Cosmos 3, billed as the first open omni-model for physical AI reasoning and action.
Google releases EmbeddingGemma 2, a sub-billion open model unifying text, image, video, and audio embeddings on-device.
“a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in a shared high-dimensional embedding space”
General-purpose LLM agents may outperform specialized models for robot control
“We are working on creating LLMs that drive robots”
Skydio uses agent orchestration to let one operator control large fleets of autonomous drones simultaneously.
“Thousands of these drones are now deployed across the country—at energy companies, public safety agencies, and construction firms.”
Dyna Robotics is building reliable, commercially deployable general-purpose manipulation policies, not just demos.
“Robot Demos Are Easy. Reliability Is Hard”
Robotics has appeared 'almost here' for 70 years despite current physical-AI hype reaching its peak.
“robotics has been "almost here" for the last 70 years.”
Perceptron AI wants to unify VLMs, VLAs, and world models into 'embodied foundational models' for real-time physical-world interaction.
“we want to move away from the distinction between VLM, VLA, world models, whatever you want to call it, to what we call embodied fundamental models”
AI researchers can prioritize experiments based on preference models.