NVIDIA releases Alpamayo 2 Super, a 34B-parameter reasoning vision model for autonomous vehicles
ICLR 2026
- Dates
- 2026-04-24 → 2026-04-28
- Location
- Rio de Janeiro, Brazil
- Ecosystem
- frontier research
- Importance
- 9/10
Related coverage & signals
AI generates a solution to the Navier–Stokes Millennium Prize Problem.
MiniMax's M3 model features multimodal capabilities with a 1 million token context.
“this model can not only work with code, but it can understand video, images, and it has a super long context of 1 million.”
OpenAI's unreleased Astra model solved 10 decade-old unsolved math problems for under $2,000 each.
“He envisions a future of large-scale, decentralized collaborations between humans and machines, where complex mathematical tasks can be diced and sliced, with humans claiming the creative parts and AI doing the lion's share of the technical grunt work.”
Moonshot AI releases Kimi K3, a 2.8T-parameter open model claiming frontier-class performance
“Open Frontier Intelligence”
Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches
“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
OpenAI model disproves 80-year-old unit distance conjecture in discrete geometry
DeepMind launches new institute for AGI research and governance.
“@demishassabis and @ShaneLegg launched the DeepMind Institute, a new in-house platform for interdisciplinary research and debate on AGI governance, economics, transparency, and human flourishing.”
NVIDIA introduces EPD disaggregation for multimodal model optimization.
AI tokenomics and understanding token value are crucial as model costs rise.
AlphaGenome Atlas predicts effects of 9 billion DNA variations.
OpenAI announces a $5 million grant for research on AI's impact on teens.
“Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.”
Hugging Face introduces NeoMME, a novel multimodal-native and multilingual encoder.
“NeoMME demonstrates a new way to seamlessly integrate multiple modalities in AI.”
Zhipu's GLM 5.3 Flash anonymously dominated Open Router and costs 40x less than Claude
“at its peak, OX Alpha accounted for nearly a third of Open Router's entire weekly traffic”
Qwen releases 2.4T open-weight model with autonomous 10-day coding and AI research capabilities
Two API settings tripled OpenAI's GPT-5.6 scores on the ARC-AGI-3 benchmark
Black Forest Labs launches FLUX 3 Video, beating Seedance 2.0, Gemini Omni, and Grok Imagine
Thinking Machines releases Inkling, a 970B-parameter Apache-licensed open-weights multimodal model.
“the weights are Apache licensed and sitting on hugging face right now”
Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.
“The North Star of our smaller models is intelligence per byte of memory footprint.”
NVIDIA's Jensen Huang frames the AI factory as the largest infrastructure buildout in history and the best enterprise investment of the next decade.
“you can't think if you don't generate words”
Google DeepMind announces Gemma 4 12B, a unified encoder-free multimodal open model.
NVIDIA launched Cosmos 3, Nemotron 3 Ultra, and previewed the RTX Spark superchip at Computex.
Mistral releases 1 trillion parameter Large 4 model, reclaiming competitive position
“it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier”
Google releases EmbeddingGemma 2, a sub-billion open model unifying text, image, video, and audio embeddings on-device.
“a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in a shared high-dimensional embedding space”
OpenAI publishes official GPT-6 family model guide for startup production workflows
Today's top vision models fail at basic visual reasoning, relying on pattern recognition over spatial understanding.
“there's actually a big gap between how these models handle visual thinking and how humans do it”
AI researchers can prioritize experiments based on preference models.
Google launches Pics, an AI image creation and editing tool for Workspace built on Nano Banana model
“Built on our latest Nano Banana model, Google Pics — our image creation and editing tool — is now available.”
Tencent releases Hy4, a 770B open-weight LLM with 1M token context window
Google DeepMind launched Gemini Omni 1.1 Flash with enhanced developer control features
Apple launches zero-cost server LLM for developers via Private Cloud Compute with no API keys
“there are no token costs to you, the developer”
DeepSeek V4 Pro 0813 launches with 1.7T open weights and reasoning-level-dependent outputs.
“I've not noticed this kind of difference from any other model”
Apple's on-device Foundation Models framework gains multimodal image understanding and Vision Framework integration.
“This opens up new categories of experiences you can build with image understanding. It's as simple as attaching an image to your prompt.”
DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.
“Throw that all away. Out. Right now.”
Google DeepMind launches Lyria 3.5 in Flow Music with improved musicality, lyrics, vocals, and creative control
HeyGen uses HTML as the native canvas for AI agents to generate full video compositions
“Uh when you try to teach a model a new DSL or even your own custom JSON structure, it's forcing it to speak another language.”
A new method lets AI agents communicate by passing raw undecoded numbers instead of English text.
“forget English. You know what? Forget letters entirely.”
Google's Gemma models run entirely on-device on phones, enabling multimodal AI, agent skills, and offline use.
“And what's also important to remember is that this is running entirely on the device. So, it will work offline or in areas of low connectivity.”
Nemotron 3.5 Content Safety launches as a customizable multimodal safety model for global enterprise AI.
Google showcases demos of newly announced Gemini Omni and Gemini 3.5 from Google I/O 2026.