The Hallway Track

ICLR 2026

research reasoning multimodal deep-learning

Dates
2026-04-24 → 2026-04-28
Location
Rio de Janeiro, Brazil
Ecosystem
frontier research
Importance
9/10

Related coverage & signals

Ten advances in mathematics and theoretical computer science

Simon Willison · Aug 01, 2026

OpenAI's unreleased Astra model solved 10 decade-old unsolved math problems for under $2,000 each.

“He envisions a future of large-scale, decentralized collaborations between humans and machines, where complex mathematical tasks can be diced and sliced, with humans claiming the creative parts and AI doing the lion's share of the technical grunt work.”
[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0

Latent Space Blog · May 20, 2026

Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches

“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Funding grants for new research into AI and teen development

OpenAI · OpenAI Blog · Sep 08, 2026

OpenAI announces a $5 million grant for research on AI's impact on teens.

“Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.”
Gemma 4 in Action: Bringing Frontier AI to the Edge

Google Developers (Google I/O) · Jun 29, 2026

Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.

“The North Star of our smaller models is intelligence per byte of memory footprint.”
Introducing Mistral Large 4: Le chonk

Simon Willison · Oct 06, 2026

Mistral releases 1 trillion parameter Large 4 model, reclaiming competitive position

“it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier”
Introducing EmbeddingGemma 2: An open model for natively multimodal embeddings

Google Developers (Google I/O) · Oct 06, 2026

Google releases EmbeddingGemma 2, a sub-billion open model unifying text, image, video, and audio embeddings on-device.

“a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in a shared high-dimensional embedding space”
Introducing Hy4 Preview

Simon Willison · Aug 29, 2026

Tencent releases Hy4, a 770B open-weight LLM with 1M token context window

Leverage Private Cloud Compute in your app

Apple Developer (WWDC) · Aug 25, 2026

Apple launches zero-cost server LLM for developers via Private Cloud Compute with no API keys

“there are no token costs to you, the developer”
DeepSeek V4 Pro 0813 (on OpenRouter)

Simon Willison · Aug 12, 2026

DeepSeek V4 Pro 0813 launches with 1.7T open weights and reasoning-level-dependent outputs.

“I've not noticed this kind of difference from any other model”
Image understanding with on-device AI

Apple Developer (WWDC) · Aug 11, 2026

Apple's on-device Foundation Models framework gains multimodal image understanding and Vision Framework integration.

“This opens up new categories of experiences you can build with image understanding. It's as simple as attaching an image to your prompt.”
DeepMind Just Changed How AI Sees The World

Two Minute Papers · Aug 07, 2026

DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.

“Throw that all away. Out. Right now.”
HTML Is All Agents Need — James Russo, HeyGen

Andrej Karpathy · AI Engineer · Jul 21, 2026

HeyGen uses HTML as the native canvas for AI agents to generate full video compositions

“Uh when you try to teach a model a new DSL or even your own custom JSON structure, it's forcing it to speak another language.”
Scientists Found A Better Language For AI Agents

Two Minute Papers · Jun 19, 2026

A new method lets AI agents communicate by passing raw undecoded numbers instead of English text.

“forget English. You know what? Forget letters entirely.”
Gemma Playground: AI Edge Gallery

Google Developers (Google I/O) · Jun 18, 2026

Google's Gemma models run entirely on-device on phones, enabling multimodal AI, agent skills, and offline use.

“And what's also important to remember is that this is running entirely on the device. So, it will work offline or in areas of low connectivity.”