MiniMax's M3 model features multimodal capabilities with a 1 million token context.
“this model can not only work with code, but it can understand video, images, and it has a super long context of 1 million.”
58 tracked signals on multimodal.
MiniMax's M3 model features multimodal capabilities with a 1 million token context.
“this model can not only work with code, but it can understand video, images, and it has a super long context of 1 million.”
Moonshot AI releases Kimi K3, a 2.8T-parameter open model claiming frontier-class performance
“Open Frontier Intelligence”
Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches
“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
NVIDIA introduces EPD disaggregation for multimodal model optimization.
Hugging Face introduces NeoMME, a novel multimodal-native and multilingual encoder.
“NeoMME demonstrates a new way to seamlessly integrate multiple modalities in AI.”
Zhipu's GLM 5.3 Flash anonymously dominated Open Router and costs 40x less than Claude
“at its peak, OX Alpha accounted for nearly a third of Open Router's entire weekly traffic”
Qwen releases 2.4T open-weight model with autonomous 10-day coding and AI research capabilities
Black Forest Labs launches FLUX 3 Video, beating Seedance 2.0, Gemini Omni, and Grok Imagine
Thinking Machines releases Inkling, a 970B-parameter Apache-licensed open-weights multimodal model.
“the weights are Apache licensed and sitting on hugging face right now”
Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.
“The North Star of our smaller models is intelligence per byte of memory footprint.”
Google DeepMind announces Gemma 4 12B, a unified encoder-free multimodal open model.
NVIDIA launched Cosmos 3, Nemotron 3 Ultra, and previewed the RTX Spark superchip at Computex.
Google releases EmbeddingGemma 2, a sub-billion open model unifying text, image, video, and audio embeddings on-device.
“a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in a shared high-dimensional embedding space”
Today's top vision models fail at basic visual reasoning, relying on pattern recognition over spatial understanding.
“there's actually a big gap between how these models handle visual thinking and how humans do it”
Google launches Pics, an AI image creation and editing tool for Workspace built on Nano Banana model
“Built on our latest Nano Banana model, Google Pics — our image creation and editing tool — is now available.”
Google DeepMind launched Gemini Omni 1.1 Flash with enhanced developer control features
Apple's on-device Foundation Models framework gains multimodal image understanding and Vision Framework integration.
“This opens up new categories of experiences you can build with image understanding. It's as simple as attaching an image to your prompt.”
DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.
“Throw that all away. Out. Right now.”
Google DeepMind launches Lyria 3.5 in Flow Music with improved musicality, lyrics, vocals, and creative control
HeyGen uses HTML as the native canvas for AI agents to generate full video compositions
“Uh when you try to teach a model a new DSL or even your own custom JSON structure, it's forcing it to speak another language.”
Google's Gemma models run entirely on-device on phones, enabling multimodal AI, agent skills, and offline use.
“And what's also important to remember is that this is running entirely on the device. So, it will work offline or in areas of low connectivity.”
Nemotron 3.5 Content Safety launches as a customizable multimodal safety model for global enterprise AI.
Google showcases demos of newly announced Gemini Omni and Gemini 3.5 from Google I/O 2026.
Google I/O 2026 keynote featured Gemini Omni and Gemini 3.5 Flash among 12 major announcements
Google launches Gemini 3.5 Flash and Omni multimodal model under 'intelligence with action' framing
“the phrase we're using to talk about Gemini 3.5, which is one of our big releases, is intelligence with action”
Google DeepMind releases new audio models including TTS, live translation, and Gemini Live upgrades
“Our new text-to-speech models are capable of doing really compelling voice personalization, which means that you can create a voice and guide it to have specific emotions, specific resonance, and even to incorporate things like pauses.”
Google DeepMind releases EmbeddingGemma 2, an open lightweight multimodal embedding model
Condé Nast cut video discovery time from 250 minutes to under 2 minutes using multimodal AI search
Gemini's agentic video understanding selectively samples transcripts and frames to cut token costs while boosting accuracy.
“not only does it reduce the number of tokens it really needs, the performance increases because it can zoom in on certain functions and really pay attention to the things in the video that are actually important to the query”
Gemini uses agentic tool-calling loops to selectively process video segments instead of entire videos
“With agentic video understanding, we don't give the model the entire video. We give the reference to the video and a model can then decide”
Qwen releases 125B MoE model with only 6B active parameters as Qwen4 architecture preview
“a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4”
Google launches Gemini 3.5 Transcribe, its first LLM-based transcription model with live streaming support.
“this model is both available for unary on the interactions API but also for life transcription on the life API”
Apple Xcode agents can now interpret hand-drawn sketches and generate matching UI code
“Xcode took my image and used it as inspiration to create this chart.”
Azure Content Understanding adds GPT-5 series support with 25% cost reduction and improved grounding accuracy.
“cost decrease by up to 25% while improving confidence scores accuracy and grounding accuracy by up to 14% and 3% respectively”
Databricks introduces a native FILE column type for storing multimodal data in tables
MiniMax-H3 omni-modal model now runs locally on Apple Silicon via MLX port
“a rainbow colored skunk leaps over a mossy log in a supermarket”
NVIDIA releases Alpamayo 2 Super, a 34B-parameter reasoning vision model for autonomous vehicles
Amazon acquired conversational AI firm Anodot, now part of Amazon Connect, as United Airlines deploys multimodal agentic automation for CX
“the ability our ability to speak to computers and be understood contextually and semantically has never been more capable than it is today”
TwelveLabs argues video AI needs a memory layer treating video as spatial-temporal volumes, not frame sequences
“video is not a bag of frames”
Pairing Amazon Nova 2 Lite with Claude Sonnet 4.6 in a two-model Bedrock pipeline cuts per-page document-digitization cost by about two-thirds.
“This two-model approach costs about two-thirds less per page than a single-model alternative that sends the entire task to one vision-language model.”
Audio-in, visuals-out AI experiences are now feasible, matching Karpathy's claim that voice is the preferred input and visuals the preferred output.
“voice is the human preferred input for AIs. But that we prefer visuals as the output.”
Hugging Face introduces MolmoMotion, a model for language-guided 3D motion forecasting.
NVIDIA launches XR AI, a reusable foundation for building AI agents on AR glasses and XR devices.
“The hardware is ready, but creating AI experiences requires integrating live camera and microphone streams, multimodal AI models, enterprise data, tool use, deployment infrastructure, and device-specific runtimes.”
MiniMax M3 launches on NVIDIA Blackwell infrastructure as a single multimodal system for long-context reasoning and agentic workflows.
“MiniMax M3—available on NVIDIA accelerated infrastructure including NVIDIA Blackwell—changes this by enabling a single multimodal system capable of long-context reasoning”
Gemma 4 E2B runs fully on-device for multimodal inference on robot ducks via Raspberry Pi 5 and Jetson Orin Nano.
“And through multimodal inputs of Gemma 4, they're able to process and understand their environment like never before.”
Real-time multimodal voice agents are bottlenecked by media infrastructure, not models, which LiveKit handles via WebRTC on Azure.
“You see models they are the easy part. Now everything around it is the hard part.”
StepFun's Step 3.7 Flash 198B multimodal model is now available on NVIDIA-accelerated infrastructure
“AI applications are moving beyond text generation to multimodal systems that can perceive, search, and reason across images, documents, video, and language in real time”
Google launches Gemini 3.1 Flash Live as its latest multimodal real-time conversational model
“Google AI Studio is really the best place to get started with all of the Google DeepMind models.”
AWS enables streaming TTS on SageMaker via vLLM-Omni DLC for real-time voice apps
AWS vLLM-Omni DLC enables unified text-to-image-to-video pipelines on SageMaker AI
SkyRL on SageMaker HyperPod trains vision-language models to 95% maze-solving accuracy via GRPO
Gemini 3.5 live translation model enables real-time multi-language broadcast via API
“This is using the Gemini API together with Life Kit and Google Cloud Run to deploy an application that can broadcast speech in many, many different languages using Gemini live translate.”
Azure Content Understanding turns messy multimodal content into clean, structured, agent-ready output via one pipeline.
“They can reason, but they cannot really read.”
Amazon Nova 2 Lite enables zero-shot object detection via natural language prompts with no model training required.
“This multimodal foundation model detects objects through natural language prompts with no training required.”
Developer built a voice-to-physical-printout lamp using Gemini audio analysis capabilities
“There's something very magical about being able to make an idea tangible so quickly, and make something entirely unique for myself.”
Claude Opus 5.5 generates pixel art animations from photos via simple prompts
“I had seen some buzz around how good Claude Opus 5.5 was at creating pixel art animations.”
Google DeepMind announces agentic video understanding capabilities for Gemini.
Developer builds open-source family morning dashboard using Gemini 3.6 Flash for daily scheduling
“I'm not a programmer. I'm not a dev. I'm a VI code guy. But with Anti-Gravity and Gemini, this stuff is way easier”