21 tracked signals on model-release.
Introducing Claude Opus 5
Simon Willison · Jul 24, 2026
Anthropic releases Claude Opus 5, topping the Artificial Analysis leaderboard at Opus 4.8 pricing
“On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly viewthe drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part.”
Quoting OpenAI
Simon Willison · Jun 26, 2026
OpenAI previews the GPT-5.6 series (Sol, Terra, Luna) with lower pricing and government-coordinated limited release.
“We previewed our plans and the models' capabilities ahead of today's launch. At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government, before releasing more broadly.”
[AINews] not much happened today
Latent Space Blog · Oct 03, 2026
Anthropic models now hold the top three Agent Arena spots as Sonnet 5.5 debuts at #3
“good, cheap AND fast”
Claude Sonnet 5.5
Simon Willison · Sep 28, 2026
Anthropic releases Claude Sonnet 5.5, 30% faster and cheaper than Sonnet 5 with better benchmarks
“runs 30%+ faster, and costs up to 30% less for most work”
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
Simon Willison · Sep 22, 2026
New model releases from OpenAI, Anthropic, xAI, and Xiaomi trigger an AI price war, with GPT-6 Luna at half its predecessor's cost.
“It's hard to overstate how competitive this pricing is.”
Did Anthropic just kill the indie hacker...?
Fireship · Jul 29, 2026
Claude Opus 5 delivers near Fable-level intelligence at half the price with autonomous error recovery
“Anthropic says it also verifies its own work and recovers from its own mistakes without human intervention.”
Introducing Claude Opus 5 on AWS: Anthropic’s most capable Opus model
AWS Machine Learning Blog · Jul 24, 2026
Claude Opus 5 launches on AWS Bedrock, matching Fable 5 intelligence at Opus-tier pricing
“Claude Opus 5 matches Claude Fable 5's top-tier intelligence in many domains at Opus-tier pricing.”
What's new in Google AI
Google Developers (Google I/O) · May 23, 2026
Google released Gemini 3.5 Flash at I/O alongside expanded Gemini 3 model lineup
“an insane amount of pace. I think if you just look at 2024 when we had the 1.5 series of models, we were just cracking the nut on multimodality and look where we are now.”
Introducing Mistral Large 4: Le chonk
Simon Willison · Oct 06, 2026
Mistral releases 1 trillion parameter Large 4 model, reclaiming competitive position
“it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier”
[AINews] Reflection Beam - 501B-A23B American Open Model
Latent Space Blog · Oct 06, 2026
Reflection AI launches Beam, a 501B/23B-active US-trained open MoE model for coding and agentic work
Introducing Claude Sonnet 5.5 on AWS
AWS Machine Learning Blog · Sep 28, 2026
Claude Sonnet 5.5 launches on Amazon Bedrock with lower cost and faster speed for coding workloads
“Claude Sonnet 5.5 takes the work where the approach is already clear and what's left is to execute it quickly, with a lower cost per task (than Opus 5.5) for most work at faster speed.”
Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock
AWS Machine Learning Blog · Sep 22, 2026
OpenAI's GPT-6 Sol and GPT-6 Luna are now generally available on Amazon Bedrock at lower pricing.
“On an internal factuality evaluation, OpenAI found that GPT-6 Sol made approximately half as many factual mistakes as GPT-5.6 Sol.”
deepseek-ai/DeepSeek-V4-Flash-0731
Simon Willison · Jul 31, 2026
DeepSeek V4 Flash offers best-in-class value at $0.14/M input tokens with 304B parameters
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google DeepMind · Google DeepMind Blog · Jul 21, 2026
Google DeepMind launches Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber models
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding
Simon Willison · Jun 29, 2026
DeepReinforce released Ornith-1.0, an MIT-licensed open-weights agentic coding model built on Gemma 4 and Qwen 3.5.
“Initial impressions are very good - it seems to be able to run the agent harness over many tool calls in a proficient way.”
Claude Opus 4.8: "a modest but tangible improvement"
Simon Willison · May 28, 2026
Claude Opus 4.8 reduces hallucinations by being four times less likely to let code flaws pass unremarked
“Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. There's still more to be done: we're working on developing and releasing models that provide many of the same capabilities as Opus at a lower cost.”
What is Gemini 3.7 Flash?
Google Developers (Google I/O) · Aug 13, 2026
Google launches Gemini 3.7 Flash as its top coding and agent workhorse model
“our most intelligent workhorse model yet for coding and agents”
Introducing Gemini 3.7 Flash
Google DeepMind · Google DeepMind Blog · Aug 13, 2026
Google DeepMind announced Gemini 3.7 Flash, a new AI model.
Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Olive Song
AI Engineer · Jul 31, 2026
MiniMax open-sourced its strongest model M3, with Together AI capturing lion's share of inference traffic.
“We do believe that the open source community as a whole is very strong and powerful.”
GLM-5.2: Built for Long-Horizon Tasks
Hugging Face · Hugging Face Blog · Jun 17, 2026
GLM-5.2 is a new model from Hugging Face optimized for long-horizon, multi-step agentic tasks.
Mistral Large 4
Simon Willison · Oct 06, 2026
Mistral Large 4 released, benchmarks criticized as saturated by frontier models
“The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.”