The Hallway Track

model-release

21 tracked signals on model-release.

Introducing Claude Opus 5

Simon Willison · Jul 24, 2026

Anthropic releases Claude Opus 5, topping the Artificial Analysis leaderboard at Opus 4.8 pricing

“On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly viewthe drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part.”
Quoting OpenAI

Simon Willison · Jun 26, 2026

OpenAI previews the GPT-5.6 series (Sol, Terra, Luna) with lower pricing and government-coordinated limited release.

“We previewed our plans and the models' capabilities ahead of today's launch. At their request, we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government, before releasing more broadly.”
[AINews] not much happened today

Latent Space Blog · Oct 03, 2026

Anthropic models now hold the top three Agent Arena spots as Sonnet 5.5 debuts at #3

“good, cheap AND fast”
Claude Sonnet 5.5

Simon Willison · Sep 28, 2026

Anthropic releases Claude Sonnet 5.5, 30% faster and cheaper than Sonnet 5 with better benchmarks

“runs 30%+ faster, and costs up to 30% less for most work”
Did Anthropic just kill the indie hacker...?

Fireship · Jul 29, 2026

Claude Opus 5 delivers near Fable-level intelligence at half the price with autonomous error recovery

“Anthropic says it also verifies its own work and recovers from its own mistakes without human intervention.”
What's new in Google AI

Google Developers (Google I/O) · May 23, 2026

Google released Gemini 3.5 Flash at I/O alongside expanded Gemini 3 model lineup

“an insane amount of pace. I think if you just look at 2024 when we had the 1.5 series of models, we were just cracking the nut on multimodality and look where we are now.”
Introducing Mistral Large 4: Le chonk

Simon Willison · Oct 06, 2026

Mistral releases 1 trillion parameter Large 4 model, reclaiming competitive position

“it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier”
Introducing Claude Sonnet 5.5 on AWS

AWS Machine Learning Blog · Sep 28, 2026

Claude Sonnet 5.5 launches on Amazon Bedrock with lower cost and faster speed for coding workloads

“Claude Sonnet 5.5 takes the work where the approach is already clear and what's left is to execute it quickly, with a lower cost per task (than Opus 5.5) for most work at faster speed.”
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding

Simon Willison · Jun 29, 2026

DeepReinforce released Ornith-1.0, an MIT-licensed open-weights agentic coding model built on Gemma 4 and Qwen 3.5.

“Initial impressions are very good - it seems to be able to run the agent harness over many tool calls in a proficient way.”
Claude Opus 4.8: "a modest but tangible improvement"

Simon Willison · May 28, 2026

Claude Opus 4.8 reduces hallucinations by being four times less likely to let code flaws pass unremarked

“Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. There's still more to be done: we're working on developing and releasing models that provide many of the same capabilities as Opus at a lower cost.”
What is Gemini 3.7 Flash?

Google Developers (Google I/O) · Aug 13, 2026

Google launches Gemini 3.7 Flash as its top coding and agent workhorse model

“our most intelligent workhorse model yet for coding and agents”
Introducing Gemini 3.7 Flash

Google DeepMind · Google DeepMind Blog · Aug 13, 2026

Google DeepMind announced Gemini 3.7 Flash, a new AI model.

Mistral Large 4

Simon Willison · Oct 06, 2026

Mistral Large 4 released, benchmarks criticized as saturated by frontier models

“The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.”