The Hallway Track

frontier-models

23 tracked signals on frontier-models.

Quoting Anthropic Frontier Red Team

Simon Willison · Sep 29, 2026

AI models cross binary exploitation threshold; earlier models including Claude Opus 4.6 had zero successes.

“a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.”
Introducing GPT-6 Sol and Luna

OpenAI · OpenAI Blog · Sep 22, 2026

OpenAI launches GPT-6 Sol and Luna, two frontier models balancing capability and cost.

GLM-5.3: How Chinese labs keep stride with the frontier

Nathan Lambert · Interconnects · Aug 14, 2026

Z.ai's GLM-5.3 surpasses Claude Fable 5 and GPT-5.6-Sol on select benchmarks

“On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.”
[AINews] SpaceXAI Grok 4.6 and Grok @Bot

Elon Musk · Latent Space Blog · Aug 13, 2026

SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents

“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
Open-weight AI just hit 2.8 trillion parameters…

Fireship · Jul 22, 2026

Moonshot AI's Kimi K3 is a 2.8T-parameter open-weight model matching frontier closed models on coding benchmarks.

“it has OpenAI and Anthropic terrified because its Trust Me Bro benchmark performance is on par with and in some cases beating Claude Fable and GPT 5.6 Soul”
Kimi K3: The open-weights escalation

Nathan Lambert · Interconnects · Jul 20, 2026

Kimi K3 is the strongest open-weights model ever, closing the US-China gap to 3-5 months

“the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated 6-9 months to something shorter, say 3-5 months.”
Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI

Jack Clark · Import AI · Jun 22, 2026

AI systems are now reliably more persuasive than expert humans in real-world text-based persuasion.

“AI systems were reliably more persuasive than expert humans, even when expert humans chose their issues, researched in advance, underwent hours of live, structured practice, and were incentivized with £1,000 cash bonuses”
Welcome to the AGI era of AI governance

Nathan Lambert · Interconnects · Jun 14, 2026

The U.S. government forced Anthropic to suspend foreign access to its Claude Fable/Mythos models, opening a new AI governance era.

“The executive branch of the United States forcing Anthropic to turn off access — both internally and externally — to their latest Claude 5 Mythos/Fable models is the starting gun of a new era in AI governance.”
Claude Fable 5 and new AI safety fables

Nathan Lambert · Interconnects · Jun 09, 2026

Anthropic released Claude Fable 5, its smartest public model, paired with uneven heavy-handed safety controls.

“Claude Fable 5 is definitely the smartest model available to the general public”
[AINews] The Future of Latent Space

Latent Space Blog · Sep 25, 2026

A Chinese open-weight model claims frontier throne as Anthropic, OpenAI cut prices and TypeSafe AI raises $10B

“It's been an absolutely MONSTER week already”
Kimi K3 Just Broke The Economics Of AI

Two Minute Papers · Jul 29, 2026

Kimi K3's open-weights 2.8T parameter model matches frontier quality at significantly lower API cost

“even if you don't ever use it, it will be pushing token prices down”
Grok 4.7 is now available on Amazon Bedrock

AWS Machine Learning Blog · Sep 28, 2026

xAI's Grok 4.7 lands on Amazon Bedrock with 500K context and self-verification for agents

“A model that checks its own output before continuing tends to fail less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step.”
Quoting Dean W. Ball

David Sacks · Simon Willison · Jun 26, 2026

AI labs face a narrow post-release window to recoup frontier model costs, and US data center buildout assumes a global market for US AI services.

“No one is building $100 billion dollar data centers to serve frontier models to whatever 100 companies the US government will allow access.”
Can AI Learn Mathematical Intuition?

a16z · Sep 01, 2026

AI models can solve hard math problems when given human-provided intuitive hints

“Maybe some of that understanding resides in model weights. To me, that's like pretty unsatisfying.”
How to make AI work in the enterprise | Ali Ghodsi Co-founder and CEO of Databricks

Ali Ghodsi · Databricks · Jun 17, 2026

Databricks aims to make enterprise AI work by feeding organizational data, processes, and tacit knowledge as context to frontier and open-source models.

“If we just give the context to these very smart models that can solve those super hard problems, they'll be able to do amazing things inside of our companies.”
Mistral Large 4

Simon Willison · Oct 06, 2026

Mistral Large 4 released, benchmarks criticized as saturated by frontier models

“The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.”