AI models cross binary exploitation threshold; earlier models including Claude Opus 4.6 had zero successes.
“a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.”
23 tracked signals on frontier-models.
AI models cross binary exploitation threshold; earlier models including Claude Opus 4.6 had zero successes.
“a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.”
OpenAI launches GPT-6 Sol and Luna, two frontier models balancing capability and cost.
Z.ai's GLM-5.3 surpasses Claude Fable 5 and GPT-5.6-Sol on select benchmarks
“On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.”
SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents
“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
Moonshot AI's Kimi K3 is a 2.8T-parameter open-weight model matching frontier closed models on coding benchmarks.
“it has OpenAI and Anthropic terrified because its Trust Me Bro benchmark performance is on par with and in some cases beating Claude Fable and GPT 5.6 Soul”
Kimi K3 is the strongest open-weights model ever, closing the US-China gap to 3-5 months
“the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated 6-9 months to something shorter, say 3-5 months.”
AI systems are now reliably more persuasive than expert humans in real-world text-based persuasion.
“AI systems were reliably more persuasive than expert humans, even when expert humans chose their issues, researched in advance, underwent hours of live, structured practice, and were incentivized with £1,000 cash bonuses”
The U.S. government forced Anthropic to suspend foreign access to its Claude Fable/Mythos models, opening a new AI governance era.
“The executive branch of the United States forcing Anthropic to turn off access — both internally and externally — to their latest Claude 5 Mythos/Fable models is the starting gun of a new era in AI governance.”
Anthropic released Claude Fable 5, its smartest public model, paired with uneven heavy-handed safety controls.
“Claude Fable 5 is definitely the smartest model available to the general public”
OpenAI frontier model solves open math problems, releases formal Lean proofs on GitHub
Claude Opus 5.5 ships, leads SimpleBench at 88.4% and dominates explainer video creation
A Chinese open-weight model claims frontier throne as Anthropic, OpenAI cut prices and TypeSafe AI raises $10B
“It's been an absolutely MONSTER week already”
Alibaba releases open weights for 2.4T-parameter Qwen3.8-Max, its largest open-weight model
Kimi K3's open-weights 2.8T parameter model matches frontier quality at significantly lower API cost
“even if you don't ever use it, it will be pushing token prices down”
xAI's Grok 4.7 lands on Amazon Bedrock with 500K context and self-verification for agents
“A model that checks its own output before continuing tends to fail less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step.”
OpenAI publishes priorities and principles for independent third-party safety assessments of frontier models.
OpenAI launches Daybreak program giving vetted partners access to frontier cyber models
AI labs face a narrow post-release window to recoup frontier model costs, and US data center buildout assumes a global market for US AI services.
“No one is building $100 billion dollar data centers to serve frontier models to whatever 100 companies the US government will allow access.”
Google DeepMind announced Gemini 4 Argon as their next frontier intelligence era
AI models can solve hard math problems when given human-provided intuitive hints
“Maybe some of that understanding resides in model weights. To me, that's like pretty unsatisfying.”
Databricks aims to make enterprise AI work by feeding organizational data, processes, and tacit knowledge as context to frontier and open-source models.
“If we just give the context to these very smart models that can solve those super hard problems, they'll be able to do amazing things inside of our companies.”
Mistral Large 4 released, benchmarks criticized as saturated by frontier models
“The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.”
Frontier models handle big problems but local models offer control and privacy
“you just want some control, you want some privacy, you just want to keep things local and make my stuff my stuff”