SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents
“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents
“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
xAI's Grok 4.7 lands on Amazon Bedrock with 500K context and self-verification for agents
“A model that checks its own output before continuing tends to fail less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step.”
OpenAI DevDay 2026 launched Dots always-on agents, GPT-6.1 Sol, and major platform APIs
AI models cross binary exploitation threshold; earlier models including Claude Opus 4.6 had zero successes.
“a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.”
OpenAI launches GPT-6 Sol and Luna, two frontier models balancing capability and cost.
Z.ai's GLM-5.3 surpasses Claude Fable 5 and GPT-5.6-Sol on select benchmarks
“On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.”
Greg Brockman confirms ChatGPT Chat and Work modes will merge by end of 2026
Moonshot AI's Kimi K3 is a 2.8T-parameter open-weight model matching frontier closed models on coding benchmarks.
“it has OpenAI and Anthropic terrified because its Trust Me Bro benchmark performance is on par with and in some cases beating Claude Fable and GPT 5.6 Soul”
Kimi K3 is the strongest open-weights model ever, closing the US-China gap to 3-5 months
“the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated 6-9 months to something shorter, say 3-5 months.”
AI systems are now reliably more persuasive than expert humans in real-world text-based persuasion.
“AI systems were reliably more persuasive than expert humans, even when expert humans chose their issues, researched in advance, underwent hours of live, structured practice, and were incentivized with £1,000 cash bonuses”
The U.S. government forced Anthropic to suspend foreign access to its Claude Fable/Mythos models, opening a new AI governance era.
“The executive branch of the United States forcing Anthropic to turn off access — both internally and externally — to their latest Claude 5 Mythos/Fable models is the starting gun of a new era in AI governance.”
Anthropic released Claude Fable 5, its smartest public model, paired with uneven heavy-handed safety controls.
“Claude Fable 5 is definitely the smartest model available to the general public”
Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches
“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
OpenAI frontier model solves open math problems, releases formal Lean proofs on GitHub
Claude Opus 5.5 ships, leads SimpleBench at 88.4% and dominates explainer video creation
Meta's Muse is the first consumer-accessible agentic AI with persistent Linux VMs per user
“It's the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it's a genuinely open question whether consumers have any understanding what this means.”
A Chinese open-weight model claims frontier throne as Anthropic, OpenAI cut prices and TypeSafe AI raises $10B
“It's been an absolutely MONSTER week already”
Alibaba releases open weights for 2.4T-parameter Qwen3.8-Max, its largest open-weight model
Kimi K3's open-weights 2.8T parameter model matches frontier quality at significantly lower API cost
“even if you don't ever use it, it will be pushing token prices down”
YC's Garry Tan claims 400x personal coding productivity gain with AI agents.
“One person does what used to take a thousand people. And I don't mean that as a metaphor. I mean that mechanically this year, the people in this room will do this.”
Databricks launched Omnigent, an open-source 'meta-harness' layer on top of the agentic stack to make agents effective at scale.
“we call it a meta harness of harnesses, if you know what an agent harness is.”
Databricks open-sources Omnigent, a meta-harness to combine, control, and share agents across Claude Code, Codex, Cursor, and more.
“CDC is brittle enough to joke that it means “continuous data corruption,””
Google DeepMind adds computer use capabilities to its Gemini 3.5 Flash model.
Clay runs over 350 million go-to-market AI agents monthly, processing trillions of tokens per week.
“We run this over 350 million times a month. It processes trillions of tokens every week.”
LangChain launched Engine, an agent that autonomously investigates traces and drafts PRs to improve other agents.
“We're working towards a future where agents improve themselves.”
Z.ai's open-weight GLM-5.2 marks a step-change for open agentic models, rivaling top labs.
“minor version numbers can have AI models crossing meaningful user experience thresholds”
Microsoft Foundry assembles agents, tuning, evals, and OpenEnv into an owned reinforcement-learning loop that improves over time.
“the durable asset is not the model you rent, it is the learning loop you own”
AI lets enterprises finally capture human capital and tacit knowledge, but humans stay valuable by finding gaps.
“Every company is going to have the human capital that is still going to be super valuable because humans and their ability to find the gaps that exist at all times is going to be the way we all will create value”
Google launches Anti-gravity agentic platform and Gemma 4 hits 100M downloads in first month
“It's our smartest open model yet. It's purposebuilt for advanced reasoning, agentic workflows, and the response has been incredible. 100 million downloads in the first month and it's pushing Gemma downloads past half a billion.”
All major AI model labs are now also building agents as their core product
LangChain achieved 64% cost reduction in their coding agent via model routing with no quality loss
“we were able to see a 64% reduction in median cost per thread with no measurable change in quality”
Claude Cowork moves model inference and VM execution to the cloud for mobile and battery improvements
“The "new" version of Cowork runs model inference and the VM in the cloud. Each session gets its own sandbox, not sharing state with other sessions.”
OpenAI's computer use agents are now '180 degrees different,' approaching superhuman software operation speed
“180 degrees different”
GEPA proposes reflective optimization in text space to overcome RL's sample inefficiency
“instead of using only a zero or one reward signal, we can make a language model or agent analyze the entire execution process to understand what worked and what didn't”
LangChain launches LangSmith fine-tuning in public beta with SmithTune, a CLI to post-train models from agent traces.
“today we're launching LangSmith fine-tuning in public beta with SmithTune, a CLI to allow you to post-train models from your LangSmith traces in one workflow”
Strands Evals and Amazon Bedrock AgentCore add skill-focused evaluators to measure agent skill selection and instruction following.
“A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain-specific task like redacting a contract, reconciling an invoice, or following a team’s pull-request conventions.”
GPT-6 Astra let Parallel's agents research labor-market data in half the time and cost.
OpenAI publishes priorities and principles for independent third-party safety assessments of frontier models.
Lambert argues true recursive self-improvement won't arrive soon; current AI-safety anxiety reflects scaled agents, not imminent superintelligence.
“they’ll turn out to be directionally correct (relative to the expectations of almost anyone not linked to the community) but factually wrong.”
Files are replacing Python for building agents.