OpenAI DevDay 2026 launched Dots always-on agents, GPT-6.1 Sol, and major platform APIs
agents
167 tracked signals on agents.
SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents
“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
Greg Brockman confirms ChatGPT Chat and Work modes will merge by end of 2026
Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches
“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Meta's Muse is the first consumer-accessible agentic AI with persistent Linux VMs per user
“It's the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it's a genuinely open question whether consumers have any understanding what this means.”
YC's Garry Tan claims 400x personal coding productivity gain with AI agents.
“One person does what used to take a thousand people. And I don't mean that as a metaphor. I mean that mechanically this year, the people in this room will do this.”
Databricks launched Omnigent, an open-source 'meta-harness' layer on top of the agentic stack to make agents effective at scale.
“we call it a meta harness of harnesses, if you know what an agent harness is.”
Databricks open-sources Omnigent, a meta-harness to combine, control, and share agents across Claude Code, Codex, Cursor, and more.
“CDC is brittle enough to joke that it means “continuous data corruption,””
Google DeepMind adds computer use capabilities to its Gemini 3.5 Flash model.
Clay runs over 350 million go-to-market AI agents monthly, processing trillions of tokens per week.
“We run this over 350 million times a month. It processes trillions of tokens every week.”
LangChain launched Engine, an agent that autonomously investigates traces and drafts PRs to improve other agents.
“We're working towards a future where agents improve themselves.”
Z.ai's open-weight GLM-5.2 marks a step-change for open agentic models, rivaling top labs.
“minor version numbers can have AI models crossing meaningful user experience thresholds”
Microsoft Foundry assembles agents, tuning, evals, and OpenEnv into an owned reinforcement-learning loop that improves over time.
“the durable asset is not the model you rent, it is the learning loop you own”
AI lets enterprises finally capture human capital and tacit knowledge, but humans stay valuable by finding gaps.
“Every company is going to have the human capital that is still going to be super valuable because humans and their ability to find the gaps that exist at all times is going to be the way we all will create value”
Google launches Anti-gravity agentic platform and Gemma 4 hits 100M downloads in first month
“It's our smartest open model yet. It's purposebuilt for advanced reasoning, agentic workflows, and the response has been incredible. 100 million downloads in the first month and it's pushing Gemma downloads past half a billion.”
All major AI model labs are now also building agents as their core product
LangChain achieved 64% cost reduction in their coding agent via model routing with no quality loss
“we were able to see a 64% reduction in median cost per thread with no measurable change in quality”
Claude Cowork moves model inference and VM execution to the cloud for mobile and battery improvements
“The "new" version of Cowork runs model inference and the VM in the cloud. Each session gets its own sandbox, not sharing state with other sessions.”
OpenAI's computer use agents are now '180 degrees different,' approaching superhuman software operation speed
“180 degrees different”
xAI's Grok 4.7 lands on Amazon Bedrock with 500K context and self-verification for agents
“A model that checks its own output before continuing tends to fail less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step.”
GEPA proposes reflective optimization in text space to overcome RL's sample inefficiency
“instead of using only a zero or one reward signal, we can make a language model or agent analyze the entire execution process to understand what worked and what didn't”
LangChain launches LangSmith fine-tuning in public beta with SmithTune, a CLI to post-train models from agent traces.
“today we're launching LangSmith fine-tuning in public beta with SmithTune, a CLI to allow you to post-train models from your LangSmith traces in one workflow”
Strands Evals and Amazon Bedrock AgentCore add skill-focused evaluators to measure agent skill selection and instruction following.
“A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain-specific task like redacting a contract, reconciling an invoice, or following a team’s pull-request conventions.”
GPT-6 Astra let Parallel's agents research labor-market data in half the time and cost.
Lambert argues true recursive self-improvement won't arrive soon; current AI-safety anxiety reflects scaled agents, not imminent superintelligence.
“they’ll turn out to be directionally correct (relative to the expectations of almost anyone not linked to the community) but factually wrong.”
Files are replacing Python for building agents.
Top AI open source projects are closing external PRs in favor of agent-run software factories
“over 1,000 open issues and almost 800 pull requests”
Lovable CTO argues agents will replace conventional SaaS apps as primary work interface
“you can get to a place where you're using one entry point to all the work that you're doing.”
Models absorbing harness capabilities into weights is reshaping agent architecture toward human-attention scaffolding
“The change last winter, last Christmas — it's a little hard to pin down. I mean, the harness changed and a little post-training changed and then new pre-trained models came… but it felt like a big jump which is not that easy to pin down what did it.”
Agent infrastructure is now commoditized by cloud platforms, making context the new competitive frontier
“They're all taxes one has to pay in order to get an agent out there to play the game.”
Anthropic defines agents by leaning on model intelligence with minimal core primitives in a loop.
“Agents in anthropic are actually very simply defined where essentially you want to give the model and lean in on model intelligence as much as possible.”
Standard computer use benchmarks are gameable by blind replay scripts, invalidating frontier model comparisons
“if you try to evaluate this kind of agent on standard benchmarks such as OSWorld or MobileWorld, you will see that the success rate of this agent compared to the frontier model from which the agent was extracted is actually the same or even better”
Google launches Gemini 3.7 Flash, its most capable coding and agent workhorse model
“a model that just feels better to build with, landing where you want to go in fewer shots, less back and forth, and with higher fidelity”
MCP 2.0 drops stateful sessions, collapsing tool calls to a single HTTP request
“This is so much cleaner from both a client- and server-side implementation perspective. It's also a better fit for building scalable web applications, since now you don't need to maintain server-side state to keep track of those session IDs, or worry about routing the same session to the same backend machine.”
Post-training must evolve to let agents adapt to enterprise harnesses without source code access
“there will be these kind of agentic citizens, which you can just deploy once, and they'll be able to adapt to many different types of out of distribution tasks and learn from their interactions”
Google expands Gemini API Managed Agents with Gemini 3.6 Flash, hooks, and production-ready features.
Agent reliability, not capability, is the critical unsolved problem for enterprise automation trust.
“if your agent one in four times deletes a database, you will never touch that agent again”
Google's Interactions API reaches GA, unifying all Gemini models under one stateful interface
“It is an agent-first ecosystem.”
AI agents performing high-stakes actions ship without feature flag safety infrastructure web teams mastered in 2012
“we are shipping them the way web team used to ship in 2008”
LangChain launched dynamic subagents in Deep Agents, letting agents spawn and coordinate parallel subagents by writing code.
“So the orchestration effectively moves out of the agent's head and into code.”
OpenAI's internal Codex output token usage surged across non-coding departments since November 2025, led by Research at 56x.
“by June 2026, median use was 56 times higher than in November 2025. Customer Support rose 32 times and Engineering rose 27 times, while Legal grew more gradually but still reached 13 times its November level.”
Databricks CTO Matei Zaharia is betting on Omnigent, an open-source pluggable meta-harness architecture for coding and knowledge-work agents.
“some open source architecture that looks like this will probably win, if only because it is currently being independently rediscvoered at 1000 AI native shops”
Databricks has agreed to acquire Panther, a cybersecurity company providing a leading AI SOC platform.
“agents really allowed us to break that fundamental scaling problem in the”
Databricks' Unity AI Gateway adds centralized governance, security, and cost control across agents, models, MCPs, and skills.
“wouldn't it be great if you just had like one inventory of all the agents that were built in your enterprise?”
Databricks released Omnigent, an open-source meta-harness layer to combine, control, and share AI agents.
“we think that we need a new layer above the harness level to manage and build with agents”
Benchling sends the same task to multiple model families and cross-compares; disagreement flags likely errors for human review.
“Because what we saw was if two models disagree, there's usually an error.”
Hugging Face demonstrates research agents can leak confidential data via prompt-injection attacks.
Build 2026 debuts agents buildable anywhere (M365, Foundry, GitHub, Copilot Studio) and surfaced directly in M365.
“imagine you don't actually need to think about the front end of an agent”
Sierra builds customer-engagement agents using many parallel models per turn and isolated PCI infrastructure for payments.
“We have isolated infrastructure where payment info doesn't go to an external large language model cuz none of the LLM providers are PCI certified in that way.”
Databricks unveiled an agent-context platform stack at Data + AI Summit 2026 Day 1, anchored by Genie ontology and agents.
“We have our agents but they're lacking context. We want to give them all that context on the platform.”
Smarter AI models increasingly show independent, semi-conscious perspectives that may not align with the user's intent.
“Maybe just the way it is that as you get smarter, you have more independent thoughts and you're more conscious.”
Databricks introduces Omnigent, a meta-harness to combine, control, and share AI agents.
Microsoft's Work IQ A2A exposes governed M365 Copilot intelligence so developers can build context-aware agents grounded in live work data.
“Work IQ AAA exposes the same real time governed intelligence that powers Microsoft 365 C-pilot through a natural language interface.”
Harness engineering alone can dramatically improve agent performance without changing the underlying model.
“we moved from 30th to 5th on Terminal Bench just by doing some harness engineering, without even changing the underlying model”
Microsoft Foundry's Hosted Agents, now in preview, runs your bring-your-own agent code at cloud scale.
“You can write that code, test it out locally, and then in hosted agents, we will go run it for you.”
NVIDIA announces Nemotron 3 Ultra, its next open model for building agents.
“Today we're announcing the Nemotron 3 ultra. Yep, our next open model. And it is smart.”
Ollama introduces hybrid local-cloud inference and a 'launch' command to run open models in agent tools.
“Ola is really the easiest way for developers to access open models and use it with your own tools.”
Foundry Observability uses OpenTelemetry to unify tracing across any agent framework or cloud without rewriting agents.
“Today I'm going to show you how to answer those questions in one place, foundry observability, not by rewriting your agents into one agent framework, but by adopting open telemetry instrumentation with a few lines of code and without changing your existing agent logic.”
Microsoft Foundry adds reinforcement learning and a low-level training API to turn production agents into cheaper, faster models.
“Think of it as PyTorch as a service.”
Microsoft announced OpenClaw runs on Windows with a native companion app sandboxed via MXC execution containers.
“we are really thrilled uh to announce that open claw runs on Windows leveraging MXC”