No good standards exist for agent development; real-world continual learning is largely absent.
“there aren't really good standards for these things, right? Like there there's not some one-size-fits-all uh solution”
No good standards exist for agent development; real-world continual learning is largely absent.
“there aren't really good standards for these things, right? Like there there's not some one-size-fits-all uh solution”
OpenAI DevDay 2026 launched Dots always-on agents, GPT-6.1 Sol, and major platform APIs
SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents
“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
Greg Brockman confirms ChatGPT Chat and Work modes will merge by end of 2026
Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches
“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Meta's Muse is the first consumer-accessible agentic AI with persistent Linux VMs per user
“It's the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it's a genuinely open question whether consumers have any understanding what this means.”
AI's implementation isn't improving productivity in development teams.
“We used AI to solve a problem that AI created.”
Restate framework helps run resilient agents in production environments.
“We want to be able to bring it back and let it continue exactly at.”
Design skills are merging with engineering skills through tools like Impeccable.
“the role of the engineer and the design are blurring.”
AI has the potential to bridge the gap between design and engineering.
“I love AI.”
1Password improves engineering productivity by 21% using Codex.
YC's Garry Tan claims 400x personal coding productivity gain with AI agents.
“One person does what used to take a thousand people. And I don't mean that as a metaphor. I mean that mechanically this year, the people in this room will do this.”
Databricks launched Omnigent, an open-source 'meta-harness' layer on top of the agentic stack to make agents effective at scale.
“we call it a meta harness of harnesses, if you know what an agent harness is.”
Databricks open-sources Omnigent, a meta-harness to combine, control, and share agents across Claude Code, Codex, Cursor, and more.
“CDC is brittle enough to joke that it means “continuous data corruption,””
Google DeepMind adds computer use capabilities to its Gemini 3.5 Flash model.
Clay runs over 350 million go-to-market AI agents monthly, processing trillions of tokens per week.
“We run this over 350 million times a month. It processes trillions of tokens every week.”
LangChain launched Engine, an agent that autonomously investigates traces and drafts PRs to improve other agents.
“We're working towards a future where agents improve themselves.”
Z.ai's open-weight GLM-5.2 marks a step-change for open agentic models, rivaling top labs.
“minor version numbers can have AI models crossing meaningful user experience thresholds”
Microsoft Foundry assembles agents, tuning, evals, and OpenEnv into an owned reinforcement-learning loop that improves over time.
“the durable asset is not the model you rent, it is the learning loop you own”
AI lets enterprises finally capture human capital and tacit knowledge, but humans stay valuable by finding gaps.
“Every company is going to have the human capital that is still going to be super valuable because humans and their ability to find the gaps that exist at all times is going to be the way we all will create value”
Google launches Anti-gravity agentic platform and Gemma 4 hits 100M downloads in first month
“It's our smartest open model yet. It's purposebuilt for advanced reasoning, agentic workflows, and the response has been incredible. 100 million downloads in the first month and it's pushing Gemma downloads past half a billion.”
All major AI model labs are now also building agents as their core product
LangChain achieved 64% cost reduction in their coding agent via model routing with no quality loss
“we were able to see a 64% reduction in median cost per thread with no measurable change in quality”
Claude Cowork moves model inference and VM execution to the cloud for mobile and battery improvements
“The "new" version of Cowork runs model inference and the VM in the cloud. Each session gets its own sandbox, not sharing state with other sessions.”
OpenAI publishes official GPT-6 family model guide for startup production workflows
OpenAI's computer use agents are now '180 degrees different,' approaching superhuman software operation speed
“180 degrees different”
xAI's Grok 4.7 lands on Amazon Bedrock with 500K context and self-verification for agents
“A model that checks its own output before continuing tends to fail less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step.”
GEPA proposes reflective optimization in text space to overcome RL's sample inefficiency
“instead of using only a zero or one reward signal, we can make a language model or agent analyze the entire execution process to understand what worked and what didn't”
LangChain launches LangSmith fine-tuning in public beta with SmithTune, a CLI to post-train models from agent traces.
“today we're launching LangSmith fine-tuning in public beta with SmithTune, a CLI to allow you to post-train models from your LangSmith traces in one workflow”
Strands Evals and Amazon Bedrock AgentCore add skill-focused evaluators to measure agent skill selection and instruction following.
“A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain-specific task like redacting a contract, reconciling an invoice, or following a team’s pull-request conventions.”
GPT-6 Astra let Parallel's agents research labor-market data in half the time and cost.
Lambert argues true recursive self-improvement won't arrive soon; current AI-safety anxiety reflects scaled agents, not imminent superintelligence.
“they’ll turn out to be directionally correct (relative to the expectations of almost anyone not linked to the community) but factually wrong.”
Files are replacing Python for building agents.
AI verification and validation is crucial in chip development.
“it's like being in the 1990s, you have the internet, but they tell you that you can only use it from two to four o'clock.”
Databricks cut $1M annually in AI agent waste within a single hour of optimization work.
Top AI open source projects are closing external PRs in favor of agent-run software factories
“over 1,000 open issues and almost 800 pull requests”
Lovable CTO argues agents will replace conventional SaaS apps as primary work interface
“you can get to a place where you're using one entry point to all the work that you're doing.”
Models absorbing harness capabilities into weights is reshaping agent architecture toward human-attention scaffolding
“The change last winter, last Christmas — it's a little hard to pin down. I mean, the harness changed and a little post-training changed and then new pre-trained models came… but it felt like a big jump which is not that easy to pin down what did it.”
Agent infrastructure is now commoditized by cloud platforms, making context the new competitive frontier
“They're all taxes one has to pay in order to get an agent out there to play the game.”
Anthropic defines agents by leaning on model intelligence with minimal core primitives in a loop.
“Agents in anthropic are actually very simply defined where essentially you want to give the model and lean in on model intelligence as much as possible.”