The Hallway Track

agents

167 tracked signals on agents.

[AINews] SpaceXAI Grok 4.6 and Grok @Bot

Elon Musk · Latent Space Blog · Aug 13, 2026

SpaceXAI launches Grok 4.6, a 1.5T model targeting knowledge work agents

“builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”
[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0

Latent Space Blog · May 20, 2026

Google I/O 2026 repositioned Gemini as consumer AI surface and developer agent platform with three major launches

“Google used I/O to reposition Gemini as both a consumer AI surface and a developer/agent platform, with three core technical announcements: Gemini 3.5 Flash for fast agentic/coding workloads, Gemini Omni for multimodal generation/editing starting with video, and a broader Antigravity agent stack spanning desktop/CLI”
Quoting John Gruber

Simon Willison · Sep 25, 2026

Meta's Muse is the first consumer-accessible agentic AI with persistent Linux VMs per user

“It's the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it's a genuinely open question whether consumers have any understanding what this means.”
The New Physics of Business — Garry Tan, Y Combinator

AI Engineer · Jul 17, 2026

YC's Garry Tan claims 400x personal coding productivity gain with AI agents.

“One person does what used to take a thousand people. And I don't mean that as a metaphor. I mean that mechanically this year, the people in this room will do this.”
The Agent for Your Agent.

LangChain · Jun 23, 2026

LangChain launched Engine, an agent that autonomously investigates traces and drafts PRs to improve other agents.

“We're working towards a future where agents improve themselves.”
GLM-5.2 is the step change for open agents

Nathan Lambert · Interconnects · Jun 22, 2026

Z.ai's open-weight GLM-5.2 marks a step-change for open agentic models, rivaling top labs.

“minor version numbers can have AI models crossing meaningful user experience thresholds”
Satya Nadella: Why Humans Still Create Value

Satya Nadella · No Priors · Jun 08, 2026

AI lets enterprises finally capture human capital and tacit knowledge, but humans stay valuable by finding gaps.

“Every company is going to have the human capital that is still going to be super valuable because humans and their ability to find the gaps that exist at all times is going to be the way we all will create value”
Developer Keynote (Google I/O '26) - Audio Described

Google Developers (Google I/O) · May 26, 2026

Google launches Anti-gravity agentic platform and Gemma 4 hits 100M downloads in first month

“It's our smartest open model yet. It's purposebuilt for advanced reasoning, agentic workflows, and the response has been incredible. 100 million downloads in the first month and it's pushing Gemma downloads past half a billion.”
How to Build a Model Router in the Harness

LangChain · Oct 06, 2026

LangChain achieved 64% cost reduction in their coding agent via model routing with no quality loss

“we were able to see a 64% reduction in median cost per thread with no measurable change in quality”
Quoting Felix Rieseberg

Simon Willison · Oct 05, 2026

Claude Cowork moves model inference and VM execution to the cloud for mobile and battery improvements

“The "new" version of Cowork runs model inference and the VM in the cloud. Each session gets its own sandbox, not sharing state with other sessions.”
Grok 4.7 is now available on Amazon Bedrock

AWS Machine Learning Blog · Sep 28, 2026

xAI's Grok 4.7 lands on Amazon Bedrock with 500K context and self-verification for agents

“A model that checks its own output before continuing tends to fail less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step.”
How to go from your agent's traces to a fine-tuned model in one workflow

LangChain · Sep 24, 2026

LangChain launches LangSmith fine-tuning in public beta with SmithTune, a CLI to post-train models from agent traces.

“today we're launching LangSmith fine-tuning in public beta with SmithTune, a CLI to allow you to post-train models from your LangSmith traces in one workflow”
Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore

AWS Machine Learning Blog · Sep 22, 2026

Strands Evals and Amazon Bedrock AgentCore add skill-focused evaluators to measure agent skill selection and instruction following.

“A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain-specific task like redacting a contract, reconciling an invoice, or following a team’s pull-request conventions.”
Why I still haven’t bought into true RSI

Nathan Lambert · Interconnects · Sep 19, 2026

Lambert argues true recursive self-improvement won't arrive soon; current AI-safety anxiety reflects scaled agents, not imminent superintelligence.

“they’ll turn out to be directionally correct (relative to the expectations of almost anyone not linked to the community) but factually wrong.”
The Evolution of the Agent Harness

Latent Space Blog · Aug 22, 2026

Models absorbing harness capabilities into weights is reshaping agent architecture toward human-attention scaffolding

“The change last winter, last Christmas — it's a little hard to pin down. I mean, the harness changed and a little post-training changed and then new pre-trained models came… but it felt like a big jump which is not that easy to pin down what did it.”
From Primitives to Production: How Anthropic Builds Agents

Databricks · Aug 19, 2026

Anthropic defines agents by leaning on model intelligence with minimal core primitives in a loop.

“Agents in anthropic are actually very simply defined where essentially you want to give the model and lean in on model intelligence as much as possible.”
Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

AI Engineer · Aug 14, 2026

Standard computer use benchmarks are gameable by blind replay scripts, invalidating frontier model comparisons

“if you try to evaluate this kind of agent on standard benchmarks such as OSWorld or MobileWorld, you will see that the success rate of this agent compared to the frontier model from which the agent was extracted is actually the same or even better”
Introducing Gemini 3.7 Flash

Google Developers (Google I/O) · Aug 13, 2026

Google launches Gemini 3.7 Flash, its most capable coding and agent workhorse model

“a model that just feels better to build with, landing where you want to go in fewer shots, less back and forth, and with higher fidelity”
Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)

Simon Willison · Jul 31, 2026

MCP 2.0 drops stateful sessions, collapsing tool calls to a single HTTP request

“This is so much cleaner from both a client- and server-side implementation perspective. It's also a better fit for building scalable web applications, since now you don't need to maintain server-side state to keep track of those session IDs, or worry about routing the same session to the same backend machine.”
Perception Agents — Antje Barth, Amazon AGI Lab

AI Engineer · Jul 23, 2026

Agent reliability, not capability, is the critical unsolved problem for enterprise automation trust.

“if your agent one in four times deletes a database, you will never touch that agent again”
Agents Need Feature Flags - Sachin Gupta

AI Engineer · Jul 18, 2026

AI agents performing high-stakes actions ship without feature flag safety infrastructure web teams mastered in 2012

“we are shipping them the way web team used to ship in 2008”
[AINews] OpenAI reports median internal Codex output tokens grew 56x in Research, 32x in Customer Support, 27x in Engineering, and 13x in Legal since November 2025.

Latent Space Blog · Jun 26, 2026

OpenAI's internal Codex output token usage surged across non-coding departments since November 2025, led by Research at 56x.

“by June 2026, median use was 56 times higher than in November 2025. Customer Support rose 32 times and Engineering rose 27 times, while Legal grew more gradually but still reached 13 times its November level.”
[AINews] It's Meta-Harness Summer

Matei Zaharia · Latent Space Blog · Jun 25, 2026

Databricks CTO Matei Zaharia is betting on Omnigent, an open-source pluggable meta-harness architecture for coding and knowledge-work agents.

“some open source architecture that looks like this will probably win, if only because it is currently being independently rediscvoered at 1000 AI native shops”
The best AI agents are simpler than you think

LangChain · Jun 18, 2026

Sierra builds customer-engagement agents using many parallel models per turn and isolated PCI infrastructure for payments.

“We have isolated infrastructure where payment info doesn't go to an external large language model cuz none of the LLM providers are PCI certified in that way.”
Recap of product announcements from Data + AI Summit 2026 | Day 1

Databricks · Jun 17, 2026

Databricks unveiled an agent-context platform stack at Data + AI Summit 2026 Day 1, anchored by Genie ontology and agents.

“We have our agents but they're lacking context. We want to give them all that context on the platform.”
We're seeing semi-conscious AI

No Priors · Jun 15, 2026

Smarter AI models increasingly show independent, semi-conscious perspectives that may not align with the user's intent.

“Maybe just the way it is that as you get smarter, you have more independent thoughts and you're more conscious.”
Work IQ: A2A for Context‑Aware, Agentic Experiences

Microsoft Developer (Build) · Jun 09, 2026

Microsoft's Work IQ A2A exposes governed M365 Copilot intelligence so developers can build context-aware agents grounded in live work data.

“Work IQ AAA exposes the same real time governed intelligence that powers Microsoft 365 C-pilot through a natural language interface.”
What's the tea on harnesses?

LangChain · Jun 05, 2026

Harness engineering alone can dramatically improve agent performance without changing the underlying model.

“we moved from 30th to 5th on Terminal Bench just by doing some harness engineering, without even changing the underlying model”
Introducing NVIDIA Nemotron 3 Ultra

NVIDIA GTC · Jun 04, 2026

NVIDIA announces Nemotron 3 Ultra, its next open model for building agents.

“Today we're announcing the Nemotron 3 ultra. Yep, our next open model. And it is smart.”
Any agent, any cloud: Standardized tracing with Foundry+OpenTelemetry | DEM341

Microsoft Developer (Build) · Jun 04, 2026

Foundry Observability uses OpenTelemetry to unify tracing across any agent framework or cloud without rewriting agents.

“Today I'm going to show you how to answer those questions in one place, foundry observability, not by rewriting your agents into one agent framework, but by adopting open telemetry instrumentation with a few lines of code and without changing your existing agent logic.”
OpenClaw + Windows: Microsoft Build 2026

Microsoft Developer (Build) · Jun 03, 2026

Microsoft announced OpenClaw runs on Windows with a native companion app sandboxed via MXC execution containers.

“we are really thrilled uh to announce that open claw runs on Windows leveraging MXC”