The Hallway Track

observability

71 tracked signals on observability.

Don't sleep on wrapture

Simon Willison · Sep 11, 2026

Wrapture is a new indispensable monkey patching package for Python developers.

“This feels like one of those Swiss Army Knife packages that, once mastered, will provide value against all sorts of problems for years to come.”
Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore

AWS Machine Learning Blog · Sep 22, 2026

Strands Evals and Amazon Bedrock AgentCore add skill-focused evaluators to measure agent skill selection and instruction following.

“A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain-specific task like redacting a contract, reconciling an invoice, or following a team’s pull-request conventions.”
Introducing: LangSmith Tuned Evaluators

LangChain · Aug 18, 2026

LangChain launches tuned evaluators that beat frontier models on agent evals at lower cost

“our post-trained perceived error evaluator outperformed all frontier closed and open models, while still remaining the most cost-effective”
Your Agents Need a Save Button - Hamza Tahir, ZenML

AI Engineer · Jul 18, 2026

AI agents lack persistent state checkpoints, making debugging and replay impossible today.

“all of that is lost and it is only stamped as a read-only trace by the end, which is sitting in another tool far away from where the actual code is.”
The Return of the Data Scientist | Interrupt 26

LangChain · Jun 12, 2026

AI engineering evals are fundamentally data science, and teams keep repeating avoidable eval mistakes.

“I'm saying that using LLM judges kind of blindly without validating your validators is bad.”
Evaluate AI agents systematically with Agent-EvalKit

AWS Machine Learning Blog · Jun 11, 2026

AWS released Agent-EvalKit, an open-source toolkit for systematically evaluating AI agents' full execution paths.

“An agent might deliver a well-structured, actionable response while hallucinating, fabricating facts because its tools returned empty results.”
Any agent, any cloud: Standardized tracing with Foundry+OpenTelemetry | DEM341

Microsoft Developer (Build) · Jun 04, 2026

Foundry Observability uses OpenTelemetry to unify tracing across any agent framework or cloud without rewriting agents.

“Today I'm going to show you how to answer those questions in one place, foundry observability, not by rewriting your agents into one agent framework, but by adopting open telemetry instrumentation with a few lines of code and without changing your existing agent logic.”
From observability to ROI for AI agents on any framework | BRK252

Microsoft Developer (Build) · Jun 03, 2026

Microsoft Foundry observability adds end-to-end tracing, evaluation, monitoring, and ROI for AI agents on any framework.

“agents are non-deterministic, creating new reliability and consistency challenges for developers and operators”
Dashboards Are Dead — Sarah Simionescu, Composio

AI Engineer · Oct 04, 2026

AI agents and MCP are making traditional dashboards and query languages obsolete

“Every dashboard you've ever used, every weird little-known query language ever written, was a means of translation between you and your data, because the machines on the other end couldn't understand what you really wanted.”
What Is LangChain, Actually?

LangChain · Sep 23, 2026

LangChain has evolved from an open source framework into a full agent development lifecycle platform anchored by LangSmith.

“Turns out building the agent is really fun. Super fun. And kind of easy now, but actually keeping it from going completely sideways in production is the hard part.”
The Agentforce Keynote in 2 Minutes #DF26

Salesforce · Sep 18, 2026

Salesforce touts Agentforce ROI playbook and launches Agent Optimizer with now-free, unmetered observability.

“But my favorite part of observability is that it's now unmetered. That means it's free, y'all.”
What does it really take to ship an AI agent?

Microsoft Developer (Build) · Aug 28, 2026

Microsoft Foundry offers end-to-end tooling for deploying, governing, and monitoring AI agents in production.

“The biggest challenge when building your AI agent is how do you get it in production?”
How Podium Traces Every AI Agent Decision

LangChain · Aug 28, 2026

Podium uses LangSmith to trace agent chain-of-thought and debug unexpected AI behavior

“The reality is when you really get into the details of what context that agent was provided, It becomes obvious the agent was behaving rationally.”
Agentic observability with Amazon OpenSearch Service MCP Apps

AWS Machine Learning Blog · Aug 25, 2026

Amazon OpenSearch Service MCP Apps embed interactive observability visualizations inside AI agent chat threads

“You're not trusting the AI's interpretation. You're seeing the actual query result rendered as an interactive chart, trace waterfall, or service map.”
Voice Agent observability with LangSmith 🌟

Google Developers (Google I/O) · Aug 05, 2026

LangSmith now provides full observability for voice agents built on Gemini Live

“To take that agent to production safely, you need visibility into what your agent is doing.”
Frontier results, on device - RL Nabors, Arize

AI Engineer · Jun 29, 2026

Local on-device models can replace frontier models like GPT-5 and Claude to cut inference costs, latency, and security risks.

“Every time you reach for foundation models like GPT-5 or Claude, it's costing you, your users, and the environment.”
AI Agent Failure Detection and Root Cause Analysis with Strands Evals

AWS Machine Learning Blog · Jun 15, 2026

AWS's Strands Evals SDK adds detectors that automatically diagnose AI agent failures and recommend fixes.

“Detectors answer “why did it fail?” by producing diagnoses at the per-span level with categorized failures, causal chains, and fix recommendations.”
Monitor GenAI applications beyond golden signals | ODSP907

Microsoft Developer (Build) · Jun 03, 2026

GenAI applications require extending traditional golden signals because they are non-deterministic, variably costed, newly attackable, and subjectively judged.

“What makes Genai applications fundamentally different from the traditional software we've been monitoring for decades? There are four key shifts.”
Evaluating Deep Agents using LangSmith on AWS

AWS Machine Learning Blog · May 28, 2026

LangSmith on AWS provides five evaluation patterns for deep agents in production

“Validating AI agent behavior before production is one of the hardest problems in applied AI.”
5 tips to creating production-ready AI agents

Google Developers (Google I/O) · May 28, 2026

Google recommends five architectural patterns for production-ready AI agents at scale

“Production agents need robust architecture, not just clever prompts.”
The 5 Levels of Self-Driving Production — Eric Schwartz, Traversal

AI Engineer · Oct 02, 2026

AI coding agents accelerate development but are shifting engineer time toward troubleshooting, not design.

“more and more time is spent on troubleshooting. Much more code is being written. People are less understanding of the code that goes into production.”
Voice Agent observability with LangSmith

Google Developers (Google I/O) · Jul 31, 2026

Google ADK voice agents using Gemini Live can be traced with LangSmith for production observability

“To take that agent to production safely, you need visibility into what your agent is doing and the ability to test its behavior in a number of different scenarios that it might encounter with real end users.”
Trace Every Cursor Agent Turn in LangSmith

LangChain · Jul 21, 2026

LangSmith now traces every Cursor agent turn with full tool call and sub-agent visibility

“If you watch the Claude Code or Codex versions of the series, you'll notice that the traces actually share a common structure under the hood, so you can compare all three agents in the same LangSmith workspace.”
Build an agentic incident triage assistant with Amazon Quick and New Relic

AWS Machine Learning Blog · Jun 09, 2026

AWS shows how to build an agentic incident triage assistant using Amazon Quick, New Relic MCP Server, and Asana.

“From a single prompt, the Amazon Quick agent investigates the incident, assembles a root cause analysis (RCA) brief with evidence links, and creates a tracked Asana task ready for handoff.”
LLM Observability, Evaluation, Experimentation Platform — Dat Ngo, Arize

AI Engineer · Jun 07, 2026

Building reliable AI systems requires observability, evaluation, and experimentation, not magic.

“It's really the same set of patterns, just maybe a different flavor coming out. And it's really it feels like magic, but it's not magic, right? It's all just engineering.”
Behind the Scenes: Accelerating the AI Agent DevOps Lifecycle with End-to-End | LIVE159

Microsoft Developer (Build) · Jun 05, 2026

Microsoft Foundry's agent platform unifies tracing, evaluation, and optimization to test non-deterministic AI agents end-to-end.

“The properties that makes agents so useful like they're stateful, they they're long horizon, they can plan, they can course correct, they interact with their environment through tools are also the ones that makes it very hard to test.”