Navan architects argue AI agents are at the same inflection point as microservices in 2015
“If you can't build a single agentic loop, why go in and try to build a multi-agent orchestrated system?”
11 tracked signals on production-ai.
Navan architects argue AI agents are at the same inflection point as microservices in 2015
“If you can't build a single agentic loop, why go in and try to build a multi-agent orchestrated system?”
LangChain CEO outlines a five-stage agent lifecycle: build, test, deploy, monitor, govern
“getting something that works, you know, initially locally or in a one-off on a Twitter demo is easy, but like shipping it reliably to 100 of users at scale, the difficulty becomes down down to like the performance and the behavior of the agent”
Most production ML security breaches stem from basic infrastructure mistakes, not exotic AI attacks
“almost everything that is breaking in the production ML security isn't some exotic AI attack. It's the same boring infrastructure mistakes that we supposedly fixed years ago.”
Best production agents use deliberate human checkpoints rather than pursuing full autonomy from day one
“The best teams, sort of like shipping the best agents, just aren't really going for a full autonomy on day one.”
Traditional retry and circuit breaker patterns fail for LLMs; per-request provider fallback is required.
“If you have a single model provider, their ceiling is your ceiling. Their outage is your outage.”
AI agents evolve rapidly but evaluation frameworks fail to keep pace with model changes
“building a demo with AI is really easy but making it production quality is really hard”
Production AI agent loops need discipline to avoid unreviable 40,000-line PRs on real teams
“we're building 40,000line PRs that just nobody wants to read”
Uber built a multimodal AI agent to enhance merchant food photos at scale without looking AI-generated
“we're threading the needle here. We need to be able to stay faithful to the original image, preserve the brand of the merchant, and avoid everything looking the same.”
Google recommends five architectural patterns for production-ready AI agents at scale
“Production agents need robust architecture, not just clever prompts.”
YouTube Ads team built agent reliability via layered evals, critique agents, and strong tool foundations.
“the reliability of your agent is basically a function of the capabilities of the agent uh the guard rails and the evals”
Agent harnesses will evolve into autonomous 'claws' — the next phase beyond current frameworks
“welcome to the harness era”