Evaluating Agents in Production: Traces, LLM-as-Judge, and Prompt Management at Wonder
Wonder uses LangSmith traces and LLM-as-judge to automate AI agent evaluation at production scale
“The feedback loop literally went from minutes, sometimes even like multiple hours, versus now it's all automated.”
Wonder, a meal planning AI startup, adopted LangSmith for agent tracing and LLM-as-judge evaluation after manual log debugging proved unscalable. The team credits prompt management and MCP integration with Claude Code for dramatically compressing their debugging feedback loop from hours to near-instant. The case study highlights emerging best practices for production AI agent observability but is primarily a vendor testimonial rather than a novel industry signal.