The Hallway Track
Engineering Insights

The Return of the Data Scientist | Interrupt 26

LangChain · Jun 12, 2026 · Engineering Insights

AI engineering evals are fundamentally data science, and teams keep repeating avoidable eval mistakes.

“I'm saying that using LLM judges kind of blindly without validating your validators is bad.”

Hamel Husain and Shreya Shankar argue that the 'harness' and eval work driving agentic AI is really data science, and they catalog common mistakes teams make—using generic off-the-shelf metrics, trusting unvalidated LLM judges, and poor experimental design. The fix is to apply rigorous data-science practices: look at real traces to name bespoke failure modes, treat LLM judges as imperfect classifiers validated with train/test splits and precision/recall, and design metrics aligned to the specific application. It matters because reliable evals are the bottleneck for shipping trustworthy AI agents at scale.

evals llm-as-judge data-science ai-engineering observability

Watch / read the original source →