The Hallway Track

llm-as-judge

9 tracked signals on llm-as-judge.

The Return of the Data Scientist | Interrupt 26

LangChain · Jun 12, 2026

AI engineering evals are fundamentally data science, and teams keep repeating avoidable eval mistakes.

“I'm saying that using LLM judges kind of blindly without validating your validators is bad.”
Evaluate AI agents systematically with Agent-EvalKit

AWS Machine Learning Blog · Jun 11, 2026

AWS released Agent-EvalKit, an open-source toolkit for systematically evaluating AI agents' full execution paths.

“An agent might deliver a well-structured, actionable response while hallucinating, fabricating facts because its tools returned empty results.”
WWDC26: Improve your prompts by hill-climbing with Evaluations | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple's new Evaluations framework lets developers hill-climb and align model judges to reduce drift in AI features.

“This discrepancy between model and human is known as drift, and it is a problem faced by all developers trying to evaluate intelligent features.”
Evaluate your Amazon Nova Sonic voice agent at scale, no microphone required

AWS Machine Learning Blog · Jun 08, 2026

AWS released the open source Nova Sonic Test Harness to automatically evaluate voice agents at scale without a microphone.

“It runs complete multi-turn conversations with Amazon Nova Sonic automatically, evaluates them using LLM-as-judge techniques, and can even detect cases where the model's audio output doesn't match its text output (audio hallucinations).”
Building a Harness with Jev

LangChain · Sep 21, 2026

LangChain demos 'Jev,' a fast, cheap System 1 classification model for agent routing, guardrails, and evals.

“System 1 models are a class of AI models built to make fast structured decisions that software can use directly.”