The Hallway Track
Engineering Insights

Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase

AI Engineer · Oct 05, 2026 · Engineering Insights

LLM judges used as agent benchmark verifiers are unreliable and overconfident evaluators

“The fact is that the web is a very open system. There is no single path to correctness. There is no absolute truth.”

BrowserBase and Microsoft researchers published a study finding that LLM-based judges—used as verifiers in leading agent benchmarks like OSWorld, WebVoyager, and Online Minor Web—fail to provide reliable evaluation signals. The team discovered this after their own deterministic evaluation environments became unscalable as agent capabilities grew and the open, dynamic nature of the web broke deterministic verification. The finding challenges a widespread industry assumption that LLM judges are a trustworthy substitute for deterministic or human verification in agentic RL training loops.

llm-evaluation agent-benchmarks web-automation rl-environments verifiers

Watch / read the original source →