Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase
LLM judges used as agent benchmark verifiers are unreliable and overconfident evaluators
“The fact is that the web is a very open system. There is no single path to correctness. There is no absolute truth.”
BrowserBase and Microsoft researchers published a study finding that LLM-based judges—used as verifiers in leading agent benchmarks like OSWorld, WebVoyager, and Online Minor Web—fail to provide reliable evaluation signals. The team discovered this after their own deterministic evaluation environments became unscalable as agent capabilities grew and the open, dynamic nature of the web broke deterministic verification. The finding challenges a widespread industry assumption that LLM judges are a trustworthy substitute for deterministic or human verification in agentic RL training loops.