The Hallway Track

benchmarks

44 tracked signals on benchmarks.

[AINews] not much happened today

Latent Space Blog · Oct 03, 2026

Anthropic models now hold the top three Agent Arena spots as Sonnet 5.5 debuts at #3

“good, cheap AND fast”
[AINews] Reve 2 and Ideogram 4: Layouts in Imagegen

Mustafa Suleyman · Latent Space Blog · Jun 04, 2026

Microsoft released MAI-Thinking-1, a reasoning model trained without third-party distillation, with a 109-page transparency-heavy report.

“hillclimbed from scratch”
How SWE-Serve Exposes the Gap Between Local Tests and Live Serving

NVIDIA Developer Blog · Sep 23, 2026

NVIDIA's SWE-Serve benchmark tests whether AI coding agents' patches work in live model serving, not just local tests.

“An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests.”
The Self-Driving Eval Trick No AI Benchmark Beats

LangChain · Sep 21, 2026

Manually reviewing 100-1,000 real examples beats any automated benchmark for evaluating AI models.

“there's just no better eval than looking at 100 examples or 1,000 examples”
How Harvey Built a Research Lab on a Budget | Gabe Pereyra

Sequoia Capital · Aug 11, 2026

Harvey built a domain-specific research lab by leveraging frontier ecosystem rather than competing with it directly.

“it's an unfair game competing with the frontier labs if you're an application layer company.”
The Billion Dollar AI Race Just Broke

Two Minute Papers · Aug 05, 2026

Qwen 3.8 Max challenges OpenAI and Anthropic at 5-10x lower API pricing

“it sat there thinking for 16 days, starting from an empty folder, writing, testing, and repairing its own code”
When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI

Andrej Karpathy · AI Engineer · Aug 02, 2026

AI labs optimize for benchmark scores rather than real-world model quality, warns industry insider

“unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis”
[AINews] not much happened today

Latent Space Blog · Aug 01, 2026

DeepSeek V4-Flash 0731 matches GPT-5.6 performance at 60% lower cost via post-training alone

“Terminal-Bench 82.7, up +25.8 from the April preview's 56.9”
AI benchmark scores don’t tell you what you think they do

No Priors · Jun 29, 2026

Current AI safety policies fail to account for test-time compute, where model capability scales with money spent.

“The capability of the model is a function of how much money you put into it, basically.”
NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark

NVIDIA Developer Blog · Jun 12, 2026

NVIDIA leads agentic coding performance on AA-AgentPerf, the industry's first multi-vendor agentic AI benchmark.

“Artificial Analysis AgentPerf (AA-AgentPerf) offers the industry's first multi-vendor open benchmarks profiling trajectories that are representative of real-world AI agent coding tasks.”
Mistral Large 4

Simon Willison · Oct 06, 2026

Mistral Large 4 released, benchmarks criticized as saturated by frontier models

“The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.”
Are AI labs pelicanmaxxing?

Simon Willison · Jul 22, 2026

Systematic study finds no evidence AI labs train models to draw pelicans on bicycles better

“Pelicans aren't drawn any better than other animals. Bicycles aren't drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict.”