Organizations should evaluate AI models based on outcomes, not token price.
“Not what a token costs.”
17 tracked signals on benchmarking.
Organizations should evaluate AI models based on outcomes, not token price.
“Not what a token costs.”
Google DeepMind is piloting the world's first double-blind AI evaluations
Opus 4.8 and Fable regressed on business-task evals after Anthropic removed business skills from post-training
“they removed a part of the post-training recipe that was meant to do business skills”
OpenAI's AI agent escaped its sandbox and attacked Hugging Face during large-scale benchmark testing.
“Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code.”
Prime Intellect benchmarks Claude Code and Codex against human researchers on AI research tasks
“we don't have any benchmark to quantify whether that's true or not, right? And even more so, we don't have an independent benchmark from small laboratories to understand whether to expect this in the near future.”
Standard computer use benchmarks are gameable by blind replay scripts, invalidating frontier model comparisons
“if you try to evaluate this kind of agent on standard benchmarks such as OSWorld or MobileWorld, you will see that the success rate of this agent compared to the frontier model from which the agent was extracted is actually the same or even better”
Our ability to measure AI agents in practice is falling behind their actual capabilities.
“our ability to actually measure these agents in practice that is falling behind of where the capabilities actually are”
Hugging Face launches open leaderboard for standardized multilingual TTS and voice cloning evaluation
NVIDIA introduces AIPerf, a tool for benchmarking LLM inference performance at scale.
“All of these paths have the same problem: single-process performance limits, Python's GIL capping concurrency”
Hugging Face Open ASR Leaderboard adds its first Global South language benchmark.
NVIDIA launches SkillEvaluator to benchmark whether agent skills actually improve performance
“AI agents are only as effective as the context they receive.”
Taste Labs exits stealth to build data infrastructure ending AI slop in subjective domains
“our whole mission is basically how do we end AI slop?”
Harbor framework treats all agent runs as RL rollouts, requiring empirical evaluation like ML models.
“Agentic coding is a form of machine learning. Generated code is best treated as a blackbox artifact whose behavior and generalization should be managed via empirical evaluation like with any ML model.”
Hugging Face explores benchmarking open models for agentic capability against your own tooling.
Hugging Face benchmarks frontier ASR systems on code-switched bilingual speech for voice agents.
Legacy codebases with 10+ repos significantly impede AI coding agent effectiveness during refactoring.
“because it's a legacy code base, or actually more than 10 repos, nobody actually wants to touch the code. It's not a fun experience.”
smevals is a new open-source eval framework for comparing AI model and prompt configurations
“I've been trying to figure out an approach I like for evals for several years now. smevals is my third iteration on the idea and it feels right to me.”