Claude Fable 5.1 sets a new standard with impressive benchmark scores.
“Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks".”
14 tracked signals on benchmark.
Claude Fable 5.1 sets a new standard with impressive benchmark scores.
“Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks".”
NVIDIA's AVO agent architecture achieves 100% on ARC-AGI-3 benchmark.
“A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks.”
Z.ai's GLM-5.3 surpasses Claude Fable 5 and GPT-5.6-Sol on select benchmarks
“On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.”
Moonshot AI's Kimi K3 is a 2.8T-parameter open-weight model matching frontier closed models on coding benchmarks.
“it has OpenAI and Anthropic terrified because its Trust Me Bro benchmark performance is on par with and in some cases beating Claude Fable and GPT 5.6 Soul”
Two API settings tripled OpenAI's GPT-5.6 scores on the ARC-AGI-3 benchmark
Moonshot AI's Kimi K3 2.8T MoE beats Opus 4.8, claiming best open-weights model title
“This is more than a model drop; it is a fairly complete recipe for large-scale agentic post-training and serving.”
OpenAI introduces MentalHealthBench, an expert-informed benchmark for evaluating safe AI responses in mental health conversations.
“MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.”
GLM 5.3 Flash delivers near-frontier AI performance for free via open weights.
“I think within this year, in a few months, Fable might be surpassed by free AI systems.”
DiG-bench shows Fable 5 displays creative intuition; current frontier models cannot beat discovery games
“The new frontier for analyzing AI systems is understanding how good they are at inferring the unwritten rules of their environment”
Gemini 3.7 Flash closes competitive gap against Claude 4.8+ and GPT 5.5+ series
Agent harness architecture drives double-digit benchmark swings independent of model choice
“Harness design alone can account for double-digit swings in benchmark results and significant differences in token cost”
Hugging Face launches the FFASR Leaderboard to benchmark automatic speech recognition on real-world audio.
OpenAI launches LifeSciBench, an expert-authored benchmark for evaluating AI on real-world life science research tasks.
“Introducing LifeSciBench, an expert-authored, expert-reviewed benchmark for evaluating how AI systems handle real-world life science research tasks and decisions.”
Databricks releases OfficeQA Pro V2 benchmark for enterprise grounded-reasoning evaluation