Traditional single-number benchmarks fail to capture that modern AI capability scales with test-time compute budget.
“The capability of the model is a function of how much money you put into it.”
8 tracked signals on reasoning-models.
Traditional single-number benchmarks fail to capture that modern AI capability scales with test-time compute budget.
“The capability of the model is a function of how much money you put into it.”
Microsoft released MAI-Thinking-1, a reasoning model trained without third-party distillation, with a 109-page transparency-heavy report.
“hillclimbed from scratch”
Microsoft AI shipped seven first-party models including MAI Thinking One, its first frontier reasoning model.
“We call our approach humanist superintelligence.”
An OpenAI reasoning model helped researchers identify 18 new diagnoses in previously unsolved rare genetic disease cases.
NVIDIA's Nemotron 3 Ultra delivers faster, more token-efficient reasoning for long-running, multi-turn AI agents.
“Single-turn chatbots are evolving into long-running agents that can reason, maintain context, use tools, and run efficiently across many turns to complete complex workflows.”
Qwen 3.8 27B defaults to extreme overthinking, consuming 22K tokens for simple tasks
“This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware.”
llm-mistral 0.16 adds reasoning model support including Mistral Large 4
llm 0.33 adds composable templates, per-call embedding keys, and reasoning summary options