The Hallway Track
Research Findings

We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect

AI Engineer · Sep 26, 2026 · Research Findings

Prime Intellect benchmarks Claude Code and Codex against human researchers on AI research tasks

“we don't have any benchmark to quantify whether that's true or not, right? And even more so, we don't have an independent benchmark from small laboratories to understand whether to expect this in the near future.”

Prime Intellect is building an open benchmark to measure how well AI models like Claude Code and Codex can conduct AI research autonomously, using the nanoGPT speedrun challenge as the evaluation environment. The work addresses a gap called out by big labs warning of recursive self-improvement: there is no independent, public benchmark to quantify whether AI-driven research is actually here. This matters because it could provide an early, open signal for when AI systems cross the threshold into meaningful autonomous scientific contribution.

autonomous-research benchmarking recursive-self-improvement nanoGPT prime-intellect

Watch / read the original source →