We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect
Prime Intellect benchmarks Claude Code and Codex against human researchers on AI research tasks
“we don't have any benchmark to quantify whether that's true or not, right? And even more so, we don't have an independent benchmark from small laboratories to understand whether to expect this in the near future.”
Prime Intellect is building an open benchmark to measure how well AI models like Claude Code and Codex can conduct AI research autonomously, using the nanoGPT speedrun challenge as the evaluation environment. The work addresses a gap called out by big labs warning of recursive self-improvement: there is no independent, public benchmark to quantify whether AI-driven research is actually here. This matters because it could provide an early, open signal for when AI systems cross the threshold into meaningful autonomous scientific contribution.