Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker
AI now completes programming tasks taking humans weeks, in hours via MirrorCode benchmark
“We also found that AI models are improving rapidly over time. Leading models from a year ago would have scored about 30%, and were limited to simpler programs, such as a calendar utility.”— Jack Clark
Epoch and METR released MirrorCode, a benchmark where AI must reimplement software programs purely through CLI access with no source code or web access. Claude Opus 4.7 solved a task in 14 hours for $251 that METR estimates would take a human 2-17 weeks, with 17 of 25 target programs achieving at least one perfect-scoring run. The deeper implication flagged by Jack Clark is that AI systems can now self-orient within novel environments and bootstrap capabilities from scratch, a potential precursor to more general autonomous learning.