[AINews] FrontierCode: Benchmarking for Code Quality over Slop
Cognition launches FrontierCode, a benchmark grading code quality and maintainability over passing-test 'slop'.
“Many SWE-bench-Passing PRs Would Not Be Merged into Main”
Cognition released FrontierCode, a benchmark inspired by FrontierMath that evaluates frontier models on code quality and maintainability rather than just passing tests, directly responding to findings that many SWE-bench-passing PRs would be rejected from real codebases. It matters because it targets the false-positive/'slop' problem undermining existing coding benchmarks and reflects the late-2025 jump in agentic engineering capability.