SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius
SWE-rebench evaluates coding agents monthly on fresh, decontaminated real-world software engineering tasks.
“if you want to build some open truly decontaminated benchmark, um time splits are the only way”
Nebius's Ibragim Badertdinov presented SWE-rebench, a leaderboard that evaluates ~30 models monthly on fresh real-world software engineering tasks collected from the prior month. The key insight is that time-based splits are the only reliable way to avoid benchmark contamination, since released benchmark data leaks into the next generation's pre-training. This matters because trustworthy, decontaminated evals are increasingly critical as many strong open and closed coding models compete.