The Hallway Track
Engineering Insights

SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius

AI Engineer · Jun 04, 2026 · Engineering Insights

SWE-rebench evaluates coding agents monthly on fresh, decontaminated real-world software engineering tasks.

“if you want to build some open truly decontaminated benchmark, um time splits are the only way”

Nebius's Ibragim Badertdinov presented SWE-rebench, a leaderboard that evaluates ~30 models monthly on fresh real-world software engineering tasks collected from the prior month. The key insight is that time-based splits are the only reliable way to avoid benchmark contamination, since released benchmark data leaks into the next generation's pre-training. This matters because trustworthy, decontaminated evals are increasingly critical as many strong open and closed coding models compete.

coding-agents evaluation benchmark-contamination swe-rebench

Watch / read the original source →