RepairBench: Leaderboard of Frontier Models for Program Repair
André Silva, Martin Monperrus · 2025
AI-driven program repair uses AI models to repair buggy software by producing patches. Rapid advancements in frontier models surely impact performance on the program repair task. Yet, there is a lack of frequent and standardized evaluations to actually understand the strengths and weaknesses of models. To that end, we propose RepairBench, a novel leaderboard for AI-driven program repair. The key characteristics of RepairBench are: 1) it is execution-based: all patches are compiled and executed against a test suite, 2) it assesses frontier models in a frequent and standardized way. RepairBench leverages two high-quality benchmarks, Defects4J and GitBug-Java, to evaluate frontier models only against real-world program repair tasks. At the time of writing, RepairBench shows that claude-3-5-sonnet-20241022 is the best model for program repair, and deepseek-v3 one of the cheapest while ranking third. We publicly release the evaluation framework of RepairBench as well as all patches generated in the course of the evaluation.