What does a passing test establish?

SWE-Bench evaluates patches against tests drawn from real software issues. Those tests can miss relevant cases, allowing an incorrect repair to receive a passing result.

Testing the evaluation

UTBoost builds on UTGenerator, which analyzes a Python project and its dependencies to create additional tests. It compares the behavior of a generated patch with the reference patch and improves the parser that reads test results.

What the paper reports

The evaluation reported in the ACL 2025 paper identified 345 erroneous patches previously marked as passing across SWE-Bench Lite and Verified.

Read the paper for the evaluation design and limitations. The released toolkit includes test-generation code, augmented tests, and instructions for re-evaluation.