Evaligo Releases SWE-Race: A Benchmark for Concurrency Bug Fixing in Multi-Agent Systems
By Mr.Xu
Published:
Summary:Evaligo has launched SWE-Race, a benchmark platform designed to evaluate the performance of multi-agent systems in fixing real-world concurrency bugs, including race conditions, deadlocks, and cancellation issues. The benchmark is built on 188 concurrency bugs extracted from over 100 Python projects. Each task is graded by the project's own tests in a network-isolated container to prevent agents from recovering fixes from Git history. Initial results show that GLM-5.3 Flash and GPT-5.6 Luna perf
Overview
SWE-Race, developed by Evaligo, is a benchmark platform designed to evaluate the performance of multi-agent systems in fixing real-world concurrency bugs. Key features of the platform include:
- Real-world Bug Data: The benchmark is built on 188 real concurrency bugs extracted from over 100 Python projects, including race conditions, deadlocks, and cancellation issues.
- Controlled Testing Environment: Each task is run in a network-isolated container with the codebase trimmed to a single commit to prevent agents from recovering fixes from Git history.
- Transparent Scoring Mechanism: Tasks are graded by the project's own tests, ensuring fairness and accuracy in evaluation.
Key Findings
- Model Performance Differences: GLM-5.3 Flash achieved an 85% success rate in single-attempt tasks, while GPT-5.6 Luna scored 81%. With multiple attempts, GLM-5.3 Flash's success rate slightly decreased to 82%, similar to GPT-5.6 Luna.
- Task Difficulty Impact: Approximately half of the tasks were relatively easy for all models, with success rates close to 100%. However, in the harder half, the models' performance varied significantly, with GLM-5.3 Flash, GPT-5.6 Luna, and other models achieving 50%, 45%, and 23% success rates, respectively.
- Network Access Restrictions: Out of 11,000 commands, 69 attempted to access the network but all failed. GLM-5.3 Flash tried 50 times to download the already-fixed library version.
- Comparison of Old and New Bugs: Older bugs (pre-2026) were solved about 9 percentage points more often than newer ones, but the confidence interval crosses zero, so no definitive conclusions can be drawn.
Industry Impact
SWE-Race provides an objective evaluation standard for the capabilities of multi-agent systems in fixing concurrency bugs. The open access to task data and testing protocols offers valuable resources for researchers and developers, driving advancements in AI-driven code repair. Furthermore, the platform sets a clear direction for future model improvements and optimizations, such as enhancing performance in complex tasks and reducing reliance on network resources.
Developer Recommendations
- Leverage Public Data: Developers can utilize the open task data and testing protocols from SWE-Race to evaluate and improve their models.
- Focus on Model Robustness: When designing AI models, special attention should be paid to their performance in complex and diverse tasks to enhance robustness.
- Explore New Methods: Researchers and developers can explore new methods and techniques, such as reinforcement learning and transfer learning, to further improve model performance in fixing concurrency bugs.
Conclusion
The release of SWE-Race marks an important milestone in AI-driven code repair. It not only provides an objective evaluation platform but also points the way for future research and technological improvements.
— END —Source: Reddit r/MachineLearning (2026-10-06)
Tags: #Evaligo #SWE-Race #Concurrency Programming #Multi-Agent Systems #AI Code Repair
Community Comments