← Back to brief
ResearchOfficialPreprintarXiv Software Engineering

SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution

Researchers introduce SWE-Milestone, a benchmark designed to evaluate AI agents on streams of milestone-level coding tasks that simulate real-world software evolution. By testing 12 advanced models across 4 agent frameworks, they observe that performance drops from over 80% on isolated tasks to 38.03% in continuous, evolving scenarios. This highlights significant challenges for AI agents in maintaining system integrity and managing error propagation over time.

Why it matters: This benchmark reveals a major limitation in current AI coding agents: their inability to reliably handle long-term software maintenance, which is essential for real-world applications.

Full story at: arXiv Software Engineering