Recent evaluations indicate that AI coding agents struggle significantly with large-scale refactoring, particularly when dealing with cross-file dependencies across diverse programming environments. A new benchmark developed by researchers from Shanghai Jiao Tong University and other institutions highlights this limitation, showing frontier models achieving only a 41.2% resolve rate on tasks designed to test genuine structural understanding of codebases.
What Changed in Evaluation Methodology
The industry has long relied on benchmarks that often fail to capture the complexity required for production-grade software engineering. The new SWE-Bench ProMax addresses this by focusing specifically on large-scale refactoring tasks derived from real-world commits across seven programming languages, including Python, Java, TypeScript, Go, C, C++, and Rust.
Unlike previous evaluations that may have been compromised by leaked solutions or overly narrow test suites, this benchmark underwent rigorous multi-stage curation. Researchers manually reviewed every instance to ensure precise issue descriptions and removed tests with insufficient complexity or limited scope. This approach ensures the results reflect genuine agent capabilities rather than memorized answers.
Engineering Implications for AI Adoption
The primary takeaway is that token proximity does not guarantee structural understanding of codebases. Engineers who treat LLMs primarily as text-processing engines risk deploying agents that fail when faced with complex, convoluted systems where changes in one file impact others.
- Context Window Limitations: Even highly optimized attention algorithms struggle to maintain focus across massive contexts, leading to missed observations and confusion during large-scale modifications.
- Race Conditions and Time Factors: Agents often fail when parallel tasks execute in different orders or when users interact with systems rapidly. This highlights a fundamental gap between static code analysis and dynamic runtime behavior involving time-dependent logic.
Vojtěch Pavlík from SUSE notes that current models lack the ability to read an entire large codebase at once, meaning they cannot truly "understand" it in its entirety without significant human oversight or architectural adjustments.
Architecture and Operational Considerations
For platform teams evaluating AI integration strategies, these findings suggest that relying solely on automated agents for refactoring tasks is currently risky. The benchmark results imply a need to maintain robust manual review processes alongside agent-assisted workflows when dealing with critical infrastructure or complex legacy systems.
The distinction between text generation and structural understanding becomes crucial here. If an LLM treats code as prose, it may generate syntactically correct diffs that break existing functionality because they fail to account for the deterministic graph of information represented by a large system's architecture.
What This Means For Practitioners
The emergence of SWE-Bench ProMax signals a shift toward more complex, less learnable tests designed to expose where agents truly struggle. As benchmarks become ubiquitous and models memorize common tasks, practitioners must look beyond headline scores that may be misleading due to flawed test suites or data leakage.
The gap between current capabilities and real-world requirements is widening as agents are pushed to solve harder problems without adequate training data or architectural support. Practitioners should expect continued limitations in automated refactoring until models evolve beyond simple text generation into true structural reasoning engines.

