Development teams transitioning from simple coding assistants to autonomous agents face a critical reality: there is no single "best" model for generating secure software, and paying more does not guarantee better security. A comprehensive evaluation by Secure Code Warrior and RMIT tested six frontier models across eleven language frameworks using over 600 codebases.
Performance Varies By Framework
The study found that a specific LLM could produce twice as much secure code in one environment while lagging significantly in another. For platform engineers, this means model selection cannot be generic; it must align with the primary stack.
- GPT 5.1 performed strongest on Software Integrity and Security Logging but struggled with Insecure Design.
- Google's Gemini 2.5 Pro excelled at handling Identification Failures and Authentication issues.
- Claude Sonnet 4.5 demonstrated dominance in the Java ecosystem, React, C#, and JavaScript environments.
The Hidden Cost of Agentic AI
While token pricing has dropped significantly following market shifts like DeepSeek's release, agentic workflows introduce a new cost equation that often goes overlooked. Unlike standard chat interactions where costs are predictable per prompt-response pair, agents execute multi-step loops involving tool calls and external application interaction.
The study highlighted massive variance in resource consumption: Claude Sonnet 4.5 generated roughly four times the output tokens of GPT 5.1 for identical tasks (176 million vs 42 million). For operations teams monitoring budgets, this implies that a "cheap" model per token might become prohibitively expensive if an agent's logic triggers excessive tool usage or verbose reasoning loops.
Security Implications
The data confirms that security is not uniform across models. GPT 5 mini scored the lowest overall (10.0), while Gemini 2.5 Flash, despite being below average in most categories, showed unexpected resilience against Server-Side Request Forgery (SSRF). Security architects must therefore treat AI-generated code as untrusted by default and rely on independent Static Application Security Testing (SAST) pipelines rather than trusting the model's inherent safety.
What This Means For Practitioners
The era of assuming a single "safe" LLM exists is over. Engineering teams must adopt a polyglot strategy, selecting models based on their specific framework strengths—such as using Sonnet for Java-heavy stacks or Gemini Pro for Python/Swift projects—and implementing strict token budgets to control the operational costs of autonomous agents.