Veracode Finds AI-Generated Code Security Has Not Improved Little Since Last Year

Veracode’s 2026 GenAI Code Security Report finds that the security of AI-generated code stagnates at a pass rate of 56 percent — with code-specific models less secure than general-purpose ones.
- AI models built into software code are less secure than general-purpose models
- Despite huge gains in AI speed, power, and reasoning, the report reveals nearly half of AI-generated code still fails security tests.
Veracode, a global leader in application risk management, today released its 2026 GenAI Code Security report which reveals that, despite rapid advances in AI coding capabilities, security has stagnated. Across the four test briefs and more than 100 models tracked since the program began, the average pass rate remains at 56 percent — virtually unchanged from last year’s report. Each model was evaluated in code generation tasks involving multiple programming languages and vulnerability categories, under standardized conditions without specific security information. AI now generates nearly half of all committed code, but the security gap is not shrinking.
The difference is clear: the models produce consistent code with a near-universal syntax pass rate of 100 percent. When it comes to security, they fail about 44 percent when they are not given security-specific guidance.
“As AI-fueled code velocity increases, developers are inundated with compliance risks, security warnings, and quality issues,” said Chris Wysopal, Co-founder and Chief Security Evangelist at Veracode. “We are seeing a rapid increase in the adoption of AI-powered tools for coding and building software. But the main problem remains: the models may be almost perfect, but they still fail in about half of all tasks where security is required. That number should be a red flag for any organization.”
GenAI Code Security Report leaderboard: Summer 2026

Figure 1: LLM Leaderboard by Security Pass Rate
OpenAI’s GPT-5.5 leads this year’s report with 68 percent, while six of the 11 models score between 50 percent and 53 percent. Alibaba’s Qwen3.7-max reaches the final 50 percent, producing a code that is vulnerable to all other results in this test. The best model available today still fails about one in three safety functions. Notably, previous editions of the report were dominated by Western AI models. Not anymore, as the Kimi-K2.6 and MiMo-V2.5 are more efficient than several Western models. For enterprise security teams, the evolution of the model is another factor that should be measured in procurement decisions, as well as the level of security effectiveness.
Built-to-Code Models and large LLMs are not safe
Two widely held assumptions do not survive the data. First, models built specifically for software coding are less secure than general-purpose AI. The special coding models had a 51 percent security pass rate, compared to 52 percent for the general purpose models. Testing was done against raw models and not agents or a production environment with additional tools, guardrails, or human review in the loop. This means that developers who choose code-optimized tools with the assumption that they will reduce security risks are actually deploying vulnerable code at the same rate as everyone else.


Figure 2: Defense Pass Rate for Coding-specific vs General-purpose LLMs
Second, the size of the model has no effect on the security performance. Large models (more than 100 billion parameters) are in the middle of 53 percent, medium models are in the middle of 51 percent, and small models are in the middle of 51 percent. One architectural feature that comes to mind: cognitive models maintain a consistent margin of safety over non-cognitive models, at 56 percent versus 51 percent. This suggests extended reasoning functions as a way to update the internal code.


Figure 3: Protection Passage Level of Models by Size
Java Still Catches You – But It’s Going in the Right Way
Language security pass rates range from Python at 63 percent down to Java at just 30 percent. Java is still the most dangerous language for generating AI code by a significant margin. Despite low performance, it is the only language with a clear, consistent upward trend over the past year.
“I’ve been talking about making advanced AI models, like Claude Fable and Mythos, available to developers and defenders alike – and I stand by that position,” Wysopal concluded. “The right answer is not restricting access; it’s clear, evidence-based security. This research makes it clear that AI-generated code needs to be treated like any unreviewed code: scan it, fix it, and never send it blindly. Until LLMs think about security the way they think about syntax, guard lines in job development are not an option.”
Managing Risk in the AI Era
Veracode recommends that organizations take the following steps to address security gaps in AI-generated code:
- Integrate AI-powered tools like Veracode Fix into developer workflows to fix security risks in real-time.
- Embed security into agent workflows to automatically implement secure coding standards.
- Use Software Architecture Analysis (SCA) to detect third-party vulnerabilities and open source dependencies in AI-generated code.
- Use Package Firewall to block vulnerable, malicious or non-compliant packages before they reach the development environment.
To download the full GenAI Code Security 2026 report, visit the Veracode website. Attendees at Black Hat in Las Vegas August 4-7 are invited to visit booth #4927 to learn more about AI code security and the report’s findings.
About Veracode
Veracode is a global leader in Application Risk Management for the AI era. Powered by billions of lines of code scanning and an AI-assisted debugging engine, the Veracode platform is trusted by organizations worldwide to build and maintain secure software from code creation to cloud deployment. Thousands of the world’s leading development and security teams use Veracode every second of every day to gain accurate, actionable visibility into exploitable risk, achieve real-time vulnerability remediation, and dramatically reduce their security liability. Veracode is an award-winning company that provides security capabilities for the entire software development lifecycle, including Veracode Fix, Static Analysis, Dynamic Analysis, Software Composition Analysis, Container Security, Application Security Posture Management, Malicious Package Detection, Package Firewall, and Penetration Testing.
Learn more at www.veracode.com, the Veracode blog, and LinkedIn and X.
Copyright © 2026 Veracode, Inc. All rights reserved. Veracode is a registered trademark of Veracode, Inc. in the United States and may be registered in certain other jurisdictions. All other product names, brands, or logos are the property of their respective owners. All other trademarks mentioned herein are the property of their respective owners.
Press and media contacts
Katy Gwilliam
Head of Global Communications, Veracode
[email protected]



