Blog · AI and security

AI writes code that compiles, but in 44% of Veracode's tests it introduced a vulnerability

Veracode's 2026 report on the security of AI-generated code found that models almost always write syntactically correct code, but about 44% of the test tasks introduced a risky vulnerability. And Verizon's 2026 DBIR, cited in the report, ranks software vulnerabilities as the top entry point in breaches, at 31%.

Code on a screen
Code on a screen. Cropped to 16:9. Photo: Sai Kiran Anagani · CC0 · Wikimedia Commons

What Veracode measured

Veracode gives each model programming tasks and checks whether the result avoids known weaknesses (CWEs). The average security pass rate was 56%, nearly the same as the 55% in its first report. Syntax, on the other hand, is practically solved.

The best model in the evaluation, GPT-5.5, reached 68%: it still fails almost one in three security tasks. More than half of the models landed between 50% and 53%. Code-specialized models averaged 51% and general-purpose models 52%. Large models scored 53%; medium and small ones, 51%. Reasoning models showed a small edge: 56% versus 51%.

Where does it fail, and why does it matter?

It depends on the type of flaw. On SQL injection, the models passed 83% of the time, and on cryptographic algorithms, 87%. On cross-site scripting they dropped to 15%, and on log injection, to 12%. According to Veracode, problems that depend on data flow and application context are the ones models don't learn as a repeatable pattern.

What changed this year is scale: in organizations that have adopted these tools, AI already writes about half of the code that gets committed. The same failure rate now applies to much more code, and review work and remediation queues grow with it. The report's conclusion is direct: no model generates secure code reliably enough to stop verifying it.

Sources

  1. Veracode, 2026 GenAI Code Security Report

Consulting: technical leadership →

← Back to the blog