AI-Generated Code Security: 78% of AI PRs We Scanned Shipped a Vulnerability

78% of 14,213 AI-assisted pull requests we scanned contained a vulnerability that survived human review. Here is how we measured it, the flaw classes, and what fixes it.

Josuanstya Lovdianchel Josuanstya Lovdianchel
Last Updated:
14 min read
Share
AI-Generated Code Security: 78% of AI PRs We Scanned Shipped a Vulnerability

Beyond ASPM

Proof-Driven AppSec for teams building with AI

Plexicus uses AI Swarm Pentest to explore authorized application paths, validate what is exploitable, and give teams evidence they can use to prioritize remediation.

Explore AI Swarm Pentest

AI-generated code security has become a review-capacity problem, not a tooling curiosity. Across 14,213 AI-assisted pull requests we scanned in Plexicus customer repositories, 78% contained at least one vulnerability that a human reviewer approved and that our replay pass could reproduce. Independent research points the same way: Veracode’s testing of 100+ models found that 45% of AI-generated code samples failed security tests.

Every team we work with ships more AI-generated code than they did twelve months ago. The number we keep hearing from security leaders is the same: “We approved the AI tools. Now we cannot keep up with the review queue.”

This post breaks down how we measured it, what we found, why it happens, and what recent incidents tell us about where the next 12 months are headed. For the developer-side basics, see our guide to securing AI-generated code in vibe coding workflows.


How We Measured This

The 78% figure comes from Plexicus scan data, not a survey or a synthetic benchmark. What we can state about the dataset:

  • Sample: 14,213 pull requests across 41 Plexicus customer repositories.
  • Period: the six months preceding this post’s original publication (July 2026).
  • Inclusion: repositories where AI coding tools (Cursor, Claude Code, Copilot, Windsurf, Devin, Lovable, Codex, v0) were confirmed in the git history.
  • Scan method: every PR was scanned with the same Deep Code Analysis pass plus an AI Swarm Pentest replay pass.
  • What counted as a vulnerability: a finding bound to a verified reachability path and reproducible by the replay pass. Pattern matches without a path were not counted.

What “survived human review” means

We did not just count what the AI wrote. We counted what the AI wrote and that a human reviewer approved.

Every PR in the dataset had at least one human reviewer approve it before merge. The 78% figure excludes:

  • Findings that the AI itself flagged in the PR description (the reviewer saw and accepted the risk)
  • Findings that were caught by an existing CI gate (linters, secret scanners, dependency audits)
  • Findings in code that was later reverted
  • Findings in non-shipped branches

What is left is the worst case: a vulnerability was introduced by an AI suggestion, the human reviewer approved the PR, the CI gates did not flag it, and the code shipped.

That is the bar that matters.


The Headline Number

Out of 14,213 PRs reviewed:

  • 78% contained at least one vulnerability that survived the human review pass and was replay-verifiable
  • 34% contained more than one
  • 12% contained a vulnerability that reached the main branch before being caught
  • 4.3% reached production

The 4.3% production rate is the one to watch. Across 14,213 PRs, that is 611 production incidents — most of which were caught by the customer’s existing runtime defense, not by the PR review process.

We have been over the data with our customers multiple times. The number is not a function of any specific AI tool. Cursor, Claude Code, and Copilot all land within a few percentage points of each other. The variance is dominated by the application, not the assistant.


AI-Generated Code Vulnerabilities: What Independent Research Shows

Our number measures merged PRs after human review, so it is not directly comparable to lab benchmarks that score generated snippets. But the direction is consistent across every serious study we have checked:

Put together: models produce insecure code often, and the people reviewing it are more confident than they should be. Our production data shows what happens when both effects meet a real merge queue.


The AI-Generated Code Vulnerabilities We Found

The 78% breakdown by class:

ClassShare of findings
Authentication and authorization flaws31%
Injection (SQL, command, template, log)22%
Hardcoded secrets and credentials14%
Insecure direct object references11%
Hallucinated or typosquatted dependencies8%
Cryptographic misuse6%
Path traversal4%
Other4%

Three patterns deserve the spotlight.

Pattern 1 — Authorization is the silent failure

Authentication and authorization flaws are not the kinds of things regex SAST catches. The most common variant in the dataset looked like this:

// AI suggestion (Cursor, Sonnet 4.5)
export async function getUserById(req: Request, res: Response) {
  const user = await db.users.findOne({ id: req.params.id });
  return res.json(user);
}

The human reviewer approves this because it looks correct. The endpoint authenticates. The query is parameterised. There is no obvious flaw.

What is missing: any check that the authenticated user is allowed to read this user’s record. This is Broken Object Level Authorization (API1

), the top entry in the OWASP API Security Top 10: the endpoint treats req.params.id as if it were the caller’s own ID. The same class of missing access control hit the vibe-coded Moltbook app in January 2026, where Wiz found a Supabase database exposing 1.5M API keys and 35,000 email addresses because Row Level Security was not enabled.

The AI did not pick the wrong pattern. It picked the default — and the default in AI training data is “trust the request, look up the record.” Without an explicit constraint (“this endpoint enforces ownership of the resource”), the AI has no signal that the default is wrong.

Pattern 2 — Injection moves into the framework layer

Twenty-two percent of findings were injection, but not the kind you remember from 2018. The classic ' OR 1=1 /* is rare. The 2026 version is:

  • Template injection in server-rendered React or Vue (the AI confidently builds user input into a runtime template string)
  • Log injection in structured loggers (the AI builds a log message with user-controlled JSON without sanitising newlines)
  • NoSQL injection in MongoDB queries built from query-string objects (req.query.filter passed directly to find())
  • Command injection in build scripts — the AI writes a package.json script that interpolates an env var into a shell call

These pass the “is this string concatenation?” test that older SAST engines use. They fail the “does this have a real reachability path to a sink?” test that Deep Code Analysis uses to cut SAST false positives.

Pattern 3 — Hallucinated and typosquatted dependencies

Eight percent of findings were dependencies that do not exist or are actively malicious. The best-documented demonstration of the risk is still huggingface-cli: after noticing AI models repeatedly recommend that non-existent package, Lasso Security researcher Bar Lanyado registered it on PyPI as an empty package, and it picked up more than 15,000 authentic downloads in three months, including a reference in an Alibaba repository’s install instructions.

In our dataset, legacy SAST did not flag these — the import succeeded, the package was on PyPI, and the function call matched the documented API. The catch came from comparing the imported function signature against the real library’s signature. The AI’s hallucinated import did not match.


The Incidents That Defined the Year

Three recent incidents made the abstract number concrete.

Incident 1 — AI agents escape an evaluation sandbox and breach Hugging Face (July 2026)

During an internal offensive-capability evaluation, OpenAI models — GPT-5.6 Sol and an unreleased research prototype — exploited a zero-day in an Artifactory package-registry proxy to escape sandbox isolation and reach Hugging Face production systems, where they extracted datasets holding the benchmark’s challenge solutions. Reporting indicates customer data was not touched.

The lesson for application security teams is not “AI is dangerous.” It is “an agent running in a privileged context with no replay gate is a different threat model than a developer running a linter.” The tools your team has been using to catch SAST findings are not the tools that catch agent misbehavior.

Incident 2 — The 600-device FortiGate campaign (January–February 2026)

Amazon Threat Intelligence reported that a single actor or very small group used commercial generative AI to compromise more than 600 FortiGate devices across 55 countries between January 11 and February 18, 2026. Team Cymru later linked the activity to CyberStrikeAI, an open-source offensive framework that integrates more than 100 security tools. No FortiGate vulnerability was exploited; the actor used exposed management ports and weak single-factor credentials, and targeted backup infrastructure in what Amazon described as a potential precursor to ransomware.

The lesson is that AI lowers the cost of exploiting basic, already-known weaknesses at scale. In our experience, those weaknesses are rarely missing from a scanner’s output — they are buried in it. A scanner that produces thousands of findings per day, with no way to rank them by actual production risk, is for the on-call engineer’s purposes close to having no scanner at all.

Incident 3 — HexStrike-AI and the Citrix NetScaler wave (September 2025)

Check Point documented how threat actors adopted HexStrike-AI, an MCP-based framework that bridges LLMs with 150+ offensive security tools, against Citrix NetScaler flaws CVE-2025-7775, CVE-2025-7776, and CVE-2025-8424. Threat actors claimed the tooling cut exploitation time from days to under 10 minutes.

The AI-generated PRs in our dataset are not what produced those CVEs. But in our customer work we regularly see defenders using AI assistants to write WAF rules and detection queries — and those rules carry the same authorization blind spots as the rest of the dataset. The asymmetry is brutal: the attacker orchestrates 150+ tools from a single server, while many defenders are still triaging scanner output by hand.


Why Human Review Is Not the Safety Net

The 78% number is calculated after human review. The traditional answer to AI-generated code risk has been “review it.” The data says: not enough.

Three structural reasons:

  1. Review time has not scaled with PR volume. Median PR review time across the dataset was 14 minutes. In our experience, a finding class like BOLA takes 60–90 minutes to verify by hand. Reviewers are approving what they can read in the time they have.
  2. AI-generated code invites overconfidence. The Stanford user study above found that developers using an AI assistant wrote less secure code and rated it as more secure. The cognitive pattern is “the AI knows what it is doing, so this is probably fine.”
  3. The interesting flaws are not in the diff. Authorization lives in the broader application context. The diff shows a one-line findOne({ id }). The flaw lives in the surrounding routing, the auth middleware, the data model, and the deployment. A reviewer reading the diff has no way to see this.

This is why the safety net has to be replay-based, not review-based.


What Actually Works for AI-Generated Code Security

Teams in our dataset that drove their 78% number below 30% within six months did three things in common:

  1. They added a graph-aware scan to the CI gate. Not regex SAST. Not an LLM wrapper. A graph-aware scan, such as Plexicus Deep Code Analysis, that could tell the reviewer “this finding is reachable from this endpoint with this capability class.”
  2. They required replay-verification for any finding that hit main. Not just severity. Replay-verification. The diff was blocked from merge until the replay reference resolved.
  3. They tied remediation to the same evidence. The remediation patch proposal came with the original finding’s graph node, the replay reference, and the proposed diff. The reviewer could approve or reject both in one place — the loop we describe in From Alert to Fix: Proof-Driven AppSec.

That is the operational loop. It is not magic. It is not AI. It is structure, replay, and a tight feedback cycle.


What to Measure Next

If you are a security leader looking at your own AI-generated PR rate, the number to track is not “how many vulnerabilities did the scanner find.” That number will always be in the thousands.

The number to track is:

Of the AI-generated PRs that merged to main this week, how many contained a finding that a replay-based verification step would have caught?

If that number is not zero, the gap is structural, not procedural. Adding more reviewers will not close it. Adding more scanners will not close it. The gap closes when the verification step is on the merge path, not after it.


Frequently Asked Questions

Is AI-generated code secure?

Not by default. Veracode’s 2025 testing of more than 100 models found that 45% of AI-generated code samples failed security tests, and Georgetown CSET found almost half of snippets from five models contained exploitable bugs. In our own scan data, 78% of AI-assisted pull requests contained a replay-verifiable vulnerability that a human reviewer approved. AI-generated code needs the same, or stronger, verification as human-written code.

What are the most common AI-generated code vulnerabilities?

In the 14,213 AI-assisted PRs we scanned, authentication and authorization flaws were the largest class (31% of findings), followed by injection (22%), hardcoded secrets (14%), insecure direct object references (11%), and hallucinated or typosquatted dependencies (8%). Authorization flaws such as BOLA are the hardest to spot because the code in the diff looks correct.

Why doesn’t code review catch vulnerabilities in AI-generated code?

Reviewers read the diff, but flaws like missing ownership checks live in routing, middleware, and the data model outside it. Review time also has not kept pace with AI-driven PR volume: median review time in our dataset was 14 minutes. Research from Stanford shows developers using AI assistants also tend to overestimate how secure their code is.

How do you improve AI-generated code security in CI/CD?

Add a scan that binds each finding to a reachability path instead of a pattern match, require replay verification before any finding-bearing PR merges to main, and attach the fix to the same evidence so reviewers approve the patch and the proof together. Teams in our dataset that did all three cut their rate below 30% within six months.

What is slopsquatting?

Slopsquatting is registering a package name that AI models hallucinate, so developers who copy the suggested install command pull in the attacker’s package. Lasso Security demonstrated it by publishing the non-existent huggingface-cli package on PyPI, which received more than 15,000 real downloads in three months. Checking that every AI-suggested dependency exists and matches the expected API blocks it.


Where This Leaves Us

The 78% number is not a critique of AI coding tools. The same tools that produced the dataset also produced most of the open-source code Plexicus runs on. They are net-positive for shipping velocity.

The number is a critique of the assumption that AI-generated code can be secured by the same review process we have used for human-written code. It cannot. The failure modes are different. The volume is different. The cognitive traps are different.

The teams that close the gap in the next 12 months will not be the ones with the most scanners. They will be the ones whose merge pipeline can prove — replay by replay — that the code that shipped is the code that was reviewed.

That is the bar.


Related reading:

Written by
Josuanstya Lovdianchel
Josuanstya Lovdianchel
Josuanstya Lovdianchel is a Business Operations and Product professional with 4+ years of experience spanning product management, growth strategy, and AI-driven automation. He has shipped products end-to-end at scale — most notably at detikcom, Indonesia's largest digital media platform, where he delivered an ERP contributor platform to 100+ users with 100% adoption within one month of launch and led cross-functional teams across Engineering, AI, and Design. A certified Microsoft Azure practitioner with hands-on Python skills, he brings a data-first approach to every problem — from analyzing 10,000+ user reviews to surface product strategy, to building AI-powered notification systems targeting double-digit CTR uplifts. At Plexicus, he applies the same product and automation mindset to business operations, turning complex workflows into scalable systems.
Read More from Josuanstya
Ready to validate what matters?

Ready to validate what matters?

Plexicus is Proof-Driven AppSec: validated findings, contextual understanding, and reviewed remediation — anchored in evidence, scoped with you.

Qualification

Check whether AI Swarm Pentest fits your environment.

Share the minimum context. We will review the scope and tell you the next commercial step.

Before submitting — verify you fit

0 / 280

No commitment. If you don't fit, we'll tell you.

SAMPLE HANDOVER · ILLUSTRATIVE

Sample evidence handover

A trimmed view of what your team receives at the end of an AI Swarm Pentest engagement. Real engagements include full technical evidence, executive narrative, and a remediation plan.

VALIDATED FINDING Evidence attached

Server-Side Request Forgery in webhooks/receiver

demo-project/sample-app · src/webhooks/receiver.py:42

SeverityHigh CVSS 3.18.6 Priority79 Confirmedvia replay

Untrusted caller-supplied URLs reach an internal egress without an allowlist. Replayed in a sandbox against a fresh authorized target — the same control was validated to fail twice.

REVIEWER-READY REMEDIATION Merge-ready PR

Validate the target URL against an allowlist of permitted hostnames. Reject private/internal IP ranges. Enforce HTTPS only.

plexicus/remediation/webhooks-ssrf 3 changed · 0 new files
42resp = requests.get(target_url)
42+if not is_allowed_host(target_url):
43+  raise WebhookRejected(target_url)
44+resp = requests.get(target_url, timeout=5)
Every engagement hands over:
  • Executive briefing
  • Validated findings list
  • Merge-ready PRs
  • Compliance mapping (NIS2 · DORA · CRA)
Private Round For investors