GuardVibe
Blog
· 12 min read

Why AI Coding Agents Keep Writing the Same Vulnerabilities

Newer models write code that works, not code that is secure. Five structural reasons the same bugs recur in AI-generated apps, and what actually stops them.

You have left the same review comment three times this month. The agent put an API key behind NEXT_PUBLIC_. The agent wrote a server action that checks the user is logged in and then updates whatever row ID the client sent. The agent built an ORDER BY from a query parameter with sql.raw. You fixed each one, and next week a different feature came back with the same three bugs in new clothes.

It is tempting to read that as a model quality problem that the next release will fix. The research says otherwise, and the reasons are more useful than the statistics: once you see why the same bugs recur, you can change the conditions that produce them instead of catching them one review at a time.

What the research actually shows

Three studies are worth knowing, with their limits stated up front.

  • Pearce et al., IEEE S&P 2022 had GitHub Copilot complete 89 security-relevant scenarios, producing 1,689 programs. Roughly 40% were vulnerable.
  • Perry et al., ACM CCS 2023 ran a user study. Participants with an AI assistant wrote significantly less secure code than those without one, and were more likely to believe their code was secure. Participants who trusted the assistant less and worked on the wording of their prompts produced fewer vulnerabilities.
  • Spracklen et al., USENIX Security 2025 analyzed 576,000 code samples from 16 models and found hallucinated package names in at least 5.2% of commercial-model output and 21.7% of open-source-model output: 205,474 unique names that do not exist.

The honest objection is that the first two studies used 2021–2022 models, and today's are far better. That is true for functionality. Veracode's 2025 GenAI Code Security Report tested over 100 models and found that 45% of generated samples introduced an OWASP Top 10 flaw (43% for JavaScript, and 86% of samples failed to defend against cross-site scripting). Their summary of the trend is blunt: models got better at writing code that works, while security performance "remained flat, regardless of model size or training sophistication."

Flat across model sizes is the important part. It means the cause is not missing capability. It is structural.

Five reasons the same bugs keep coming back

1. "It works" is the goal, and insecure code works

Every signal an agent gets during a task rewards working code: the prompt describes behavior, the test checks behavior, the error message disappears when the behavior is right. A missing ownership check does not break anything. It only changes who the code works for. From inside the task, a vulnerable solution and a secure one are indistinguishable, and the vulnerable one is usually shorter.

2. The prompt leaves security unsaid, so the model fills in the median

"Let users edit their posts" says nothing about who owns a post. When a requirement is missing, the model supplies the most typical version of the code, and the most typical version of web code in public is tutorial code: short, optimized to demonstrate one idea, and routinely without auth, input validation or secret handling, because those would distract from the lesson. A model that writes the median of what it has seen writes the median tutorial's shortcuts.

Perry et al.'s finding that prompt engagement reduced vulnerabilities is the flip side: when the requirement is in the prompt, it tends to be in the code.

3. The agent sees one file; authorization lives in the whole app

Whether a route is protected depends on the middleware matcher (proxy.ts in Next.js 16), on which data helpers check ownership, on whether a helper is async, on the other twelve routes that do it correctly. An agent editing one handler sees a fraction of that. The code it writes is locally consistent and globally wrong.

This week's fail-open authorization bugs are good examples of the shape: a permission check that is present, reads correctly in isolation, and does nothing because of a fact that lives somewhere else (that hasPermission() returns a Promise).

4. The writer is reviewing the writer

Asking the same model, in the same context, whether its code is secure checks the code against the same assumptions that produced it. If the model did not consider that a server action can be called directly, the review will not consider it either. Perry et al.'s confidence result applies to the human in the loop too: the output looks competent, so it gets less scrutiny than code a colleague wrote.

5. Training data has a cutoff; advisories do not

A model knows the vulnerable versions and bypasses that were public when it was trained. The advisory published last month is not in its weights. It will pin a version it remembers, reach for an API pattern that has since been deprecated for security reasons, and it cannot know about a package name that was registered yesterday by someone waiting for exactly the name it is about to hallucinate.

The three bugs you will see first

These three are worth learning first: each is easy to generate, invisible in a happy-path test, and serious when shipped. The examples use a stack agents commonly generate: App Router, Clerk, Drizzle on Postgres.

Secrets that end up in the browser

// lib/openai.ts (vulnerable)
import OpenAI from "openai";

export const openai = new OpenAI({
  apiKey: process.env.NEXT_PUBLIC_OPENAI_API_KEY,
  dangerouslyAllowBrowser: true,
});

This usually starts with a client component that calls the model directly. The key is undefined in the browser, so the agent renames it to NEXT_PUBLIC_; the SDK then refuses to run in a browser, and the agent sets the flag that overrides the refusal. Each step fixes an error. The result is your OpenAI key in the JavaScript bundle, readable by anyone who opens dev tools. Next.js inlines every NEXT_PUBLIC_ variable into client code by design.

// lib/openai.ts (fixed)
import "server-only";
import OpenAI from "openai";

export const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

import "server-only" turns "this module leaked into the client" from a silent runtime exposure into a build error. The client component calls a server action or route handler, which calls the model.

Authenticated is not authorized

// app/posts/actions.ts (vulnerable)
"use server";

import { auth } from "@clerk/nextjs/server";
import { eq } from "drizzle-orm";
import { db } from "@/db";
import { posts } from "@/db/schema";

export async function updatePost(postId: string, title: string) {
  const { userId } = await auth();
  if (!userId) throw new Error("Unauthorized");

  await db.update(posts).set({ title }).where(eq(posts.id, postId));
}

This passes every test the agent is likely to write, because in those tests the logged-in user edits their own post. But the Next.js data security guide is explicit that a server action is reachable by a direct POST request, not only through your UI. Any logged-in user can call updatePost with someone else's postId. That is an insecure direct object reference, and it is exactly what you get when the prompt said "users can edit posts" and never said "their own".

// app/posts/actions.ts (fixed): same imports, plus `and` from drizzle-orm
export async function updatePost(postId: string, title: string) {
  const { userId } = await auth();
  if (!userId) throw new Error("Unauthorized");

  const updated = await db
    .update(posts)
    .set({ title })
    .where(and(eq(posts.id, postId), eq(posts.authorId, userId)))
    .returning({ id: posts.id });

  if (updated.length === 0) throw new Error("Not found");
}

Scoping the query by owner means there is no window between checking and writing, and "not yours" and "doesn't exist" return the same answer, so the endpoint does not confirm which IDs exist.

Raw SQL inside a safe ORM

// vulnerable: the sort column comes straight from the URL
const sort = searchParams.get("sort") ?? "created_at";

const rows = await db.execute(
  sql.raw(`SELECT id, title FROM posts ORDER BY ${sort} DESC`),
);

Drizzle's sql template parameterizes every interpolated value, which is exactly why it cannot be used for a column name: ORDER BY $1 sorts by a constant. The agent hits that, finds sql.raw, and moves on. The Drizzle docs are clear that sql.raw does no parameterization at all. Now ?sort= is a SQL injection point in an app whose authors believe the ORM protects them.

// fixed: identifiers come from an allowlist, never from input
const SORTABLE = {
  created: posts.createdAt,
  title: posts.title,
} as const;

const key = searchParams.get("sort");
const column =
  key && Object.hasOwn(SORTABLE, key)
    ? SORTABLE[key as keyof typeof SORTABLE]
    : posts.createdAt;

const rows = await db
  .select({ id: posts.id, title: posts.title })
  .from(posts)
  .orderBy(desc(column));

Object.hasOwn rather than key in SORTABLE matters: "toString" in SORTABLE is true, because in walks the prototype chain.

Finding them in a codebase you already have

Start with a few greps. They are crude and produce false positives, but they run in seconds and point you at the files worth reading:

# Secrets exposed to the client (some public keys, like Firebase's, will match: check each)
grep -rnE "NEXT_PUBLIC_[A-Z0-9_]*(SECRET|PRIVATE|SERVICE_ROLE|API_KEY)" --include="*.ts" --include="*.tsx" .
grep -rn "dangerouslyAllowBrowser" --include="*.ts" --include="*.tsx" .

# SQL that bypasses parameterization
grep -rnE 'sql\.raw\(|\$queryRawUnsafe|\$executeRawUnsafe' --include="*.ts" .

# Server action files that never call auth() (misses actions that delegate to a helper)
grep -rl '"use server"' --include="*.ts" --include="*.tsx" app src 2>/dev/null | xargs grep -L "auth("

Be clear about what automation can and cannot find. Pattern rules reliably catch the first and third bugs: a secret behind a public prefix, a raw SQL call with interpolation. They structurally cannot catch the second in general, because no scanner knows that authorId is your ownership column, or that a particular table is meant to be public. That takes a convention the tooling can check (below), an LLM-assisted pass focused on business logic, and tests that encode who is allowed to do what. GuardVibe splits the work the same way: deterministic rules for the patterns, an auth_coverage map of which routes have any guard at all, and an opt-in deep-scan that uses a model specifically for IDOR and logic flaws. Whatever you use, don't let a clean pattern scan stand in for the authorization question.

Stopping them, in order of leverage

1. Change the shape of the code so the safe way is the short way

The agent writes the shortest thing that works, so make the secure version the shortest. The Next.js docs recommend a Data Access Layer, and it is the single highest-leverage change:

  • All database access lives in src/data/*, and each file starts with import "server-only".
  • Every function there takes the current user and scopes its query by it. getPost(userId, postId), never getPost(postId).
  • Only that layer reads process.env.
  • Authorization helpers throw instead of returning booleans, which removes the whole class of checks that silently pass (the fail-open write-up covers why).

Once this exists, the agent imports getPost and gets ownership for free. Writing the unscoped query means going around the codebase's conventions, which models do far less often than they skip a requirement nobody stated.

2. State the requirements where the agent will read them

A rules file (CLAUDE.md, .cursor/rules, AGENTS.md) is read before every task. Make it specific. "Write secure code" changes nothing; rules that name your paths and patterns do:

## Security rules
- Every "use server" function and route.ts handler gets the user from auth() and
  goes through src/data/*. No direct db imports anywhere else.
- Every query in src/data/* is scoped by the caller's userId.
- Never create a NEXT_PUBLIC_ variable for a secret. Only src/data/* reads process.env.
- Never use sql.raw, $queryRawUnsafe or string-built SQL. Dynamic column names
  come from an allowlist object.
- Authorization helpers throw; they never return booleans.
- Before adding a dependency, confirm it exists on npm and report its weekly downloads.

The last rule is cheap insurance against slopsquatting: a hallucinated name either 404s or turns out to be a days-old package with no downloads, and either answer is worth seeing before npm install runs its scripts.

3. Put an independent check inside the loop, not after it

Reason four says the writer cannot review itself, so the check has to come from somewhere else, and it has to run while the agent is still working, when a fix costs one edit rather than a review cycle. Three layers, all deterministic so the same code gets the same answer every time:

  • a scan after each edit, via your agent's hook system;
  • a diff-aware pre-commit gate, so legacy findings don't block new work;
  • CI with results in a format your code host understands, such as SARIF.

Add one negative test per protected route: user B cannot read, update or delete user A's record. It is the test the agent never writes on its own, and it is the only thing that verifies authorization end to end.

4. Treat new dependencies as untrusted until checked

Verify the package exists and has a history before installing it, commit your lockfile, and run npm audit in CI. Reason five means your agent cannot know about last week's advisory; your tooling can, if it pulls advisory data at run time instead of relying on anyone's memory.

The short version

Models will keep getting better at writing code that runs. None of the five reasons above is about capability: the objective rewards working code, the prompt omits requirements, the context is one file, the reviewer is the author, and the training data stops. Each one is a property of how code gets generated, which is why the same three bugs survive every model upgrade.

You can't fix that inside the model. You can fix it around the model: code conventions that make the secure path the default, requirements written where the agent reads them, and a check that isn't the author, running while the code is being written.

If you want that last piece without assembling it yourself, npx guardvibe init claude wires GuardVibe into Claude Code as an MCP server with a post-edit scan hook and a starter set of security rules. It runs locally and needs no account.

Get the next one in your feed reader

Follow GuardVibe in your feed reader. No account, no email.