Why vm2 and node:vm Can't Sandbox LLM-Generated Code
node:vm is not a security boundary and vm2 has 65 advisories this year. How to run model-written JavaScript in Node.js behind a process and microVM boundary.
You asked the agent for a "run this snippet" feature. Maybe it is a code-interpreter tool for your chatbot, a formula field in a spreadsheet product, or a workflow step where users (or the model) write a little JavaScript to transform data. The agent came back with something like this, and it works on the first try:
// app/api/run/route.ts (vulnerable)
import vm from "node:vm";
export async function POST(req: Request) {
const { code } = await req.json();
const result = vm.runInNewContext(code, {}, { timeout: 1000 });
return Response.json({ result: String(result) });
}
It has a timeout. It runs in a "new context". It even has a fresh, empty object as the global. Every test you write against it passes, because every test you write sends it 1 + 1.
Now send it this:
this.constructor.constructor("return process")().env
That returns your server's environment: database URL, API keys, everything. One more step and the snippet calls require("child_process") and runs shell commands as your app. If you replaced node:vm with vm2, you are in a better position, but not a safe one: GitHub's advisory database lists 65 vm2 advisories published in 2026 alone, 29 of them in the last week. This guide explains why in-process JavaScript sandboxes keep breaking, and how to build a boundary that holds when the code you are running was written by a model, or by someone talking to one.
Why agents reach for node:vm and vm2
When a prompt says "let users run JavaScript" and says nothing about who those users are, the model writes the shortest code that runs JavaScript. In Node.js that is eval, new Function, or node:vm. A model that has absorbed a little security folklore upgrades to vm2, because years of Stack Overflow answers and blog posts say vm2 is "the sandbox for Node".
None of those choices fail a functional test. A sandbox that leaks only fails when someone tries to escape it, and nobody in the happy-path test suite does. This is the same structural problem behind most AI-written security bugs, which we covered in why AI coding agents keep writing the same vulnerabilities: insecure code works, so "it works" is no signal.
There is also a training-data lag. A model trained before this year has seen vm2 described as abandoned (two escapes in July 2023 shipped with "Patches: None") and also as revived (3.10.0 landed in October 2025). It has no way to know that dozens of new escapes were published since. It will happily pin whatever version it remembers.
What node:vm actually is
The Node.js documentation is unambiguous. The first thing on the node:vm page is:
The
node:vmmodule is not a security mechanism. Do not use it to run untrusted code.
A context created by node:vm is a separate set of JavaScript globals inside the same V8 isolate, the same heap, and the same process as your server. It exists so tools like test runners and REPLs can evaluate code with a different global object. It was never designed to stop that code from reaching the host.
How the escape works
The payload above uses no bug. In vm.runInNewContext(code, {}), the {} you passed is created in your realm. Inside the context it becomes the global object, so this refers to a host object. this.constructor is the host's Object, and Object.constructor is the host's Function. A function built by the host's Function constructor runs with the host's globals, where process exists.
You might try to close that path with a null-prototype global:
import vm from "node:vm";
const ctx = vm.createContext(Object.create(null));
vm.runInContext(`this.constructor.constructor("return process")()`, ctx);
// ReferenceError: process is not defined
That particular payload now fails. But the moment you expose anything useful, the leak returns, because every host object and function you hand in carries a reference back to the host realm:
import vm from "node:vm";
const ctx = vm.createContext(Object.create(null));
ctx.log = (...args) => console.log(...args); // one host function is enough
vm.runInContext(`log(log.constructor("return process")().env.HOME)`, ctx);
// prints the host's HOME
We ran all three snippets on Node.js 24; they behave as shown. A sandbox you cannot pass a logger, a data object, or a callback into is not a sandbox anyone ships, which is why "harden the context object" is a losing game.
The timeout does not hold either
The timeout option looks like protection against while (true) {}, and for synchronous code it is. Promises are different. The Node docs have a section on how microtasks scheduled inside the context can run after runInNewContext returns, outside the timeout. This snippet blocks the host's event loop for 1.5 seconds even though the timeout is 50 ms:
import vm from "node:vm";
vm.runInNewContext(
`Promise.resolve().then(() => { const end = Date.now() + 1500; while (Date.now() < end) {} })`,
{},
{ timeout: 50 },
);
Passing microtaskMode: "afterEvaluate" brings those microtasks back inside the timeout (we measured it throwing ERR_SCRIPT_EXECUTION_TIMEOUT after about 50 ms). That fixes the denial of service, not the escape. Change the loop to while (true) and, without that option, your whole server stops answering.
Why vm2 keeps breaking
vm2 wraps node:vm and adds a membrane: proxies around every object that crosses between host and sandbox, so the sandbox never holds a raw host reference. The idea is sound. The difficulty is that the membrane must be complete. Every built-in, every error path, every promise species lookup, every Node API that hands back an object, has to be wrapped correctly, on every Node.js and V8 version. An attacker needs one gap.
The 2026 advisories show what those gaps look like. A selection from the batch published on October 1 (all fixed in 3.11.7) and October 5 (fixed in 3.11.8 and 3.12.2):
- GHSA-647f-g98j-qq25 (CVE-2026-92937, critical): a bypass of an earlier fix through
call/applyindirection, leading to host code execution. - GHSA-27g9-p43v-cw3v (CVE-2026-92944, critical): an escape that only works on Node.js 26, through a stale V8 protector. Your sandbox's safety depended on your Node version.
- GHSA-8686-vhfx-7r3j (CVE-2026-92957, critical): a
NodeVMdeny rule fornode:child_processdid not block plainchild_process. - GHSA-fcqc-726x-5wfc (CVE-2026-92947, critical): sandboxed code could read and write host memory through Node's shared
Bufferpool. - GHSA-88hf-g992-jg85 (CVE-2026-92955, critical): an escape built on obtaining the host
__proto__accessor, a primitive the report notes had appeared in several earlier reports. - GHSA-5h3f-q97h-ccvc (CVE-2026-100721, critical): a custom module resolver stored allowed paths as a prefix with no separator boundary, so
foo2/passed an allowlist written forfoo.
Read the list as a pattern, not a checklist. These are not one sloppy function. They are the attack surface you inherit whenever the untrusted code shares a heap with your secrets. Our news write-up of the October 1 batch has the full list.
The current vm2 README is honest about this. It says researchers "continuously discover new ways to escape the vm2 sandbox," and recommends it only where you need tight host integration and the code is relatively trusted. LLM output that a user can steer with a prompt is not relatively trusted.
What about isolated-vm?
isolated-vm is a real step up: each Isolate is a separate V8 heap, so there are no shared object references to walk. But read its README before you treat it as the answer:
- The project is in maintenance mode.
- It states that using it "to run untrusted code does not automatically make your application safe," and advises running isolates that handle untrusted code in separate processes.
memoryLimitis "more of a guideline instead of a strict limit"; a determined attacker can use two to three times the limit.- On Node.js 20 and later you must start node with
--no-node-snapshot.
isolated-vm is a reasonable inner layer. It is still the same process as your server, and a V8 bug is a host compromise.
The boundary that holds: the operating system
The rule that falls out of all this is simple: untrusted code and your secrets must not share a process. Put the code somewhere that, if it fully escapes the JavaScript engine, finds nothing worth having and no way out.
Layer 1: a separate, stripped-down process
This runs the code in a child Node.js process with an empty environment, its own temp directory, a hard kill on timeout, a memory cap, and capped output. It uses the Node.js permission model to deny file access outside the work directory, child processes, worker threads, native addons and WASI:
// lib/run-untrusted.ts
import "server-only";
import { spawn } from "node:child_process";
import { mkdtemp, realpath, rm } from "node:fs/promises";
import { tmpdir } from "node:os";
import path from "node:path";
const MAX_OUTPUT = 64 * 1024;
export type RunResult = {
exitCode: number | null;
signal: NodeJS.Signals | null;
stdout: string;
stderr: string;
};
export async function runUntrusted(
code: string,
{ timeoutMs = 2000 }: { timeoutMs?: number } = {},
): Promise<RunResult> {
// realpath: on macOS tmpdir() is a symlink, and grants are matched on real paths
const workdir = await realpath(await mkdtemp(path.join(tmpdir(), "run-")));
try {
return await new Promise((resolve) => {
const child = spawn(
process.execPath,
[
"--permission", // fs, child_process, workers, addons, WASI denied
`--allow-fs-read=${workdir}`,
`--allow-fs-write=${workdir}`,
"--max-old-space-size=64",
"-", // read the program from stdin
],
{
cwd: workdir,
env: {}, // no API keys, no DATABASE_URL
stdio: ["pipe", "pipe", "pipe"],
timeout: timeoutMs,
killSignal: "SIGKILL",
},
);
let stdout = "";
let stderr = "";
const collect = (chunk: string, stream: "out" | "err") => {
if (stream === "out") stdout += chunk;
else stderr += chunk;
if (stdout.length + stderr.length > MAX_OUTPUT) child.kill("SIGKILL");
};
child.stdout.setEncoding("utf8").on("data", (c: string) => collect(c, "out"));
child.stderr.setEncoding("utf8").on("data", (c: string) => collect(c, "err"));
child.on("close", (exitCode, signal) =>
resolve({
exitCode,
signal,
stdout: stdout.slice(0, MAX_OUTPUT),
stderr: stderr.slice(0, MAX_OUTPUT),
}),
);
child.stdin.end(code);
});
} finally {
await rm(workdir, { recursive: true, force: true });
}
}
We exercised this on Node.js 24: console.log(6 * 7) returns 42; reading /etc/hosts and calling child_process.execSync both fail with ERR_ACCESS_DENIED; while (true) {} is killed with SIGKILL; an allocation loop dies at the heap cap; writing inside the work directory succeeds.
Two limits you must know about, both from the Node docs:
- The permission model is a seat belt, not a sandbox. The documentation says it "does not provide security guarantees in the presence of malicious code" and that "malicious code can bypass the permission model." It also notes that symbolic links are followed outside granted paths. Treat it as defense in depth.
- Network is not covered on Node.js 24. The
--allow-netflag arrived in Node.js 25.0.0. On the current LTS line, the child above can still callfetch(); we confirmed it reaches the internet. With an empty environment there are no keys to send, but the code can still reach your internal network and cloud metadata endpoints. That is an SSRF problem waiting to happen.
What layer 1 buys you is real: the code no longer shares a heap with your secrets, and an engine-level escape lands in a process with an empty environment. It is not enough on its own for code anyone on the internet can submit.
Layer 2: a container or microVM
The layer that actually contains a hostile payload is kernel-level isolation with no network. With Docker, the flags that matter are:
docker run --rm -i \
--network none \
--read-only \
--tmpfs /tmp:rw,size=16m \
--cap-drop ALL \
--security-opt no-new-privileges \
--pids-limit 64 \
--memory 128m --cpus 0.5 \
--user 65534:65534 \
node:24-slim node -
--network none removes the network interface entirely, which closes the SSRF path layer 1 leaves open. --read-only plus a small tmpfs gives the code scratch space and nothing else. --cap-drop ALL, no-new-privileges and a non-root user shrink what a kernel-level exploit has to work with, and the pids/memory/CPU limits stop fork bombs and resource exhaustion. Run the layer 1 runner inside that container and you have two independent boundaries.
Plain containers share the host kernel. For code anyone can submit, a user-space kernel such as gVisor or a microVM such as Firecracker puts a much smaller interface between the payload and the host. If you deploy on serverless, you probably cannot run Docker from a function at all; managed microVM sandboxes exist for exactly this. Vercel Sandbox, for example, runs each sandbox in a Firecracker microVM with its own filesystem and network.
What the model's code should be allowed to see
Whatever the runtime, decide the inputs and outputs explicitly:
- Inputs: pass data in as serialized JSON over stdin or a file, never as live objects or callbacks. A callback is a reference back into your process.
- Outputs: treat stdout as untrusted text. Cap its size, parse it with a schema, and never feed it back into
eval, a shell, SQL, or an HTML response without the same escaping you would apply to user input. - Secrets: none, ever. If the code needs to call an API, it asks your server through a narrow, authenticated endpoint, and your server holds the key.
How to find this in your codebase today
Start with where dynamic code runs. From the repo root:
grep -rnE "from ['\"](node:)?vm2?['\"]|require\(['\"](node:)?vm2?['\"]\)" \
--include=*.ts --include=*.tsx --include=*.js --include=*.mjs \
--exclude-dir=node_modules .
grep -rnE "\beval\(|new Function\(|runIn(New|This)?Context\(|compileFunction\(|new vm\.Script\(|new (NodeVM|VM)\(" \
--include=*.ts --include=*.tsx --include=*.js --include=*.mjs \
--exclude-dir=node_modules .
Then check what is in your dependency tree, since a library may be doing this for you:
npm ls vm2 isolated-vm
For each hit, answer three questions in review:
- Where does the code string come from? A user, a model, a database row a user can edit? If any of those, it is untrusted, whatever the UI says.
- What process does it run in? If the answer is "the same one that holds
process.env," it is a finding regardless of library. - What can it reach? Network, filesystem, other tenants' data, cloud metadata.
A static scanner can find the first and part of the second: GuardVibe's VG014 flags eval, new Function and the vm.run* family, VG1036 flags code-execution tools configured with sandbox protections turned off, and its CVE rules flag vulnerable vm2 pins in package.json. What no pattern-based scanner can do is prove where a string came from across your whole app, or tell you whether your container has a network. Those are review questions.
How to prevent it
In order of leverage:
- Say it in the prompt. "Users can submit JavaScript; treat it as hostile; run it in an isolated process with no network and no secrets; do not use node:vm or vm2." Put that line in your
AGENTS.mdor rules file, so every feature that touches code execution starts from it rather than from the median tutorial. - Question the feature. Many "run code" features are really expression evaluation. A formula language, JSONata, or a JSON-logic interpreter that cannot express
processis safer than any sandbox. - Give the codebase one way to run untrusted code. A single
runUntrusted()module, markedserver-only, that every caller uses. Agents copy the patterns they find nearby; make the safe one the nearby one. - Gate it in CI. Fail the build on new
node:vm,vm2,evalornew Functionimports outside that module, and on vulnerable vm2 versions. If you keep vm2 anywhere, keep it at the latest release (3.12.2 at the time of writing) and treat every advisory as urgent. - Keep Node current. isolated-vm's own security advice starts with V8 patches, and one of this year's vm2 escapes depended on the Node major.
The takeaway
If untrusted code runs in your server's process, the only question is how long until someone finds the way out. node:vm says so in its first line of documentation; vm2's 2026 advisory count says so in practice. Move the code into a separate process with nothing in its environment, put that process in a container or microVM with no network, and pass data in and out as plain text.
If you want to see where your codebase stands first, npx guardvibe scan . lists the dynamic-execution sites and vulnerable sandbox pins. Then the review questions above take it the rest of the way. For supply-chain risks that sit right next to this one, see slopsquatting and hallucinated package names.