wGrow
menu
Running Agent Scratchpads in WebAssembly Isolates
Infra & Security 6 October 2026 · 7 min

Running Agent Scratchpads in WebAssembly Isolates

By wGrow Project Team ·

An internal data extraction agent was asked to clean a CSV. It decided the job needed a Python library, wrote the import, and ran pip install for it. The library did not exist. The model had invented the name.

The install failed and nothing happened. That was luck, and I don’t count luck as a control.

The Hallucinated Package Incident

IT professional analyzing data on dual monitors in a modern office.

This was an early wGrow prototype, and the task was ordinary: normalise dates, strip stray whitespace, fix a few column types. The agent had a generic Python interpreter and a shell, so it used both.

The failure mode matters more than the failure. A model that invents a package name is likely to invent the same name again, because the invention is a plausible token sequence, not a random one. Anyone who registers that name on PyPI turns a hallucination into a delivery channel. The next agent run resolves the name, pip install succeeds, and the attacker’s code runs inside your infrastructure with whatever credentials the process holds.

I’m not claiming this happened to us. I’m claiming the path was open. The only thing that closed it was that nobody had registered the name.

Here is the position I’ll defend in any architecture review: handing a generic Python interpreter, or an OS-level container with a package manager, to a language model is a liability that grows with every run. Agent runtimes need zero-trust execution. The model is an untrusted author, and its output is untrusted input to your runtime.

The Blast Radius of Docker Containers

The default industry answer is an ephemeral Docker container per execution. That beats running code on the host. For this job it’s still a poor fit, for two reasons.

Latency. Booting a locked-down Python container for a single intermediate transformation was slow enough to notice in our prototype. One transformation is tolerable, but agent loops rarely stop at one. A multi-step loop might clean, reshape, validate, and aggregate, and every step pays the boot cost again. Four steps means four boots before any useful work happens, and users feel it. You can cut the cost by reusing a warm container, but then state persists between calls. That’s the very property we’re trying to remove.

Mismatch of capability. A container gives you process isolation. It also gives you an OS userspace, a filesystem view, a package manager if the image includes one, and a network namespace. Each of those is a surface you then have to strip out or lock down through configuration. Get one wrong, whether an egress rule, a mounted volume, or a writable layer, and the agent has a way out.

A CSV date-normalisation step needs none of that. It needs a string in and a string out. Giving it a whole operating system is like hiring a locksmith to change a lightbulb, then spending your afternoon supervising the locksmith.

A container’s blast radius is defined by everything you forgot to remove. I want one defined by what I explicitly granted. For scratchpad work, that should be nothing.

The Execution Bridge: Shifting to V8 and JavaScript

Execution Path
step 01
LLM Generates JS
step 02
Init V8 Context
step 03
Execute Code
step 04
Return String

Take away the host Python interpreter and you have to decide what the agent writes and runs instead. We settled the prompt first and the runtime second.

The prompt change. We stopped asking the model to write Python for data cleaning and told it to generate JavaScript. The work in an intermediate scratchpad is JSON parsing, string manipulation, date handling, filtering, grouping, and arithmetic. In our experience, current models write JavaScript for all of it about as readily as they write Python. We saw no quality penalty worth reporting for this class of work, though we never ran a formal benchmark. If your workload looks different, measure before you copy us.

The runtime change. The generated JavaScript executes inside a V8 isolate. The host passes in one string and gets one string back. There is no require, no fetch, no filesystem handle, and no child_process. Those globals aren’t blocked. They’re never provided. That difference matters, because a blocklist fails the moment you miss an entry, and an empty environment has no entries to miss.

It closes the original hole too. There’s no pip in the isolate. A hallucinated import throws, the loop catches the error, and the model gets a chance to rewrite the code. The failure costs one retry and exposes nothing.

The Python fallback. Some work does need Python, usually numerical libraries with no good JavaScript equivalent. For those cases we run Pyodide, which is CPython compiled to WebAssembly, inside the isolate. The security boundary doesn’t move: the same kinds of memory and CPU limits apply, and there are still no sockets. Pyodide needs a bigger budget than plain JavaScript, so its limits sit higher than the ones below, but the host still fixes them. We treat it as the exception path. It initialises slowly, so we reach for it only when the task demands it.

Enforcing Hard Limits at the Isolate Level

Technical diagram illustrating isolated memory and CPU boundaries.

Constraint Profile
DOCKER CONTAINER
V8 ISOLATE
Network Access
virtualized
none
File System
ephemeral root
none
Memory Bounds
cgroups
strict heap
Host Surface
kernel / OS
V8 process

An isolate is not a lightweight container, and the distinction is structural. Isolates share a single host process but keep separate heaps and execution contexts. There’s no guest kernel to boot, no filesystem to mount, and no network namespace to configure. You create a context, run the code, and throw the context away.

That structure produced the difference that changed our design. In our prototype, isolate startup was roughly two orders of magnitude faster than starting a locked-down Python container. We haven’t published logs or run a controlled benchmark, so measure this on your own runtime before adopting the pattern. Even so, the gap was enough that a four-step loop went from paying a visible startup cost to paying one users don’t notice.

The limits are where the security value sits. For plain JavaScript, we configure each execution context with:

  • A hard memory cap of 64MB. An agent that writes an accidental infinite array, or a deliberate allocation bomb, hits the ceiling and the isolate is terminated. The isolate is terminated before unbounded heap growth, and the host enforces a per-execution ceiling instead of trusting the model to stop.
  • A strict CPU timeout of 50ms. A runaway loop gets killed by the clock, not by the model’s good intentions.

Treat these as example settings and tune them to your workload. 50ms suits small transformations and will be too tight for large inputs. What matters is that the limits belong to the runtime. The model can’t negotiate with them, and no prompt can talk them away.

Network isolation is the part I care about most. An isolate has no system socket access unless the host injects a capability, and we inject none. The agent can’t open its own outbound network path, reach the internet, or download a package. Its only channel is the returned string, so the orchestrator still has to validate, redact, and control where that output is logged or forwarded. Say an attacker poisons the agent through a malicious document inside the CSV. The poisoned code still has nowhere to send anything directly.

The residual risk is real. Isolates share a process, so the boundary is only as strong as the V8 engine and the embedding code around it, and a V8 vulnerability could become an escape. That’s why the isolate’s process environment holds no secrets and no credentials worth stealing, and why we keep the engine patched. Defence in depth still applies. But the attack surface shrinks from “an operating system” to “a JavaScript engine with nothing attached”.

The Standard for Agent Runtimes

The heavy OS container is a poor default for an agent scratchpad. It still earns its place when the workload genuinely needs an OS, like building software or running a browser. For intermediate data transformations, it’s usually the wrong tool.

The standard I’d hold agent loops to is simple. Each execution runs in an ephemeral, capability-minimal isolate. If you need deterministic replay, remove or seed nondeterministic globals such as Date and Math.random, and keep host-injected functions pure. It takes one input string, performs one transformation, returns one output string, and ends within milliseconds. Nothing persists between calls. State lives in the orchestrator, where you can inspect and audit it.

The bigger lesson is about where security lives. A system message that says “do not run dangerous code” is a request, not a control. Models hallucinate package names, and they can follow instructions injected into the documents they read. Neither failure responds to politeness.

Security in agentic systems has to be enforced by the runtime. That means memory ceilings, CPU budgets, and an environment with no network. If the agent misbehaves, the limits should make the misbehaviour harmless, not just unlikely. Build the sandbox so the model’s worst output is a string that comes back and gets discarded.