Est.

WASM Sandboxing Guarantees at the Database Engine Boundary

WebAssembly's three primitives create provable isolation for untrusted code at database boundaries.

Contributing Engineer & Writer · · 10 min read
Cover illustration for “WASM Sandboxing Guarantees at the Database Engine Boundary”
Foreign Function and Runtime Extension Interfaces · October 8, 2026 · 10 min read · 2,194 words

A database that runs a user-defined function, hands a tool call to an AI agent, or loads a third-party extension is making a bet: that the code it just let in can't reach anything it wasn't given. WebAssembly's sandbox is the mechanism that makes that bet calculable instead of blind, and it rests on three primitives that together form a specific, bounded contract. WASM started as a way to run fast code in web browsers. At the database engine boundary, its job is different: contain code the engine did not write and cannot fully trust.

The first primitive is software fault isolation. The host's address space sits structurally outside it, not just logically off-limits but unreachable by the mechanics of how the memory is laid out. That's why this model fits the database boundary so well: UDFs, agentic tool calls, and third-party extensions are precisely the kind of untrusted, bounded workload that software fault isolation was built to contain.

The second primitive is capability-gated imports, delivered through WASI. Compare that to a normal POSIX process, which inherits broad ambient authority the moment it starts running. The database host has to hand the module a specific file path or a specific port, not open-ended access to the machine. That turns the attack surface of a UDF or an agentic tool call into something you can list and reason about, bounded by explicit grants.

The third primitive is the size limit on linear memory itself. The strength of the claim is that WebAssembly has a formal specification and a machine-checked proof of type safety behind it, which puts it on firmer ground than container-level isolation, where the guarantees are closer to convention than proof.

Why these guarantees matter at the database engine boundary

A broken sandbox in most contexts means a crashed process. A broken sandbox at the database boundary means direct access to production data, often at the scale of the whole system. That difference is why the precision of the WASM contract carries operational weight at this boundary, not just academic interest.

Traditional isolation methods at the database boundary, process separation, IAM policy, were built around a caller that behaves predictably: a human, or code a human wrote, making calls that map cleanly onto an identity, an action, and a resource. Docker's limitation instead raises cold-start latency that piles up inside agentic loops where many short tool calls happen in sequence.

An LLM-routed agent generating SQL on the fly, or deciding which tool to call next based on a model's output, doesn't reduce to a fixed identity-action-resource triple the way a traditional application does. Here, you often can't, because the shape of the next request depends on what the model decides, not on a fixed code path. That's a structural mismatch, and it's why a model that doesn't need to guess the agent's next move, deny-by-default capability grants, fits the problem better than a model built to anticipate it.

The pgwasm project shows this working in practice: Wasm functions run inside the database engine in a sandboxed space for custom data transformations, without the UDF getting ambient access to the operating system underneath it. And the performance profile lines up with what agentic database workloads actually need. Function-call overhead stays under 1 microsecond, startup runs under 10 milliseconds, and resource-monitoring overhead is negligible. For a workload that fires many small, short-lived queries, rather than one long sustained job, that startup profile means the sandbox doesn't introduce enough latency to undercut the reason for using it.

Sandbox costs in memory and throughput at database scale

None of this comes free, and the costs are knowable rather than surprising, which matters for anyone planning a deployment around them. The sandbox works well for the short-lived, per-query workloads it was built for, and starts to strain under sustained, parallel, or CPU-heavy database compute.

Memory overhead is fixed per instance and adds up fast. Typical overhead runs in the tens of megabytes for the runtime itself, on top of whatever memory the module declares it needs. Scale to hundreds or thousands of concurrent executor instances, the kind of concurrency an active database workload can generate, and that overhead becomes a real chunk of RAM before a single byte of actual application state gets counted. Planning for that allocation belongs in the deployment design from day one, before memory pressure appears in production.

Throughput costs depend on the workload, not a fixed rule. A rough way to hold these three options in mind: WASM wins on startup speed and isolation, containers win on maturity and general-purpose flexibility, and native code wins when every CPU cycle counts.

In the DuckDB-WASM path, scalar UDF calls cross the WASM boundary synchronously, a more specific friction worth understanding for agentic analytics. If one of those calls is a long-running awaited operation, an LLM inference call, for instance, it stalls that execution path. Other rows don't get to proceed while it waits, a structural trade-off that matters directly for anyone building latency-sensitive analytics pipelines on top of this stack.

Threading remains the sharpest limit, and it doesn't have a clean resolution date attached to it. Shared-memory threading needs safety guarantees that the broader systems community is still working out for WASI as of early 2026. Until that work lands, parallel compute and a wide swath of high-throughput database workloads sit outside what WASM can reasonably handle.

The JIT compiler as enforcer and failure point

Everything described so far holds only as long as the compiler enforcing it does its job correctly, and that compiler has turned out to be the place where the contract actually breaks, demonstrated in Critical-severity CVEs found and published in April 2026. The sandbox is enforced by real software, the just-in-time compiler that turns WASM bytecode into machine instructions, and that software can have bugs like any other, not a flaw in WebAssembly's design itself.

Cranelift, the code generator at the center of this, has real verification work behind it: its instruction-lowering rules are under an ongoing formal verification effort using solver-based semantics, and its register allocation gets checked by a symbolic verifier under continuous fuzzing. Those efforts raise confidence in the guarantee. They don't remove the risk that an implementation bug slips through anyway, and in April 2026, two did.

On April 9, 2026, the Bytecode Alliance released 12 security advisories against Wasmtime, two of them rated Critical at a CVSS score of 9.0. CVE-2026-34971 came from a miscompilation in Cranelift's aarch64 backend. The bug produced two different computations for what should have been the same heap address: one used for the bounds check that's supposed to catch an out-of-range access, and a separate one used for the actual memory load. Because the two didn't agree, guest WebAssembly code could get an arbitrary read and write into host memory, a complete escape from the sandbox. CVE-2026-34987 hit the Winch baseline compiler: properly constructed guest code could reach host memory entirely outside its assigned linear-memory sandbox, and a proof-of-concept showed access to roughly 32 KiB before the start of memory and close to 4 GiB after it, enough to exfiltrate data or potentially execute code on the host.

CVE-2026-34971 specifically affects systems running on aarch64, the ARM64 architecture that cost-optimized server infrastructure increasingly runs on, including a lot of the infrastructure analytics platforms favor for its price-to-performance ratio.

The scale of the April release says something on its own. Twelve advisories in one batch is triple the total number published across all of 2025, and the two Critical-rated bugs double the total count of Critical-severity advisories in Wasmtime's entire history. It points toward compiler-level sandbox escapes getting found faster going forward, not less often.

The hardware side-channel gap that logical isolation cannot close

A WASM compiler with zero bugs still can't stop every kind of leak, because some of the leakage happens below the level where software isolation operates. Speculative execution, the CPU's habit of guessing ahead and running instructions before it knows whether they're needed, doesn't respect the logical boundaries a WASM runtime draws.

Spectre-class attacks can pull information across two WASM instances running inside the same operating system process, because the CPU's speculation doesn't know or care where the runtime's logical memory lines are drawn. For a multi-tenant analytics database where different customer sessions share a process, that means side-channel analysis needs to happen before deployment, as a deliberate step, rather than something assumed to be covered by the sandbox already in place.

Cloudflare's response to this class of problem, freezing timer resolution so that Date.now() stays locked to the moment a request arrived instead of advancing during execution, is a real operational mitigation. It also only closes part of the gap, since timing is one channel among several an attacker could try to use.

The distinction between this and the compiler CVEs matters for how teams respond. A compiler bug gets fixed with a patch. A hardware side channel needs an architectural answer: separating processes, controlling what timers expose, or moving to a trusted execution environment. WASM running inside SGX enclaves is one such answer. WAVEN, presented at NDSS 2025, ran a Wasm-ported version of SQLite inside SGX enclaves and measured a geometric mean overhead of 10.42% on the PolyBench benchmark suite, a result that suggests the TEE route is realistic for database environments where the security bar is high enough to justify the added cost.

The MCP layer's attack surface before WASM is invoked

For agentic database access built on the Model Context Protocol, the attack surface doesn't start at the WASM sandbox. It starts at the MCP server that sits in front of it, and a sandbox wrapped around tool execution does nothing to stop an attack that lands before the sandbox is ever reached.

A documented vulnerability in an MCP client called mcp-remote makes the point concrete. A malicious server could pass an unsanitized authorization_endpoint URL through an OAuth flow, and that string went straight into the system shell. The WASM sandbox protecting tool execution never came into play, because the payload did its damage upstream of where that sandbox sits. One might argue this shows the sandbox failed. It shows something more specific: the sandbox was never positioned to catch this kind of attack.

Prompt injection, manipulated OAuth callbacks, and malformed MCP payloads all reach the database engine boundary before any WASM enforcement gets a chance to apply. An agent that's been prompt-injected into requesting a capability it shouldn't have defeats the sandbox at the moment that capability gets granted, not at runtime once the module is already executing.

Microsoft's Wassette, released in August 2025, builds a response around exactly this gap. The design draws a clear line between what an agent can ask for and what the host actually grants, putting the first real control point at the instantiation step.

Cosmonic Control and wasmCloud take a related approach from a different angle: a distributed Wasm execution model with sandboxed MCP server execution, paired with an observability stack running through OTLP, Prometheus, Loki, and Tempo. The answer to this layer is sandboxing paired with a full record of what ran and when.

What a defensible deployment posture looks like in practice

Put the pieces together into a workable deployment posture built on layered controls, with WASM's sandbox earning its place at the database boundary for UDFs, agentic tool calls, and third-party extensions, understood as one layer among several.

At the runtime level, that means staying current with Wasmtime patches as a standing practice, not a reaction. The April 2026 advisories showed that compiler bugs capable of full sandbox escapes can sit undiscovered for a long time, and the pace of discovery looks set to pick up now that LLM-based tooling is part of how these bugs get found. Teams running workloads on aarch64 infrastructure have a direct reason to track CVE-2026-34971 specifically.

At the hardware level, multi-tenant deployments need an honest answer about side-channel exposure before launch, not after. That might mean separating processes for different tenants, applying timer controls like one vendor's frozen-clock approach, or accepting the overhead of a TEE-based deployment like the SGX enclave approach WAVEN demonstrated, where the cost ran to a noticeable but modest share of performance in exchange for closing a gap no compiler patch can touch.

At the protocol level, the MCP layer is part of the attack surface in front of it. Wassette's deny-by-default model and Cosmonic's observability stack both point toward the same requirement: capability grants need to be explicit and auditable, and what actually ran needs to be logged and reviewable after the fact, not just contained while it's running.

None of these controls substitute for the others. Software fault isolation stops a module from reaching host memory it was never supposed to touch, but it says nothing about a compiler bug that breaks the bounds check itself. Capability gating stops ambient access, but it can't catch an attacker who gets a legitimate-looking capability granted through a prompt injection upstream. The strongest deployments treat each layer as covering what the others don't, building in the auditability to know, after the fact, what ran, what it touched, and what it was allowed to touch.

Sources

  1. WAVEN: WebAssembly Memory Virtualization for Enclaves Weili Wang∗
  2. Bytecode Alliance — Wasmtime's April 9, 2026 Security Advisories
  3. Swivel: Hardening WebAssembly against Spectre Shravan Narayan†

More in Foreign Function and Runtime Extension Interfaces