Three Reasons Our WebGL Engines Never Died — and Why Chrome Hid the Worst One


Klooboo, our game platform, is a feed of live 3D worlds you swipe through in the browser. Each game runs its own three.js engine. You swipe to a game, an engine starts. You tap Go, a world loads. You go back, it’s torn down.

After enough of that on an iPhone, the screen went white. That’s what iOS does when it kills a WebView’s content process for using too much memory: no crash dialog, no error. The page is just gone.

We’d been chasing this for a while. A previous fix capped a cache that turned out not to be holding the memory. Another was backwards and got reverted. We even built an instrument for it: __kbmem(), a console command reporting how many 3D engines were still alive, tracked through WeakRefs so a collected engine would drop out on its own. It said one runner. It kept saying one runner while phones died.

This is how it turned out that no engine had ever been freed, why our own meter couldn’t see it, and why the worst of it only happened on iOS.

Measure survivors, not intentions

The first thing I stopped trusting was the meter. It reported what each engine registered, and the runner unregistered itself on teardown, whether or not it was actually released. That’s a record of what we meant to free, not of what was freed.

So I built a probe that doesn’t ask the app anything. A Playwright init script wraps WebGL before the page loads and holds a WeakRef to every canvas that ever gets a context:

// injected before any app code runs
const canvases = [];
for (const P of [WebGLRenderingContext.prototype, WebGL2RenderingContext.prototype]) {
  const texImage2D = P.texImage2D;
  P.texImage2D = function (...args) {
    if (!this.__seen) { this.__seen = true; canvases.push(new WeakRef(this.canvas)); }
    return texImage2D.apply(this, args);
  };
}
window.__alive = () => canvases.filter((r) => r.deref()).length;

Then it drives the real app at phone size: switch games, enter a world, leave it. Before every reading it forces a full garbage collection over the Chrome DevTools Protocol, so anything still alive is actually held, not just waiting to be collected:

await cdp.send("HeapProfiler.collectGarbage");
const alive = await page.evaluate(() => window.__alive());

The first run settled it. Twelve game switches later, 9 engines were alive, one more every switch or two, and every one of them detached from the page. The right answer is one or two (the current game, plus the feed’s preview while it hands over), so everything above that is leak.

Then a heap snapshot, and a short script that walks retainer edges from each surviving WebGL2RenderingContext back to the window. Two things needed skipping to get true paths: weak edges, and the edges V8 labels “part of key → value”. Those are WeakMap entries. A snapshot draws them from the key even when the map itself is unreachable, so they look like retainers when they aren’t. My first pass followed one and pointed me at the wrong object.

With those skipped, three separate causes turned up.

Cause 1: the leak meter was a leak

The first path ran through requestAnimationFrame, into the registry behind __kbmem(), and out to a renderer. Here’s the helper that was supposed to hold each engine weakly:

function weak(probe: () => GlMemory) {
  if (!globalThis.WeakRef) return { deref: () => probe };  // fallback for old browsers
  const r = new WeakRef(probe);
  return { deref: () => r.deref() };
}

It looks right, but it isn’t. In V8, all closures created in one function invocation share a single scope object. The fallback arrow captures probe, so probe lives in that shared scope, and the WeakRef branch’s deref closure holds the same scope. The weak reference sits right next to a strong one. And probe reads renderer.info.memory, so it holds the renderer. Every engine the meter measured, it pinned.

The fix is to put each closure in its own function, so they don’t share a scope:

function weak(owner: object) {
  return globalThis.WeakRef ? weakRef(new WeakRef(owner)) : strongRef(owner);
}
function strongRef(owner: object) { return { deref: () => owner }; }
function weakRef(r: WeakRef<object>) { return { deref: () => r.deref() }; }

Fixing that exposed a second flaw. With a genuinely weak reference, live read 0 while an engine was on screen. The registry was the only thing holding the probe closure, so the collector took it straight away. The meter had never been tracking the engine at all, only a closure about the engine. It now holds the canvas weakly, which lives exactly as long as the engine, and keeps the probe in a WeakMap keyed by that canvas. A WeakMap value can reference its own key without keeping it alive.

That brought game switching down a little: 9 alive became 6. Most of the leak was still there.

Cause 2: the two geometries three.js never lets go of

The next retainer path ended at an object in module scope. Reading its properties out of the snapshot: type: "BufferGeometry", an index, an interleaved position + uv attribute, and a _listeners.dispose array with one entry per dead engine.

That’s THREE.Sprite’s geometry. three.js creates one quad at module scope and hands it to every sprite in the app.

Here’s the mechanism. When a renderer first draws a geometry, it registers for that geometry’s dispose event so it can free its GPU buffers later:

// three.js, WebGLGeometries
geometry.addEventListener("dispose", onGeometryDispose);

onGeometryDispose is a closure over that renderer’s internals. And renderer.dispose() doesn’t remove it. For an ordinary geometry that’s harmless: you dispose the geometry, the event fires, every renderer lets go. But a geometry nobody ever disposes keeps a listener, and through it a whole renderer, for every renderer that has ever drawn it. Our hero picker and one of our games use sprites, so every engine that drew one was pinned for the life of the page.

Disposing that quad on teardown (fix below) flattened game switching: 2 or 3 engines alive, however many switches. Then I measured the path I hadn’t yet: entering a world and leaving it. That was worse than anything so far. After 8 round trips, 17 canvases were alive and the JS heap had gone from 8 MB to 33 MB, still climbing.

The next snapshot found the sprite quad’s twin. Post-processing passes (bloom, colour grading, output) render through a FullScreenQuad, and Pass.js shares one module-level fullscreen triangle between all of them, in every EffectComposer on the page. Our runner never disposed its composer, so every runner that ever ran was held by that triangle. The feed’s stage did dispose its composer, but it could render one more frame afterwards, and that frame registered the listener again.

The fix is to dispose both shared geometries when an engine is torn down, after its last frame:

import { FullScreenQuad } from "three/examples/jsm/postprocessing/Pass.js";

export function releaseSharedThreeGeometry(): void {
  new THREE.Sprite().geometry.dispose();  // every Sprite shares this quad
  new FullScreenQuad().dispose();         // every pass shares this triangle
}

Constructing a throwaway Sprite or FullScreenQuad is simply the way to get at the shared instance. Disposing it fires the event, and every renderer, dead or alive, drops its listener and its buffer. A live engine re-uploads a few vertices on its next frame. That’s a buffer upload, not a shader compile, so it doesn’t bring back the freeze we’d fixed the day before. Alongside that, the runner now disposes its composer and passes, and the stage’s render() refuses to draw once it’s disposed.

In Chrome, that flattened everything: 17 canvases alive → 1, and the JS heap held at 12 MB.

So I put it on the iOS simulator.

Cause 3: iOS was waiting for a restore we’d asked for

On the simulator, after a normal play session and a forced collection:

__kbmem() → live: 24  (stage: 23, runner: 1)

Twenty-three dead game engines, still alive. Chrome had just said one.

Two things made this hard to read. First, Safari’s Web Inspector keeps more alive while it’s attached, so numbers taken with it open are suspect. Second, the simulator never runs short of memory, so WebKit rarely has a reason to collect. To take both out of the picture, I ran the same probe in Playwright’s WebKit build (26.5, the same engine generation as iOS) with no inspector attached. WebKit has no command to force a collection, so the probe allocates about a gigabyte of garbage before each reading instead. It reproduced cleanly: 11 engines alive after 5 enter/leave rounds, 7 after 6 game switches.

One thing in the teardown code stood out. Both renderers had this, to survive a real GPU loss:

canvas.addEventListener("webglcontextlost", (e) => e.preventDefault());

preventDefault() on a context loss tells the browser “I’ll handle this, restore my context when you can.” That’s the right thing for a GPU hiccup mid-game. But the handler was still attached when our own teardown did this:

renderer.dispose();
renderer.getContext().getExtension("WEBGL_lose_context").loseContext();  // three's forceContextLoss()

So every deliberate teardown also announced that a restore was coming. Chrome frees the canvas anyway once nothing references it. WebKit keeps it, context and all, waiting for a restore that never arrives.

Before changing any source, I checked that this one thing was the cause. The probe patched Event.prototype.preventDefault to do nothing for webglcontextlost events only, and ran the unchanged build again. 11 alive became 1. That’s a clean yes.

The fix is to stop asking for a restore just before losing the context on purpose:

const keepRestorable = (e: Event) => e.preventDefault();
canvas.addEventListener("webglcontextlost", keepRestorable);

// ...in dispose():
canvas.removeEventListener("webglcontextlost", keepRestorable);
renderer.forceContextLoss();

A real GPU loss during play still gets restored. Only our own deliberate loss stops asking.

The numbers

Three play sessions in the iOS simulator, each ending with a forced collection:

no fixescauses 1–2 fixedall three fixed
engines alive (__kbmem())meter couldn’t tell243 (35 created, 32 freed)
GPU process peak2,988 MB2,109 MB786 MB
page process~1.1 GB1,259 MB and climbing~575 MB, flat

The middle column is the one to look at: the Chrome fixes were real, but on iOS they weren’t enough.

The simulator can’t reproduce the kill itself; only a real phone runs out of memory. What it shows is the growth that leads there, and that growth is gone.

Takeaways

  • An instrument that trusts intent will agree with your bugs. Our meter counted engines that registered and hadn’t unregistered. The leak lived exactly in the gap between “I disposed it” and “it’s gone”. Count survivors directly, with WeakRefs the app knows nothing about, after a forced collection.
  • Closures in one function share a scope. A strong fallback written beside a WeakRef makes the weak one strong. If a closure must not capture something, give it its own function.
  • renderer.dispose() doesn’t undo what drawing did. Every geometry, texture and material a renderer touched keeps a listener pointing back at it. Anything that outlives the renderer (a cache, a shared asset, three’s own module-level geometries) keeps it alive.
  • Skip ephemeron edges when you read retainer paths. WeakMap “part of key → value” edges look like retainers in a heap snapshot and usually aren’t.
  • Test memory in the engine your users run. Every Chrome measurement looked clean, and iOS held 23 dead engines. Playwright ships a real WebKit; it’s a one-line install and it reproduced the iOS behaviour without a phone.
  • Isolate the variable before you fix it. Neutralising one browser behaviour from the test harness, with no source change, turned a plausible theory into a measured cause in one run.