Why Your WebGL Game Records as a Black Rectangle in Session Replay
Klooboo is a 3D game platform. You swipe through a feed of live worlds, tap one, and play it in the browser. Every pixel that matters is drawn by three.js into a <canvas>.
So when I set up PostHog session replay to find out how people actually play — do they find the controls? do they crash immediately? do they even understand what the game is? — I was hoping for the single most useful thing an analytics tool can give you: a video of a stranger struggling with your product.
What I got was the header, the “Level 1” chip, the back button, and between them a perfect, pristine black rectangle.
The chrome recorded flawlessly. The game — the only part worth watching — did not exist.
The first cause: rrweb doesn’t record canvas at all
PostHog’s replay is built on rrweb, which doesn’t record video. It serialises the DOM, then streams mutations, and the player reconstructs the page from that description. This is why replays are so small compared to screen recordings, and it’s also the root of the problem: a <canvas> element serialises as an element, not as pixels. The player faithfully reconstructs a canvas of exactly the right size, in exactly the right place, containing exactly nothing.
There’s a documented fix, and it’s one config block:
posthog.init(key, {
session_recording: {
captureCanvas: {
recordCanvas: true,
canvasFps: 4,
canvasQuality: "0.6",
},
},
});
We’d already hit this once before, on a different product — a kids’ colouring page, where the child paints into a 2D canvas. Turning on captureCanvas fixed that completely. So I shipped the same config to the 3D game, felt good about a quick win, and went to look at the next replay.
Still black.
The second cause: the frame is already gone
Here’s the part that isn’t in the getting-started docs.
rrweb snapshots canvases from its own requestAnimationFrame loop, which knows nothing about your game’s render loop. It wakes up on some frame, walks the DOM for canvas elements, and reads each one.
And a WebGL drawing buffer, by default, is cleared the moment it has been composited. This is spec behaviour, not a bug. Unless you explicitly ask the browser to preserve it, the contents are considered consumed once the frame is on screen — the browser is free to hand that memory straight back to the driver. It’s a meaningful optimisation, which is why it’s the default.
So the sequence is:
- Your game renders a frame.
- The browser composites it. The drawing buffer is now empty.
- rrweb’s rAF fires, reads the canvas, and gets black.
Every snapshot, forever. Canvas capture was working perfectly. It was faithfully recording nothing, four times a second.
The fix is one WebGL context attribute:
new THREE.WebGLRenderer({ preserveDrawingBuffer: true });
The part that made me feel silly
While writing this up, I went looking through our own repo and found this comment, in a screenshot-testing script I’d written months earlier:
Read inside a double rAF, not a bare
evaluate(): the games’ WebGLRenderer uses the defaultpreserveDrawingBuffer:false, so the browser clears the drawing buffer right after compositing. A plainpage.evaluate()is a separate round-trip outside the game’s own rAF loop — it lands after that clear and reads back an all-black buffer even when the game is rendering fine (verified: a bare read measured mean 0 against a visibly correct screenshot; a double-rAF read of the same frame measured mean ~28).
And in a video-capture script, the same thing again:
drawImage()on a WebGL canvas returns an empty buffer unless the context was created withpreserveDrawingBuffer:true— the drawing buffer is cleared after compositing.
I had diagnosed this exact failure mode twice, written it down twice, and still didn’t recognise it when it came back wearing a different hat. The lesson I’m taking isn’t “write more comments” — the comments existed and were excellent. It’s that a symptom shows up in a new subsystem and your brain files it as a new problem. “Session replay is broken” and “screenshot tool reads black” don’t feel related. They’re the same sentence.
The race nobody mentions
Here’s where it gets genuinely subtle, and where I nearly shipped a fix that only half worked.
posthog-js knows about preserveDrawingBuffer. If you read its recorder bundle, you’ll find it monkey-patching the canvas API on startup:
// posthog-js, paraphrased from the minified recorder
patch(HTMLCanvasElement.prototype, "getContext", (original) => function (type, ...args) {
if (["webgl", "webgl2"].includes(type)) {
if (args[0] && typeof args[0] === "object") {
args[0].preserveDrawingBuffer ||= true;
} else {
args.splice(0, 1, { preserveDrawingBuffer: true });
}
}
return original.apply(this, [type, ...args]);
});
Any WebGL context you create gets the flag forced on, whether you asked for it or not. Which sounds like it should make my fix unnecessary.
Except: that patch only exists from the moment the recorder runs. posthog-js lazy-loads its recorder as a separate script over the network. A context created before that script lands never goes through the patched getContext, and keeps the default.
Our feed stands up its preview engine on a 200ms timer after mount. A network fetch loses that race essentially always. So the feed — the first thing every single visitor sees — would have kept replaying black, while the in-game canvas (created seconds later, after a tap) would have recorded fine. A fix that works everywhere except the landing screen is arguably worse than no fix, because you’d stop looking.
There’s also a fallback in rrweb for contexts it didn’t patch:
if (context?.getContextAttributes()?.preserveDrawingBuffer === false) {
context.clear(context.COLOR_BUFFER_BIT);
}
Reading that as “we handle it” is how you lose an afternoon. It’s a hack to make the canvas readable at all, and what you read is a cleared buffer. Black.
Proving it, and the probe that lied
I wanted evidence before touching a renderer that had been carefully tuned for weak Android GPUs. So I wrote a probe: two WebGL contexts, one with preserveDrawingBuffer and one without, draw bright orange into both on one frame, read both back on a later frame exactly the way rrweb does.
In headless Chrome, both returned orange. No difference at all.
If I’d trusted that, I’d have concluded the flag was unnecessary and shipped the half-fix. Headless Chrome renders through SwiftShader, a software rasteriser with no real compositor — there’s no present step, so nothing ever clears. The probe couldn’t reproduce the bug because the environment didn’t have the mechanism that causes it.
The honest version of that result isn’t “the flag doesn’t matter.” It’s “this test can’t see the thing I’m testing for.”
What finally produced clean evidence was the real app in a real browser. Same page, same read path, one flag different:
| canvas luminance | |
|---|---|
without preserveDrawingBuffer | 0 (pure black) |
with preserveDrawingBuffer | 182.6 (rendering) |
That’s exactly what rrweb sees.
Don’t make everyone pay for it
preserveDrawingBuffer: true isn’t free. It stops the driver discarding the framebuffer after present, which on the tile-based GPUs in cheap Android phones means an extra full-screen copy every frame. These renderers are tuned around devices that already lose their WebGL context under load — we count those losses in analytics precisely because they’re a real failure mode.
So we don’t set it globally. We set it only for visitors who are actually being recorded:
export function canvasOptionsFor(replay: boolean): { preserveDrawingBuffer?: true } {
return replay ? { preserveDrawingBuffer: true } : {};
}
// at every renderer:
new THREE.WebGLRenderer({ antialias: false, ...replayCanvasOptions() });
The neat part is that this costs recorded visitors nothing extra. Once canvas capture is on, posthog’s own patch forces the same flag on them anyway — we’re just closing the window before its recorder arrives. Everyone else keeps the cheaper default.
What it actually costs in bytes
“Record the 3D” sounds expensive, so I measured it against the live game at a phone-sized viewport — 390×529 CSS pixels, which is the size rrweb downscales to before encoding.
| quality | per frame | at 4 fps |
|---|---|---|
| 0.4 | 17.5 KB | 4.1 MB/min |
| 0.6 | 23.6 KB | 5.5 MB/min |
| 0.8 | 35.4 KB | 8.3 MB/min |
Roughly a 480p YouTube stream, out of your player’s mobile data.
The framerate is the linear knob; quality barely moves the needle (0.6 → 0.4 saves only ~26%). Two things also make it cheaper than it sounds: the JPEG encode runs in a Web Worker with the bitmap transferred rather than copied, and rrweb skips a canvas whose previous snapshot is still in flight, so a device too slow to keep up drops frames instead of stalling your game loop. Frame rate stayed pegged at 60 with the snapshot loop running.
One warning if you’re extrapolating from a 2D canvas: rrweb’s worker drops any frame identical to the previous one. An idle drawing costs nothing. A 3D game never repeats a frame, so you pay for every single snapshot. The two cases are not comparable.
I shipped 2 fps first, because that was the arithmetic answer. Then I watched a replay back and it was too choppy to follow a run. We’re at 4. Measurement was an input to that decision, not the decision.
Three things I didn’t expect
Images are recorded by URL, not by value. rrweb embeds canvas as pixels but records <img> as a reference. Recording from a local dev server, every thumbnail was blank — the player runs on https:// and won’t load an http:// asset. I nearly debugged a bug that only exists on localhost.
You can’t verify this with an automated browser. posthog-js bot-detects Playwright and silently drops all capture. The console says [SessionRecording] starting and then not one byte is sent. Client-side state is inspectable that way; ingestion isn’t. A human has to click it.
We couldn’t watch our own sessions at all. We filter out internal traffic by timezone and localhost — which covers every device the team owns, so there was no way to check any of this. The documented workaround was faking your location in Chrome DevTools, which is impossible on a phone. We added ?replay=1: stored in localStorage, survives reloads, works against production on a real device. A forced session skips our marketing analytics entirely and tags itself so test runs stay identifiable.
If you build anything canvas-heavy, that last one is worth doing on day one. A verification story you can’t run on the device you’re holding isn’t a verification story.
The short version
If your WebGL app replays as a black box:
- Turn on canvas capture. rrweb records
<canvas>as an empty element otherwise. - Set
preserveDrawingBuffer: trueon the renderer. The buffer is cleared at composite, and rrweb reads on a later frame. - Mind the ordering. posthog patches
getContextto force the flag, but only for contexts created after its recorder downloads. Anything you stand up during startup loses that race. - Gate it. Only recorded visitors need the flag; it costs a buffer copy per frame on weak GPUs.
- Check your probe can see the bug. Headless Chrome has no compositor, so it cannot reproduce this one — and it’ll tell you everything is fine.