How we render HTML and CSS as 3D in WebGL
Every browser already ships one of the most sophisticated layout engines ever written. It resolves the cascade, media queries, flexbox, grid, fonts and line breaks, and it has been tuned for decades. When we set out to put web pages into 3D, the obvious move was to build a 3D UI toolkit. We decided not to. Fold.page lets the browser keep doing all of that work, then reads the result and redraws the page as an interactive WebGL scene.
This post walks through how that works, the parts that turned out to be hard, and what still isn't solved. If you're new to the idea of web pages in 3D, what is the spatial web? is the background.
The pipeline
When Fold.page starts on a page, it runs the same six steps every time:
- Target selection: find the elements marked with the
foldpageclass (or the whole body). - Discovery: collect the elements inside them that Fold knows how to draw: text, images, backgrounds, borders, form controls, video, SVG, and anything carrying a
data-foldpage-*attribute. - Style extraction: read each element's computed styles and its box on the page.
- DOM hiding: make the original elements invisible and turn off their pointer events. They stay in the document.
- 3D rendering: build a matching Three.js object for each element.
- Tracking: follow scroll, resize and style changes, and update the scene.
The rest of this post is mostly about steps 3, 5 and 6, and about what step 4 makes possible.
Let the browser do layout
Fold.page never computes layout itself. For every element it asks the browser for two things: getBoundingClientRect() for where the element sits, and getComputedStyle() for how it looks. Pixel positions become world positions through a single scale factor (0.0007 world units per CSS pixel), so a page that is 1,280 pixels wide is about 0.9 units wide in the scene.
That decision is what makes responsive design work in 3D for free. Media queries, clamp(), flex wrapping, grid tracks and container sizes are all resolved before Fold sees anything. Resize the window and the browser re-lays out the page; Fold reads the new boxes and moves the meshes.
It also shapes how the reads are cached. A full scan calls getComputedStyle() from dozens of places, so every read goes through one small layer. The detail that makes caching safe is in the spec: getComputedStyle() returns a live CSSStyleDeclaration, so a cached object still reports current values after classes change. getBoundingClientRect() returns a frozen DOMRect, so rects can't be cached the same way and are re-read on resize and layout changes.
// One cached CSSStyleDeclaration per element (and pseudo-element).
// Safe across renders: the object is live, so reads always see current styles.
const cache = new Map<Element, CSSStyleDeclaration>();
function computedStyle(el: Element) {
let style = cache.get(el);
if (!style) cache.set(el, (style = getComputedStyle(el)));
return style;
}
Text is the hard part
Boxes are easy to draw. Text is not. Fold renders text with Troika, which draws signed-distance-field glyphs that stay sharp at any distance, which matters a lot when a reader can lean in towards a headset panel.
A paragraph in HTML is rarely one style. It has bold words, links and inline code in it. Drawing each run as its own mesh is slow and makes line-wrapping impossible to match, so each block of text is one mesh with style ranges: bold, italic and colour runs inside a single string. Troika didn't support that, so we maintain a fork that adds it. Underlines are drawn from Troika's selection rectangles for the underlined range, and text-shadow is parsed into extra text meshes behind the main one.
Links inside text are hit-tested against the bounding boxes of their character range. When you click one in 3D, Fold calls click() on the real <a> element, which matters later.
What isn't solved: the browser and Troika measure text separately. With the same font and width they almost always agree on where a line breaks, but not quite always, and a near-miss shows up as one line too few or too many. Matching the browser's line breaking exactly is still open work.
Boxes, borders and stacking
Backgrounds, borders, border-radius and box-shadow become meshes behind each element's content. z-index becomes a small depth offset within each stacking group, so overlapping elements stack the way they do in the browser. Keeping that true once elements are rotated in 3D, or pinned with position: fixed, is where many of the interesting bugs live.
CSS transition carries over too. When a class change alters an element's colour, opacity, shadow or radius, Fold reads the transition timing from the computed style and tweens the 3D material instead of snapping.
overflow: hidden in 3D
In the browser, overflow: hidden clips a box's children to its padding box. In a 3D scene there is no "box" to clip against, and the children might have depth of their own.
Earlier versions used stencil-buffer masks. The current version clips in the fragment shader instead. Each overflow: hidden element defines a rounded-box clip volume, and every fragment inside it is tested against the stack of volumes from its ancestors, which is how nested overflow behaves in the browser too.
Two trade-offs came out of that:
- GLSL arrays have a fixed length, so the shader keeps the four nearest clipping ancestors. Deeper nesting drops the outermost ones.
- Clipping in X and Y is what authors mean; clipping in Z almost never is. Volumes get a generous margin in depth, so an element with
data-foldpage-depthisn't sliced in half by its card.
Making :hover work without hovering
Step 4 hides the real elements and turns off their pointer events, so the browser never sees the mouse over them, and :hover never matches. But every site's design depends on :hover.
The fix is to mirror the stylesheets. Fold walks every CSS rule, including the ones inside @media, @supports and @layer, and for each rule whose selector contains :hover (or :active) it writes a copy where the pseudo-class is replaced by a class:
/* Your stylesheet */
@layer components {
.btn:hover { background: var(--accent-2); }
}
/* Fold's mirror, wrapped in the same at-rules */
@layer components {
.btn.foldpage-hover { background: var(--accent-2); }
}
When a ray from the pointer hits an element in 3D, Fold adds the class to the real element, re-reads its computed style and animates the change. Ancestors get numbered classes too, so .card:hover .title still works when the pointer is over the title.
Wrapping each mirror in the same at-rules is what keeps the cascade intact. A reset stylesheet's a:hover inside a low-priority layer must still lose to your own unlayered CSS, exactly as the original does.
One Chrome quirk to know about: rule.style.cssText comes back empty when a declaration uses a shorthand with an unresolved var(), because the browser can't serialise the longhands. The fix is to fall back to slicing the declaration block out of rule.cssText. Single-page apps add and remove stylesheets as you navigate, so the mirror is rebuilt whenever the set of stylesheets changes.
Clicks, focus and forms still go to the real page
Because the DOM is still there, Fold doesn't have to reimplement interaction. It forwards it.
- Clicks call
click()on the real element, so framework handlers, router links, analytics and form submits all run exactly as before. - Keyboard focus moves through the real DOM in its normal tab order. Fold draws the focus ring in 3D from the element's computed
outline, and when a page is laid out as 3D views, focusing an element moves the camera to it. - Form controls stay native where the platform needs them to. On iOS, tapping a 3D
<select>callsfocus()on the real one inside the tap gesture, which opens the system picker.
The same property is what keeps a Fold page accessible and indexable. Screen readers, keyboard navigation and search engines all work with the real document, not the canvas.
Rendering only when something changes
A full-page WebGL scene that redraws 60 times a second would drain a phone battery while you read. Fold renders on demand: the render loop runs only while something is changing. React updates and text layout request their own frames; everything Fold drives itself (scroll easing, CSS transition tweens, shader time, camera moves) asks for the next frame from inside the code that advances it, and stops asking when it's done. A static page costs no CPU or GPU between interactions.
Scrolling is cheap for a related reason. Fold doesn't move the page through the scene; it moves the camera. The content's world matrices stay frozen, and only subtrees that actually changed (an animating element, an element mid-transition) are recomputed, so a page with a hundred animated elements doesn't pay for a full-tree update every frame.
On slower devices Fold also measures frame times as it runs and steps down through quality tiers, dropping the most expensive effects (shaders, animations, repeated depth layers) before the page starts to stutter.
The same page in WebXR
On a headset, the same scene enters an immersive session through the WebXR Device API. There's no scrollbar in VR, so the controller thumbstick (or a finger drag in phone AR) scrolls instead, and stops at the top and bottom of the page the way a browser does. The page re-centres in front of you when you move or turn, so it stays at a comfortable reading distance.
What doesn't change is the content. There's no separate XR build and no second URL: the page you shared is the page that opens in the headset. The XR features docs cover the details.
What it doesn't do (yet)
A few honest limits, because they're the first questions people ask:
- Not every CSS property maps to 3D. The common ones do; the CSS properties reference lists exactly which.
- Text wrapping can differ from the browser on rare lines, as described above.
- Very large pages cost more to scan and draw. The quality tiers help, but a page with thousands of elements is still heavier than the same page flat.
- It's progressive enhancement, not a requirement. Without WebGL, or without JavaScript, visitors get the normal HTML page. Nothing breaks, but nothing is 3D either.
Try it
The input is the HTML and CSS you already write:
<main class="foldpage">
<h1 data-foldpage-depth="10">Hello, spatial web</h1>
<p>Regular HTML and CSS, rendered in 3D.</p>
</main>
// npm install fold-page
import Fold from 'fold-page';
Fold.init();
The live demos show each feature on real pages, and the getting started guide covers installation for plain HTML, Angular, React and Vue. If you build something with it, or find a page it renders wrong, we'd like to hear about it on Discord.