Applied Intuition · Research
S2Tok: Streaming 3D Gaussian reconstruction with persistent spatial tokens
Introduction
A narrated walkthrough of S2Tok
Turn on the sound to hear the narration of the video.
Method
The scene memory lives inside the backbone
Tokenized streaming reconstruction
The scene is a set of latent spatial tokens. After every frame they are decoded into 3D Gaussians: a reconstruction of everything seen so far, at every step.
One image per step, no KV cache
Each step sees only the current frame. Past frames are not kept or re-attended; the latent spatial tokens are the only state carried to the next step.
Adaptive token length
The number of latent spatial tokens is not fixed: each step updates the existing tokens and admits only the selected tokens of the new frame.
Interactive
Explore the reconstructed scenes
Left of the divider: each color marks the Gaussians decoded from one latent spatial token (32 per token). Right: the same Gaussians in their RGB colors, i.e. the reconstructed scene.
Drag the divider to compare · drag the scene to orbit · right-drag or shift-drag to pan · click, then scroll to zoom · double-click to reset



















