Applied Intuition · Research

S2Tok: Streaming 3D Gaussian reconstruction with persistent spatial tokens

1Applied Intuition 2University of Illinois Urbana-Champaign 3Purdue University 4University of California, Berkeley

*Work done during an internship at Applied Intuition †Corresponding author

Code (on release) Paper →

Introduction

A narrated walkthrough of S2Tok

Turn on the sound to hear the narration of the video.

Method

The scene memory lives inside the backbone

S2Tok pipeline: encoder, stream-spatial transformer with scene tokens, admission, multi-scale decoder and Gaussian head
1

Tokenized streaming reconstruction

The scene is a set of latent spatial tokens. After every frame they are decoded into 3D Gaussians: a reconstruction of everything seen so far, at every step.

2

One image per step, no KV cache

Each step sees only the current frame. Past frames are not kept or re-attended; the latent spatial tokens are the only state carried to the next step.

3

Adaptive token length

The number of latent spatial tokens is not fixed: each step updates the existing tokens and admits only the selected tokens of the new frame.

Results

One pass over the video, a scene after every frame

Mip-NeRF 360 · Room Mip-NeRF 360 · Counter Mip-NeRF 360 · Bonsai Mip-NeRF 360 · Garden DL3DV · 001 DL3DV · 005 DL3DV · 006 DL3DV · 007 RealEstate10K · 000 RealEstate10K · 002 RealEstate10K · 003 RealEstate10K · 004 RealEstate10K · 007 ScanNet++ · 01 ScanNet++ · 03 Castle Knight Courthouse Living Room Mid Century

Interactive

Explore the reconstructed scenes

Left of the divider: each color marks the Gaussians decoded from one latent spatial token (32 per token). Right: the same Gaussians in their RGB colors, i.e. the reconstructed scene.

Drag the divider to compare · drag the scene to orbit · right-drag or shift-drag to pan · click, then scroll to zoom · double-click to reset