S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens

arXiv Project Page Code License

S2Tok reconstructs a scene as 3D Gaussians from an uncalibrated image stream, one frame at a time. It keeps a persistent, size-adaptive set of spatial tokens as its scene state: each new frame updates the existing tokens, and a learned admission module adds new tokens only for content that is not yet represented. A hierarchical decoder turns the tokens into non-pixel-aligned Gaussians, so the scene can be rendered from new views at any point of the stream without keeping past frames.

⚠️ Research release. This model is released for research purposes under the terms below. It is not a production system and does not include any Applied Intuition proprietary data, product code, or production checkpoints.

Model Details

Developed by Applied Intuition — AI Research
Model type Streaming feed-forward transformer (DA3-Giant ViT-G/14 backbone) with a persistent token memory, a learned admission module and a 3D Gaussian decoder
Paper S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens
Code https://github.com/Applied-Intuition-Open-Source/S2Tok
Project page https://s2tok.github.io/
Initialized from ZipSplat (CC BY-NC 4.0), itself initialized from DA3-Giant (CC BY-NC 4.0)
Training data DL3DV-10K and RealEstate10K
License (weights) CC BY-NC 4.0 — see WEIGHTS_LICENSE.md
Contact GitHub Issues on the code repo

Checkpoints

Checkpoint Description Size Link
s2tok_252.pt 1.72B params. 252x252 input, trained on 12–24 frame sequences from RealEstate10K and DL3DV-10K. 6.87 GB s2tok_252.pt

The file holds model (inference weights) and config (architecture settings), and loads with torch.load(..., weights_only=True). SHA-256 checksums are in SHA256SUMS.

Intended Use & Limitations

Intended use: research on streaming and feed-forward 3D reconstruction and novel-view synthesis.

Out of scope / limitations:

  • Frames are center-cropped and resized to 252x252. Gaussians are expressed in the camera frame of the first image, up to a global scale.
  • The memory holds at most 4,000 tokens (128k Gaussians); once it is full, no new content is admitted.
  • Trained on 12–24 frame sequences of static scenes and evaluated at 12 and 50 views; retention over much longer streams is not established.

Commercial use is not permitted under CC BY-NC 4.0. The weights are also subject to the terms of their upstream model and datasets (see License).

How to Use

from huggingface_hub import hf_hub_download
from s2tok import load_model

path = hf_hub_download("AppliedIntuitionResearch/S2Tok", "s2tok_252.pt")
model = load_model(path, device="cuda")

memory = model.init_memory()
for image in frames:  # [3, 252, 252] RGB tensors in [0, 1]
    out, memory = model.step(image.cuda(), memory)
out.gaussians.save_ply("scene.ply")  # Gaussians after the last frame (3DGS PLY)

Streaming demo with an interactive viewer:

python demo.py --model_path s2tok_252.pt --seq_path path/to/video.mp4 --output_dir outputs/my_scene

Full installation and usage instructions: see the GitHub repository.

License

The weights are released under CC BY-NC 4.0, non-commercial use only (WEIGHTS_LICENSE.md, full text in LICENSE). The non-commercial restriction is also required by their upstream sources:

  • fine-tuned from the released ZipSplat checkpoint (CC BY-NC 4.0), which is initialized from DA3-Giant (CC BY-NC 4.0); depth labels for training were produced with DA3-Nested-Giant-Large (CC BY-NC 4.0);
  • trained on DL3DV-10K (DL3DV-10K Terms of Use and CC BY-NC 4.0: non-commercial research and education only) and RealEstate10K (Google; frames from YouTube videos).

Use of the weights must also comply with the terms of these models and datasets. No training data is included in this repository.

Citation

@article{li2026s2tok,
  title   = {S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens},
  author  = {Li, Fang and Yenphraphai, Jiraphon and Herau, Quentin and Meng, Depu and Hu, Yihan and Xu, Tianshuo and Ahuja, Narendra and Zhan, Wei},
  journal = {arXiv preprint arXiv:2610.08978},
  year    = {2026}
}

Acknowledgments

This model builds on ZipSplat (code Apache-2.0, weights CC BY-NC 4.0), Depth Anything 3 and DINOv2 (Apache-2.0). It is trained on DL3DV-10K and RealEstate10K. We thank the authors for making their work available.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AppliedIntuitionResearch/S2Tok

Base model

veichta/zipsplat
Finetuned
(2)
this model

Paper for AppliedIntuitionResearch/S2Tok