---
license: cc-by-nc-4.0
library_name: gaussianformer
tags:
- 3d-gaussian-splatting
- neural-rendering
- novel-view-synthesis
- transformer
- pytorch_model_hub_mixin
- model_hub_mixin
- gaussianformer
pipeline_tag: image-to-image
datasets:
- ShapeSplats/Objaverse_Splats
base_model:
- microsoft/renderformer-v1-base
---
# GaussianFormer
GaussianFormer renders 3D Gaussian Splatting scenes with a transformer, without per-scene optimization. It adapts
[RenderFormer](https://huggingface.co/microsoft/renderformer-v1-base) (SIGGRAPH 2025), a transformer renderer for
triangle meshes: each Gaussian's 14 parameters `[pos(3), scale(3), quat(4), rgb(3), opacity(1)]` become one scene
token, and the two-stage architecture (view-independent scene encoder, view-dependent ray decoder with a DPT head)
is kept and warm-started from RenderFormer's weights.
*Eight objects never seen in training: ground truth (full splat) and GaussianFormer, PSNR per view.*
*Two held-out objects, rasterized (left) and rendered by GaussianFormer (right): a camera orbit and a tumbling object.*
**Code:** [github.com/SVLwoof/gaussianformer](https://github.com/SVLwoof/gaussianformer)
## How to use
Install the [code](https://github.com/SVLwoof/gaussianformer) with [uv](https://docs.astral.sh/uv/) (Linux, NVIDIA GPU,
Python 3.12); this also brings PyTorch for CUDA 12.8 and a prebuilt Flash Attention wheel:
```bash
uv add "gaussianformer[flash] @ git+https://github.com/SVLwoof/gaussianformer" # or: uv pip install "..."
```
With pip, install Flash Attention separately (pip does not read uv's sources):
```bash
pip install "gaussianformer @ git+https://github.com/SVLwoof/gaussianformer"
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.8.3/flash_attn-2.8.3+cu12torch2.9cxx11abiTRUE-cp312-cp312-linux_x86_64.whl
```
Then, from a 3D Gaussian Splatting PLY:
```python
import torch
from gaussianformer import GaussianFormerRenderingPipeline, load_ply
from gaussianformer.utils.cameras import orbit
pipeline = GaussianFormerRenderingPipeline.from_pretrained("shahafvl/gaussianformer").to("cuda")
gaussians = load_ply("object.ply", up="z") # [N, 14]; up = the file's up axis
c2w = torch.from_numpy(orbit(8, 1.7)) # [V, 4, 4] camera-to-world (-Z forward, +Y up)
images = pipeline(gaussians, c2w, fov=45.0) # [V, 512, 512, 3]
```
The model takes each Gaussian as 14 activated values, pos(3) | scale(3) | quat wxyz(4) | rgb(3) | opacity(1): linear
scales, a unit quaternion, colors and opacity in [0, 1], for an object centred, Y-up and scaled into [-0.45, 0.45]^3,
about 20k Gaussians. A PLY stores log scales, raw quaternions, spherical-harmonic coefficients and opacity logits;
`load_ply` converts them, normalises the object, prunes it to 20k Gaussians and fine-tunes the kept ones
against the full splat (about a minute on a GPU).
From the command line, in a clone of the repository (`uv sync`):
```bash
uv run python -m tools.ply_to_h5 --ply object.ply --out object.h5 --up z
uv run python infer.py --model shahafvl/gaussianformer --h5 object.h5 --out renders/ # one PNG per camera
```
## Results
300 objects held out from training, 4 views each (views 0, 4, 7 and 11 at camera distance 1.7), PSNR on the
object's bounding box against the full-splat ground truth:
| | PSNR (dB) | LPIPS |
|---|---|---|
| gsplat rasterization of the same 20k-Gaussian input | 44.97 | reference |
| **GaussianFormer** | **37.08** | +0.009 over rasterization |
- PSNR quantiles over the 300 objects (10 / 50 / 90 %): 33.3 / 37.5 / 40.9 dB.
- Closer in (distance 1.15, the same four angles): 36.87 dB, against 42.24 dB for rasterization.
- On 300 training objects the same metric is 37.19 dB, so seen and unseen objects differ by 0.07 dB.
*A held-out object up close: ground truth, rasterization of the 20k-Gaussian input, GaussianFormer, and its absolute
error.*
## Changes over RenderFormer
Besides the Gaussian input tokens and the training data:
- **Projected 2-D RoPE** (`proj_rope_2d`): each Gaussian is projected into the view, and its image coordinates enter
the view decoder's cross-attention as a 2-D rotary embedding matched against each patch's position.
- **Windowed cross-attention** (`xattn_window=8`): each tile of 8x8 patches attends only to the Gaussians whose
projected footprint reaches it. Needs Flash Attention on GPU.
- **4 px patches** (`patch_size=4`, `ray_embed_patch=8`): RenderFormer uses 8 px patches; 4 px gives a 128x128
token grid at 512 px, affordable because of the windowing. The pretrained 8 px ray embedding is reused.
## Training
- **Data:** all 26,820 training objects of
[Objaverse_Splats](https://huggingface.co/datasets/ShapeSplats/Objaverse_Splats), each pruned
from ~50k to 20k Gaussians (LightGaussian importance score, then a short recovery fine-tune of the kept Gaussians).
Targets are gsplat renders of the full, unpruned splat at 512 px: 28 views per object at camera distances 1.15,
1.7 and 2.45.
- **Schedule:** initialised from RenderFormer's weights and trained at 256 px, then 385k steps at 512 px on
8 GPUs (one view per GPU per step; 4 passes over every view): a warmup-stable-decay learning rate at 5e-5, with a
cosine decay to 5e-7 over the last 10k steps. About 4.5 days on 8 L40S.
- **Loss:** log-HDR L1 (background pixels weighted 0.05) and LPIPS-VGG, 0.5 each.
- The 256 px checkpoint the last two stages start from is published as
[gaussianformer-256px](https://huggingface.co/shahafvl/gaussianformer-256px); `scripts/train.sh` in the code repository
runs all stages.
## Limitations
- Trained on isolated single objects centred in [-0.45, 0.45]^3, 45° FOV, camera distances 1.15 to 2.45. The
held-out objects come from the same distribution; multi-object scenes and real captures are not evaluated.
- Colors are view-independent (only the DC color of each Gaussian is used).
- Inputs of about 20k Gaussians; scene-encoder attention is quadratic in the number of Gaussians.
- Fine detail (thin strokes, engravings, small text) is softer than rasterizing the same splat.
## License
Trained on Objaverse_Splats, a subset of [Objaverse](https://objaverse.allenai.org/) whose terms restrict commercial
use; this checkpoint is CC-BY-NC-4.0, research and non-commercial use only.
## Citation
A paper on GaussianFormer is in preparation and will be presented at a later date; its citation will be added here.
GaussianFormer builds on RenderFormer; please also cite:
```bibtex
@inproceedings{zeng2025renderformer,
title = {RenderFormer: Transformer-based Neural Rendering of Triangle Meshes with Global Illumination},
author = {Chong Zeng and Yue Dong and Pieter Peers and Hongzhi Wu and Xin Tong},
booktitle = {ACM SIGGRAPH 2025 Conference Papers},
year = {2025}
}
```