GrapeSplat reconstructs a renderable 3D Gaussian scene from unposed, uncalibrated images in one forward pass. Gaussians live on a scene-level voxel grid instead of on pixels, so the grid sets the primitive count rather than the image resolution and the view count.
NOTICE.md lists the third-party components and what we changed in them.
Python 3.12, CUDA 12.8, uv.
git clone --recurse-submodules https://github.com/VAISR/GrapeSplat.git
cd GrapeSplat
MAX_JOBS=1 uv syncKeep MAX_JOBS=1. Several packages under gitmodules/ build CUDA extensions on the first sync, and compiling them in parallel exhausts host memory. Raising it is at your own risk.
Configs reach large assets through project-relative entries you create. None of them is tracked.
.shared/projectsholds pretrained weights in Hugging Face{owner}/{repo}layout.shared/datasetsholds datasets, one directory each.modelholds checkpoints loaded byckpt=.logsholds training logs.dataholds cached split files
One link covers the first two: ln -s <storage root> .shared.
Git LFS is the fastest transport, so skip the contents on clone and pull them after:
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/{owner}/{repo} .shared/projects/{owner}/{repo}
cd .shared/projects/{owner}/{repo} && git lfs pullfacebook/VGGT-1Bis the geometry backbonefacebook/dinov3-vith16plus-pretrain-lvd1689mis the semantic backbone, used by the ablation onlylhjiang/anysplatis a baselinedepth-anything/DA3-GIANT-1.1is a baselineJeasco/SplatWeaveris a baselineDazzlingSun/structsplatis a baselineHwasikJeong/2Xplatis a baseline, file2xplat_dl3dv_hr.pt
Both facebook repositories gate on accepting their license first. Through the Hugging Face CLI instead, disable Xet or each connection caps near 1 MB/s:
HF_HUB_DISABLE_XET=1 HF_HUB_ENABLE_HF_TRANSFER=1 hf download {owner}/{repo}Trained GrapeSplat checkpoints live at asherchen/grapesplat, one .ckpt per config name under config/grapesplat/module/pipeline/. Download by the LFS recipe above, then point a run at one:
uv run run_once test module/pipeline=our_gauss_clustered ckpt=.model/our_gauss_clustered.ckptour_gauss_clustered.ckptis the main paper modelour_voxel_affine.ckptandour_voxel_peach.ckptare the ablation ladder below itour_sem_encoded.ckptandour_black_bgcolored.ckptare the supplementary variantsdefault.ckptcarries the initialization state for training from scratch
Place each under .shared/datasets/ with the name below; config/grapesplat/data/*.yaml expects them.
ARKitScenescomes from the snippet belowHypersimcomes fromuv run gitmodules/hypersim/code/python/tools/dataset_download_images.pyScanNetppcomes from https://kaldir.vc.in.tum.de/scannetpp/TartanAirV2comes fromuv run gitmodules/tartanair_v2/examples/download_dips_example.py, withDATA_ROOTin it pointed at your copyWildRGBDcomes fromuv run gitmodules/wildrgbd/download.py --cat allNRGBDcomes fromwget https://kaldir.vc.in.tum.de/neural_rgbd/neural_rgbd_data.zipSevenScenescomes from the 7-Scenes release, prepared with the Spann3R preprocessing scriptDTUcomes frombash src/script/down_dtu.shETH3Dcomes frombash src/script/down_eth3d.shDL3DV-Benchmarkcomes from https://huggingface.co/datasets/DL3DV/DL3DV-Benchmark by the LFS recipe above;benchmark-meta.csvlists the 140 test scenes
ARKitScenes reads six assets across both splits, and its metadata needs the corrected copy from this repository:
for split in Training Validation; do
uv run gitmodules/arkitscenes/download_data.py raw \
--split $split \
--download_dir .shared/datasets/ARKitScenes \
--raw_dataset_assets lowres_wide lowres_wide.traj lowres_depth confidence vga_wide vga_wide_intrinsics
done
cp doc/data/arkitscenes/metadata.csv .shared/datasets/ARKitScenes/raw/metadata.csvThat file is the official metadata from https://docs-assets.developer.apple.com/ml-research/datasets/arkitscenes/v1/raw/metadata.csv with sky_direction corrected for 852 of 5071 scenes, by the classifier in src/script/patch_arkitscenes_metadata.py followed by manual review. Our splits do not reproduce without it.
TartanAirV2 downloads only the front-left camera; the example script already sets that.
A Hydra config name, then overrides.
uv run run_once train module/pipeline={variant}Swap train for test to evaluate. {variant} is one of our_voxel_affine, our_voxel_peach, our_gauss_clustered, or our_sem_encoded for the ablation chain, or a ref_* entry to run a baseline in the same harness. data= picks the benchmark and +view@data=cXnY takes X context views out of Y frames.
Metrics over a finished sweep, then the tables and figures:
uv run run_eval
uv run run_eval_logrun_once submits through SLURM. Other schedulers need the srun call at the bottom of src/script/run_once.py adapted.
@misc{grapesplat2026,
title={GrapeSplat: Geometry-Grounded Reconstruction via Amalgamated Pose-Free Encoding for Feed-Forward 3D Gaussian Splatting},
author={TODO(release): author list in publication order},
year={2026},
eprint={TODO(release): arXiv identifier},
archivePrefix={arXiv}
}