NemoSplat: Feed-Forward 4D Gaussian Splatting for Media-Aware Underwater Reconstruction
We present NemoSplat, a novel feed-forward model that directly reconstructs photorealistic underwater scenes from uncalibrated image sequences, robustly handling numerous dynamic objects and water-induced degradation.
Abstract
Reconstructing photorealistic scenes in unconstrained underwater environments remains challenging due to severe media-induced light scattering and unpredictable dynamic objects. Recent feed-forward visual foundation models have demonstrated remarkable capabilities in generalized novel view synthesis and tracking. However, when directly applied to aquatic videos, optical attenuation and motion interference fatally corrupt their feature aggregation, leading to severe tracking and reconstruction failures. To overcome these limitations, we present NemoSplat, the first feed-forward 4D Gaussian Splatting framework tailored for media-aware dynamic reconstruction directly from uncalibrated marine videos. Beyond providing robust estimations of camera poses and dense scene depth, we devise a Promptable Dynamic Disentangler that utilizes a confidence-aware fusion strategy of learned dynamic probabilities and optional semantic text priors, effectively isolating massive transient entities. Furthermore, to counteract visual degradation, a Media-Aware Gaussian Predictor is formulated to jointly estimate intrinsic 3D Gaussian attributes alongside physical media parameters, rendering pristine scene appearance in a single forward pass. Additionally, we introduce a large-scale underwater dataset with massive dynamic elements to facilitate training and evaluation. Extensive experiments on our dataset demonstrate that NemoSplat achieves state-of-the-art tracking accuracy and high-fidelity rendering.
Method Overview
Given uncalibrated marine videos, NemoSplat integrates geometry prediction, dynamic-static separation and physical media restoration into a single forward pass. The model leverages a Streaming Geometry Transformer (DINOv2 tokenizer + 24-layer causal alternating attention) to extract robust spatiotemporal features. A shared feature encoder feeds specialized decoder heads to jointly estimate camera poses, dense depth, dynamic masks, water medium parameters, and intrinsic 4D Gaussians.
Overview of the NemoSplat framework: camera pose & depth heads anchor the geometry; the Promptable Dynamic Disentangler separates moving objects; the Media-Aware Gaussian Predictor jointly estimates 3D Gaussian attributes and physical water parameters (attenuation βD, backscatter βB, veiling light B∞).
Experiments
We compare NemoSplat with six state-of-the-art baselines spanning feed-forward foundation models (VGGT, StreamVGGT), feed-forward 3DGS (YoNoSplat, AnySplat), and optimization-based dynamic GS-SLAM systems (Droid-W, WildGS-SLAM) on our synthetic benchmark (BoulderShore / Coral / Deepsea) and 14 real-world aquatic sequences.
Synthetic sequences: tracking (ATE, m) and rendering (PSNR / SSIM / LPIPS)
| Method | BoulderShore | Coral | Deepsea | Average | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ATE↓ | PSNR↑ | SSIM↑ | LPIPS↓ | ATE↓ | PSNR↑ | SSIM↑ | LPIPS↓ | ATE↓ | PSNR↑ | SSIM↑ | LPIPS↓ | ATE↓ | PSNR↑ | SSIM↑ | LPIPS↓ | |
| VGGT | 0.19 | - | - | - | 0.23 | - | - | - | 5.93 | - | - | - | 2.12 | - | - | - |
| StreamVGGT | 0.40 | - | - | - | 0.79 | - | - | - | 5.01 | - | - | - | 2.06 | - | - | - |
| WildGS-SLAM | 0.12 | 18.80 | 0.60 | 0.45 | 4.97 | 17.01 | 0.55 | 0.52 | 0.81 | 29.86 | 0.81 | 0.35 | 1.96 | 21.89 | 0.65 | 0.44 |
| Droid-W | 0.11 | 20.04 | 0.71 | 0.44 | 4.65 | 18.56 | 0.58 | 0.56 | 0.78 | 31.93 | 0.83 | 0.30 | 1.85 | 23.51 | 0.71 | 0.43 |
| YoNoSplat | 0.39 | 17.81 | 0.46 | 0.55 | 0.27 | 17.63 | 0.44 | 0.56 | × (OOM) | × | ||||||
| AnySplat | 0.54 | 20.05 | 0.56 | 0.38 | 2.59 | 20.98 | 0.52 | 0.38 | 5.78 | 24.24 | 0.41 | 0.40 | 2.97 | 21.76 | 0.50 | 0.39 |
| Ours | 0.24 | 21.71 | 0.57 | 0.31 | 0.54 | 20.41 | 0.56 | 0.37 | 4.85 | 29.83 | 0.55 | 0.41 | 1.88 | 23.98 | 0.56 | 0.36 |
Bold: best. Underlined: second best. “-”: method does not support rendering. ×: out of memory. NemoSplat achieves the best overall PSNR (23.98 dB) and LPIPS (0.36), with a highly competitive average ATE of 1.88 m.
Real-world sequences (14 scenes): novel view synthesis quality (PSNR / SSIM / LPIPS)
| Method | Metric | 01 | 02 | 03 | 04 | 05 | 06 | 07 | 08 | 09 | 10 | 11 | 12 | 13 | 14 | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| WildGS-SLAM | PSNR↑ | 19.31 | 18.13 | 20.85 | 19.92 | 19.24 | 21.75 | 19.79 | 16.08 | 21.40 | 19.38 | 20.79 | 21.64 | 16.29 | 13.50 | 19.15 |
| SSIM↑ | 0.56 | 0.65 | 0.74 | 0.62 | 0.52 | 0.74 | 0.62 | 0.42 | 0.68 | 0.49 | 0.69 | 0.78 | 0.58 | 0.37 | 0.60 | |
| LPIPS↓ | 0.50 | 0.52 | 0.38 | 0.37 | 0.44 | 0.42 | 0.44 | 0.50 | 0.35 | 0.36 | 0.28 | 0.25 | 0.41 | 0.58 | 0.41 | |
| Droid-W | PSNR↑ | 20.59 | 19.83 | 18.61 | 18.73 | 17.77 | 20.50 | 17.05 | 16.71 | 18.43 | 19.05 | 17.90 | 21.05 | 13.20 | 14.32 | 18.12 |
| SSIM↑ | 0.73 | 0.69 | 0.51 | 0.55 | 0.40 | 0.68 | 0.49 | 0.50 | 0.60 | 0.46 | 0.63 | 0.72 | 0.50 | 0.43 | 0.56 | |
| LPIPS↓ | 0.37 | 0.39 | 0.46 | 0.37 | 0.52 | 0.51 | 0.58 | 0.42 | 0.43 | 0.39 | 0.36 | 0.27 | 0.50 | 0.57 | 0.44 | |
| YoNoSplat | PSNR↑ | 15.39 | 16.73 | 15.44 | 16.71 | 14.83 | 20.06 | 16.43 | 13.05 | 18.17 | 16.58 | 12.22 | 15.05 | 13.75 | 14.52 | 15.64 |
| SSIM↑ | 0.44 | 0.35 | 0.33 | 0.42 | 0.33 | 0.66 | 0.54 | 0.38 | 0.52 | 0.37 | 0.35 | 0.41 | 0.34 | 0.35 | 0.41 | |
| LPIPS↓ | 0.67 | 0.64 | 0.63 | 0.62 | 0.69 | 0.44 | 0.56 | 0.58 | 0.53 | 0.56 | 0.64 | 0.57 | 0.60 | 0.62 | 0.60 | |
| AnySplat | PSNR↑ | 20.43 | 18.28 | 17.69 | 19.80 | 18.83 | 22.04 | 20.19 | 16.44 | 21.54 | 19.84 | 20.45 | 20.86 | 15.14 | 14.80 | 19.22 |
| SSIM↑ | 0.68 | 0.47 | 0.46 | 0.63 | 0.55 | 0.75 | 0.72 | 0.48 | 0.73 | 0.55 | 0.69 | 0.68 | 0.46 | 0.42 | 0.60 | |
| LPIPS↓ | 0.35 | 0.50 | 0.47 | 0.48 | 0.49 | 0.40 | 0.39 | 0.45 | 0.35 | 0.40 | 0.32 | 0.30 | 0.49 | 0.56 | 0.42 | |
| Ours | PSNR↑ | 20.98 | 22.35 | 18.21 | 21.67 | 20.10 | 22.79 | 22.89 | 19.99 | 24.84 | 21.45 | 20.72 | 21.53 | 21.21 | 22.13 | 21.58 |
| SSIM↑ | 0.60 | 0.69 | 0.48 | 0.68 | 0.57 | 0.82 | 0.67 | 0.75 | 0.73 | 0.63 | 0.69 | 0.61 | 0.70 | 0.75 | 0.68 | |
| LPIPS↓ | 0.31 | 0.23 | 0.36 | 0.24 | 0.35 | 0.25 | 0.30 | 0.18 | 0.23 | 0.27 | 0.29 | 0.24 | 0.16 | 0.21 | 0.26 |
Bold: best per scene. Underlined: second best. NemoSplat achieves the best average on all three metrics (PSNR 21.58 / SSIM 0.68 / LPIPS 0.26), +2.36 dB PSNR over the second-best AnySplat and a 36.6% relative LPIPS reduction compared to WildGS-SLAM.
Qualitative Comparison
On highly degraded real-world aquatic sequences, conventional feed-forward models and SLAM-based baselines suffer from severe ghosting, topological blurring, and visual artifacts caused by unconstrained dynamic entities (e.g., schools of swimming fish). NemoSplat explicitly decouples transient motion from the static topology, synthesizing crisp, temporally consistent, artifact-free novel views.
Reconstruction Results
NemoSplat disentangles static Gaussian models from dynamic underwater scenes, providing comprehensive outputs including RGB renderings, depth maps, and dynamic masks.
Dataset
We construct a large-scale underwater dataset with 256 training sequences (155K frames) and 20 evaluation scenes. The training corpus spans 9 diverse geographic domains and employs a 13-class taxonomy to systematically capture complex marine life behaviors. Training data includes depth, water, and dynamic masks annotated via a human-refined SAM3 + SAHI pipeline. For evaluation, 6 synthetic UE5 sequences provide precise ground truth.
Example sequences of our training dataset, covering 9 diverse geographic domains with a 13-class taxonomy of marine life behaviors.
Example sequences of our proposed evaluation dataset, including synthetic scenes with precise ground-truth annotations and diverse in-the-wild underwater sequences.
BibTeX
@article{guo2026nemosplat,
title = {NemoSplat: Feed-Forward 4D Gaussian Splatting for Media-Aware Underwater Reconstruction},
author = {Guo, Xiaopeng and Tse, Wai Chung and Zhu, Yipeng and Zhang, Hanwen and Huang, Huajian and Yeung, Sai-Kit},
journal = {arXiv preprint},
year = {2026}
}