NemoSplat: Feed-Forward 4D Gaussian Splatting for Media-Aware Underwater Reconstruction

1The Hong Kong University of Science and Technology
2Beijing Institute of Technology

Corresponding Author

NemoSplat teaser

We present NemoSplat, a novel feed-forward model that directly reconstructs photorealistic underwater scenes from uncalibrated image sequences, robustly handling numerous dynamic objects and water-induced degradation.

Abstract

Feed-Forward 4D Gaussian Splatting Underwater Reconstruction Novel View Synthesis Dynamic Scenes

Reconstructing photorealistic scenes in unconstrained underwater environments remains challenging due to severe media-induced light scattering and unpredictable dynamic objects. Recent feed-forward visual foundation models have demonstrated remarkable capabilities in generalized novel view synthesis and tracking. However, when directly applied to aquatic videos, optical attenuation and motion interference fatally corrupt their feature aggregation, leading to severe tracking and reconstruction failures. To overcome these limitations, we present NemoSplat, the first feed-forward 4D Gaussian Splatting framework tailored for media-aware dynamic reconstruction directly from uncalibrated marine videos. Beyond providing robust estimations of camera poses and dense scene depth, we devise a Promptable Dynamic Disentangler that utilizes a confidence-aware fusion strategy of learned dynamic probabilities and optional semantic text priors, effectively isolating massive transient entities. Furthermore, to counteract visual degradation, a Media-Aware Gaussian Predictor is formulated to jointly estimate intrinsic 3D Gaussian attributes alongside physical media parameters, rendering pristine scene appearance in a single forward pass. Additionally, we introduce a large-scale underwater dataset with massive dynamic elements to facilitate training and evaluation. Extensive experiments on our dataset demonstrate that NemoSplat achieves state-of-the-art tracking accuracy and high-fidelity rendering.

Method Overview

Given uncalibrated marine videos, NemoSplat integrates geometry prediction, dynamic-static separation and physical media restoration into a single forward pass. The model leverages a Streaming Geometry Transformer (DINOv2 tokenizer + 24-layer causal alternating attention) to extract robust spatiotemporal features. A shared feature encoder feeds specialized decoder heads to jointly estimate camera poses, dense depth, dynamic masks, water medium parameters, and intrinsic 4D Gaussians.

NemoSplat pipeline overview

Overview of the NemoSplat framework: camera pose & depth heads anchor the geometry; the Promptable Dynamic Disentangler separates moving objects; the Media-Aware Gaussian Predictor jointly estimates 3D Gaussian attributes and physical water parameters (attenuation βD, backscatter βB, veiling light B).

Experiments

We compare NemoSplat with six state-of-the-art baselines spanning feed-forward foundation models (VGGT, StreamVGGT), feed-forward 3DGS (YoNoSplat, AnySplat), and optimization-based dynamic GS-SLAM systems (Droid-W, WildGS-SLAM) on our synthetic benchmark (BoulderShore / Coral / Deepsea) and 14 real-world aquatic sequences.

Synthetic sequences: tracking (ATE, m) and rendering (PSNR / SSIM / LPIPS)

Method BoulderShore Coral Deepsea Average
ATE↓PSNR↑SSIM↑LPIPS↓ ATE↓PSNR↑SSIM↑LPIPS↓ ATE↓PSNR↑SSIM↑LPIPS↓ ATE↓PSNR↑SSIM↑LPIPS↓
VGGT0.19---0.23---5.93---2.12---
StreamVGGT0.40---0.79---5.01---2.06---
WildGS-SLAM0.1218.800.600.454.9717.010.550.520.8129.860.810.351.9621.890.650.44
Droid-W0.1120.040.710.444.6518.560.580.560.7831.930.830.301.8523.510.710.43
YoNoSplat0.3917.810.460.550.2717.630.440.56× (OOM)×
AnySplat0.5420.050.560.382.5920.980.520.385.7824.240.410.402.9721.760.500.39
Ours0.2421.710.570.310.5420.410.560.374.8529.830.550.411.8823.980.560.36

Bold: best. Underlined: second best. “-”: method does not support rendering. ×: out of memory. NemoSplat achieves the best overall PSNR (23.98 dB) and LPIPS (0.36), with a highly competitive average ATE of 1.88 m.

Real-world sequences (14 scenes): novel view synthesis quality (PSNR / SSIM / LPIPS)

MethodMetric 01020304050607 08091011121314 Avg.
WildGS-SLAM PSNR↑19.3118.1320.8519.9219.2421.7519.7916.0821.4019.3820.7921.6416.2913.5019.15
SSIM↑0.560.650.740.620.520.740.620.420.680.490.690.780.580.370.60
LPIPS↓0.500.520.380.370.440.420.440.500.350.360.280.250.410.580.41
Droid-W PSNR↑20.5919.8318.6118.7317.7720.5017.0516.7118.4319.0517.9021.0513.2014.3218.12
SSIM↑0.730.690.510.550.400.680.490.500.600.460.630.720.500.430.56
LPIPS↓0.370.390.460.370.520.510.580.420.430.390.360.270.500.570.44
YoNoSplat PSNR↑15.3916.7315.4416.7114.8320.0616.4313.0518.1716.5812.2215.0513.7514.5215.64
SSIM↑0.440.350.330.420.330.660.540.380.520.370.350.410.340.350.41
LPIPS↓0.670.640.630.620.690.440.560.580.530.560.640.570.600.620.60
AnySplat PSNR↑20.4318.2817.6919.8018.8322.0420.1916.4421.5419.8420.4520.8615.1414.8019.22
SSIM↑0.680.470.460.630.550.750.720.480.730.550.690.680.460.420.60
LPIPS↓0.350.500.470.480.490.400.390.450.350.400.320.300.490.560.42
Ours PSNR↑20.9822.3518.2121.6720.1022.7922.8919.9924.8421.4520.7221.5321.2122.1321.58
SSIM↑0.600.690.480.680.570.820.670.750.730.630.690.610.700.750.68
LPIPS↓0.310.230.360.240.350.250.300.180.230.270.290.240.160.210.26

Bold: best per scene. Underlined: second best. NemoSplat achieves the best average on all three metrics (PSNR 21.58 / SSIM 0.68 / LPIPS 0.26), +2.36 dB PSNR over the second-best AnySplat and a 36.6% relative LPIPS reduction compared to WildGS-SLAM.

Qualitative Comparison

On highly degraded real-world aquatic sequences, conventional feed-forward models and SLAM-based baselines suffer from severe ghosting, topological blurring, and visual artifacts caused by unconstrained dynamic entities (e.g., schools of swimming fish). NemoSplat explicitly decouples transient motion from the static topology, synthesizing crisp, temporally consistent, artifact-free novel views.

Qualitative comparison of novel view synthesis

Reconstruction Results

NemoSplat disentangles static Gaussian models from dynamic underwater scenes, providing comprehensive outputs including RGB renderings, depth maps, and dynamic masks.

Reconstruction results on real-world sequences

Dataset

We construct a large-scale underwater dataset with 256 training sequences (155K frames) and 20 evaluation scenes. The training corpus spans 9 diverse geographic domains and employs a 13-class taxonomy to systematically capture complex marine life behaviors. Training data includes depth, water, and dynamic masks annotated via a human-refined SAM3 + SAHI pipeline. For evaluation, 6 synthetic UE5 sequences provide precise ground truth.

Training dataset examples

Example sequences of our training dataset, covering 9 diverse geographic domains with a 13-class taxonomy of marine life behaviors.

Evaluation dataset examples

Example sequences of our proposed evaluation dataset, including synthetic scenes with precise ground-truth annotations and diverse in-the-wild underwater sequences.

BibTeX

@article{guo2026nemosplat,
  title   = {NemoSplat: Feed-Forward 4D Gaussian Splatting for Media-Aware Underwater Reconstruction},
  author  = {Guo, Xiaopeng and Tse, Wai Chung and Zhu, Yipeng and Zhang, Hanwen and Huang, Huajian and Yeung, Sai-Kit},
  journal = {arXiv preprint},
  year    = {2026}
}