Dynamic-Robust Photometric–Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/dmucby/SPAR
Area: 3D Vision
Keywords: Dynamic Scene Reconstruction, Open-Vocabulary 3D Scene Understanding, Novel View Synthesis, Dynamic Region Prediction, Self-Supervised Learning
TL;DR¶
To tackle spatial feature misalignments caused by moving foregrounds in dynamic environments, SPAR introduces a Cross-View Dynamic Region Predictor alongside a joint semantic-geometric encoder-decoder that filters transient motion noise before latent aggregation and performs self-supervised, dynamic-region-aware reconstruction in a single forward pass.
Background & Motivation¶
The fusion of novel view synthesis (NVS) and open-vocabulary segmentation (OVS) has driven the emergence of powerful feed-forward 3D foundation models. Contemporary frameworks such as LSM and SIU3R infer geometry, appearance, and open-vocabulary semantic fields directly from sparse and unposed multi-view observations, bypassing the substantial computational overhead of per-scene neural radiance field (NeRF) or 3D Gaussian Splatting (3DGS) optimization. However, existing feed-forward 3D reconstruction systems rely on a strict fundamental assumption: all captured input views must depict an entirely static 3D world.
In in-the-wild video acquisitions, this static assumption is routinely violated by transient dynamic objects such as moving pedestrians, domestic pets, and passing vehicles. Because an object in motion occupies distinct 3D positions across different temporal timestamps, its 2D projected regions cannot be mapped onto a single, consistent static world coordinate frame. When conventional models aggregate these unisolated dynamic features into global latent spaces, cross-view geometric and semantic collisions inevitably emerge, causing severe ghosting artifacts in RGB view synthesis and fragmenting the downstream 3D semantic fields. Although recent efforts like WildRayZer attempt to mitigate dynamic distractors via mask suppression, they restrict their scope to photometric reconstruction alone and decouple geometry from high-level semantics, thereby forfeiting the potential structural regularization that semantic representations offer to geometric modeling.
This paper tackles the challenge by unifying dynamic motion disentanglement, photometric reconstruction, and semantic field inference into an integrated framework. The core idea is to build SPAR, a joint semantic-geometric encoding architecture that isolates transient motion tokens prior to latent aggregation via a Cross-View Dynamic Region Predictor, and trains the model end-to-end using a dynamic-region-aware self-supervised paradigm where predicted static probability masks spatially gate reconstruction errors.
Method¶
Overall Architecture¶
SPAR operates on an unposed sparse multi-view input depicting dynamic environments, synthesizing target-view RGB images alongside continuous open-vocabulary semantic feature maps in a single forward pass. The workflow comprises four synchronized phases: input tokenization and dynamic filtering, joint photometric-semantic scene encoding, dual-head rendering decoding, and dynamic-region-aware self-supervised optimization. First, camera poses are estimated to construct Plücker ray-conditioned photometric and semantic tokens from raw images and 2D semantic foundations. Next, a Cross-View Dynamic Region Predictor detects cross-view motion inconsistencies, predicting transient foreground masks to physically discard corrupted tokens. The cleaned static tokens are then aggregated into learnable global scene tokens via a Transformer scene encoder. Finally, a rendering decoder queries the unified scene tokens with target-view Plücker rays, driving dual MLP heads to synthesize target-view RGB appearance and continuous high-dimensional semantic embeddings.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Dynamic Multi-View Inputs<br/>Unposed RGB & Semantics"] --> B["Cross-View Dynamic Region Predictor<br/>Multimodal Concat & Cross-View Attention"]
B --> C["Dynamic Token Spatial Exclusion<br/>Static Probability & Hard Filtering"]
C --> D["Joint Pho-Sem Scene Encoder<br/>Global Scene Tokens & Cleaned Tokens Attention"]
D --> E["Dual-Head Rendering Decoder<br/>Target Plücker Ray Query & Dual MLPs"]
E --> F["Dynamic-Region-Aware Gated Optimization<br/>Static Probability Weighted Pho-Sem Loss"]
Key Designs¶
1. Cross-View Dynamic Region Predictor: Isolating Transient Noise Prior to Latent Aggregation Single-view observations lack depth and temporal parallax baselines to distinguish moving foreground objects from complex static geometry. To eliminate transient interference before spatial encoding, SPAR introduces the Cross-View Dynamic Region Predictor (CV-DRP). For an input view \(I_i \in \mathbb{R}^{H \times W \times 3}\), the model concatenates its pre-extracted LSeg semantic features \(F_{sem}^{(i)}\), photometric patch tokens \(F_{pho}^{(i)}\), and Plücker ray tokens \(T_{ray}^{(i)}\) along the channel dimension into a unified representation \(F^{(i)} \in \mathbb{R}^{N \times 3D}\). Paired source and reference features \((F^{(i)}, F^{(j)})\) are fed into stacked Transformer self-attention layers to identify cross-view epipolar violations and appearance-semantic discrepancies. An MLP and bilinear upsampling stage then output a full-resolution dynamic probability map \(P_m^{(i)} \in \mathbb{R}^{H \times W}\), yielding a binary mask via thresholding: $\(M^{(i)} = \mathbb{I}(P_m^{(i)} > \tau)\)$ During inference, SAM2 can optionally refine these spatial priors. During training and feed-forward execution, the mask directly excises dynamic tokens before scene encoding, ensuring the global latent representation remains uncorrupted by transient motion.
2. Joint Pho-Sem Scene Encoder and Dual-Head Decoder: Deep Cross-Modal Scene Fusion Rather than compressing semantic features into individual discrete 3D Gaussian primitives or rasterized surfels, which inevitably degrades semantic granularity, SPAR adopts a centralized latent token interaction scheme. The network initializes a compact set of learnable global scene tokens \(T_{scene} \in \mathbb{R}^{S \times D}\) (\(S=768\)) and concatenates them with the masked static photometric tokens \(\tilde{T}_{pho,in}\) and semantic tokens \(\tilde{T}_{sem,in}\) to form the unified sequence \(T_{all}\). An \(L=8\) layer Transformer encoder performs dense self-attention, allowing global scene tokens to distill cross-modal appearance details and high-level categorical co-occurrence priors into latent representations \(T_{scene}'\). In the decoding stage, target-view Plücker rays are projected into photometric and semantic query tokens (\(T_{ray\_pho}^{tgt}\) and \(T_{ray\_sem}^{tgt}\)), concatenated with \(T_{scene}'\), and processed by a \(12\)-layer Transformer decoder. Two distinct MLP heads subsequently decode the updated features into target-view RGB values and continuous semantic vectors: $\(R_{pho} = \text{MLP}_{pho}(T_{out\_pho}) \in \mathbb{R}^{M \times 3}, \quad R_{sem} = \text{MLP}_{sem}(T_{out\_sem}) \in \mathbb{R}^{M \times C}\)$ The continuous semantic embeddings preserve the full zero-shot generalization capabilities of open-vocabulary vision-language models like CLIP.
3. Dynamic-Region-Aware Optimization: Reconstruction-Driven Self-Supervised Learning Supervising motion masks with ground-truth temporal pixel annotations is prohibitively expensive and hard to scale across web video corpora. SPAR solves this by turning multi-view photometric and semantic rendering consistency into a self-supervising signal. Because moving objects cannot be explained by static multi-view epipolar geometry, unmodeled dynamic regions naturally produce elevated reconstruction residuals. Using the dynamic probability map \(P_m\) predicted by CV-DRP for each target view, the network computes the static probability map \(P_s = 1 - P_m\) to spatially gate both photometric and semantic reconstruction errors: $\(\mathcal{L}_{pho}^{dyn} = \frac{1}{\sum P_s} \sum P_s \odot \left( \| R_{pho} - I_{tgt} \|_1 + \lambda_{percp} \cdot \mathrm{Percep}(R_{pho}, I_{tgt}) \right)\)$ $\(\mathcal{L}_{sem}^{dyn} = \frac{1}{\sum \hat{P}_s} \sum \hat{P}_s \odot \left( 1 - \frac{R_{sem} \cdot F_{sem}^{tgt}}{\|R_{sem}\|_2 \|F_{sem}^{tgt}\|_2} \right)\)$ where \(\hat{P}_s\) denotes the downsampled static map matching the semantic resolution. To prevent the predictor from collapsing into a trivial all-dynamic solution, a regularization penalty \(\mathcal{L}_{reg} = \frac{1}{HW}\sum_{i,j} \text{BCE}(P_m^{(i,j)}, 0)\) enforces sparsity. Additionally, a copy-paste augmentation strategy pastes random object instances onto static scenes to simulate dynamic foregrounds, generating synthetic ground-truth masks for an auxiliary binary cross-entropy loss \(\mathcal{L}_{mask}\). This closed-loop formulation enables mutual reinforcement: rendering errors train the dynamic predictor, while predicted masks insulate static scene learning.
Loss & Training¶
The overall training objective is a weighted combination of dynamic-gated and regularization losses: $\(\mathcal{L}_{total} = \lambda_{pho} \mathcal{L}_{pho}^{dyn} + \lambda_{sem} \mathcal{L}_{sem}^{dyn} + \lambda_{reg} \mathcal{L}_{reg} + \lambda_{mask} \mathcal{L}_{mask}\)$ Hyperparameters are set to \(\lambda_{pho} = 1.0\), \(\lambda_{percp} = 0.2\), \(\lambda_{sem} = 0.1\), \(\lambda_{mask} = 1.0\), and \(\lambda_{reg} = 0.01\). Training is performed on \(8 \times\) NVIDIA A800 GPUs using a cosine learning rate scheduler. Pre-training randomly masks 10% of photometric input tokens to build robustness against spatial missingness; end-to-end training runs for 20K iterations with a learning rate of \(2 \times 10^{-4}\) at \(256 \times 256\) resolution with \(16 \times 16\) patch size.
Key Experimental Results¶
Main Results¶
On the challenging D-RE10K dynamic indoor benchmark, SPAR is evaluated across sparse input view regimes (\(n = 2, 3, 4\)) under unposed conditions. The results demonstrate that SPAR establishes state-of-the-art novel view synthesis fidelity, successfully eliminating motion ghosting.
| Method | Type | 2 Views PSNR↑ | 2 Views LPIPS↓ | 3 Views PSNR↑ | 3 Views LPIPS↓ | 4 Views PSNR↑ | 4 Views LPIPS↓ |
|---|---|---|---|---|---|---|---|
| NeRF On-the-go | Per-scene Opt. | 15.90 | 0.582 | 18.45 | 0.446 | 19.52 | 0.443 |
| 3DGS | Per-scene Opt. | 13.49 | 0.605 | 14.92 | 0.531 | 16.28 | 0.490 |
| Spotless-Splats | Per-scene Opt. | 16.45 | 0.468 | 17.77 | 0.394 | 18.05 | 0.390 |
| LSM | Feed-forward | 10.92 | 0.636 | 11.48 | 0.638 | 11.42 | 0.644 |
| SIU3R | Feed-forward | 13.01 | 0.577 | 13.22 | 0.564 | 13.25 | 0.553 |
| WildRayZer | Feed-forward | 21.78 | 0.308 | 21.98 | 0.314 | 22.38 | 0.290 |
| SPAR (Ours) | Feed-forward | 19.97 | 0.339 | 22.15 | 0.283 | 23.33 | 0.263 |
In dynamic motion mask evaluation, SPAR delivers segmentation accuracy superior to both self-supervised and fully supervised baselines.
| Method | Supervision | n=2 mIoU↑ | n=2 Recall↑ | n=3 mIoU↑ | n=3 Recall↑ | n=8 mIoU↑ | n=8 Recall↑ |
|---|---|---|---|---|---|---|---|
| Co-segmentation | Self-supervised | 9.6% | 45.0% | 13.7% | 53.2% | 16.3% | 45.5% |
| Segment Any Motion | Supervised | 31.9% | 47.2% | 41.2% | 57.1% | 50.9% | 70.8% |
| WildRayZer | Self-supervised | 53.9% | 85.1% | 52.1% | 84.3% | 54.2% | 87.7% |
| Ours w/o Refine | Self-supervised | 75.9% | 85.7% | 75.7% | 86.4% | 76.6% | 86.6% |
| Ours w/ Refine | Self-supervised+SAM2 | 87.8% | 84.7% | 88.5% | 86.2% | 88.3% | 85.6% |
On the ScanNet semantic understanding benchmark (2 input views, 1 novel target view): - Semantic Segmentation: SPAR achieves 0.5271 mIoU and 0.8061 Accuracy, outperforming LSM (0.5078 mIoU, 0.7686 Accuracy). - Novel View Synthesis: SPAR attains a PSNR of 26.43 dB, significantly exceeding LSM (24.39 dB), Feature-3DGS (24.49 dB), and SIU3R (25.10 dB).
Ablation Study¶
Ablation experiments validate the distinct performance contributions of dynamic-region-aware optimization, the parallel semantic branch, and auxiliary components.
| Config | PSNR↑ | SSIM↑ | LPIPS↓ | Note |
|---|---|---|---|---|
| Unconstrained Baseline | 15.66 | 0.489 | 0.532 | Neglecting dynamics causes severe ghosting and misalignment |
| Baseline + Random Masking | 17.10 | 0.536 | 0.485 | Random masking provides only limited regularization (+1.44 dB) |
| + Dynamic-Region-Aware Opt. (Ours) | 19.97 | 0.627 | 0.339 | Dynamic exclusion removes motion conflicts, boosting PSNR by +4.31 dB |
| Full model w/o semantic branch | 19.77 | 0.615 | 0.356 | Pure geometric-appearance modeling without semantic priors |
| Full model w/ semantic branch (Ours) | 19.97 | 0.627 | 0.339 | Semantic learning acts as a structural regularizer (+0.20 dB PSNR) |
| Full model w/o pre-training | 18.66 | 0.578 | 0.381 | Lacking general feed-forward 3D priors drops PSNR by 1.31 dB |
| Full model w/o CPA | 19.82 | 0.625 | 0.348 | Minor drop due to absence of synthetic motion boundary supervision |
| Full model w/o SAM2 refinement | 19.90 | 0.624 | 0.341 | Demonstrates the raw self-supervised predictor is already highly robust |
Key Findings¶
- Inter-Task Synergy: Introducing the semantic understanding branch does not compete with appearance rendering capacity; rather, it provides vital structural guidance that improves photometric PSNR by 0.20 dB (19.77 dB to 19.97 dB) and lowers LPIPS from 0.356 to 0.339.
- Unsupervised Dynamic Motion Discovery: Driven entirely by multi-view reconstruction loss gating, the raw predictor without SAM2 post-processing achieves 75.7% mIoU under 3 views, surpassing WildRayZer's 52.1% by over 23 percentage points.
- Sparse-to-Dense View Scalability: With 3 and 4 input views, SPAR exhibits superior scaling over existing approaches, reaching 23.33 dB PSNR at 4 views (+0.95 dB over WildRayZer), proving the effectiveness of the global scene token interaction.
Highlights & Insights¶
- Latent Isolation Before Aggregation: Discarding dynamic tokens at the input level prevents corrupting the latent scene representation \(T_{scene}\), addressing ghosting artifacts at their architectural origin.
- Mutual Reinforcement of Geometry and Semantics: Rather than treating novel view synthesis and semantic segmentation as conflicting tasks, SPAR demonstrates that high-level semantic fields act as spatial regularizers for sparse-view photometric reconstruction.
- Zero-Shot Outdoor Generalization: Trained on indoor datasets, SPAR generalizes zero-shot to complex unbounded outdoor scenes in SpatialVID, reliably segmenting moving cyclists and pedestrians.
Limitations & Future Work¶
- Geometric Degradation in Extreme Two-View Baselines: When only two views are available with large camera baselines, cross-view epipolar correspondences become sparse, occasionally causing non-rigid deformations to be misclassified as static occlusions.
- Complex Non-Rigid Environmental Dynamics: Natural dynamic background elements such as water ripples, swaying foliage, or moving cloud shadows may induce false-positive dynamic mask detections under strict photometric consistency assumptions.
- Real-Time Edge Deployment: The dense Transformer cross-attention mechanisms demand significant compute; adapting the latent scene tokens to feed-forward 3D Gaussian representations could facilitate real-time robotics applications.
Related Work & Insights¶
- vs LSM (Fan et al., 2024): LSM achieves feed-forward unposed 3D semantic Gaussians but assumes static scenes and suffers catastrophic ghosting on dynamic data; SPAR isolates dynamic foregrounds and avoids semantic compression bottlenecks via latent scene tokens.
- vs WildRayZer (Chen et al., 2026): WildRayZer tackles dynamic feed-forward NVS but omits semantic understanding and yields lower motion mask mIoU (~52%); SPAR achieves 75.7% self-supervised mIoU (88.5% with refine) and simultaneous open-vocabulary segmentation.
- vs SIU3R (Wei et al., 2026): SIU3R introduces alignment-free 3D multi-task representation learning but cannot process moving objects; SPAR provides the requisite robustness for unconstrained dynamic videos.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering unification of open-vocabulary 3D scene understanding and dynamic-robust feed-forward reconstruction.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks across D-RE10K, ScanNet, and zero-shot SpatialVID, coupled with rigorous ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Cohesive formulation, clear mathematical derivation, and structured presentation.
- Value: ⭐⭐⭐⭐⭐ Offers a practical feed-forward foundation for embodied navigation and 3D perception in real-world dynamic environments.