Skip to content

Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation

Conference: ECCV 2026
Paper: ECCV Official
Full-text Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5627.txt
Area: 3D Vision
Keywords: large view synthesis models, multi-view panoptic segmentation, implicit cross-view correspondence, reconstruction-free label propagation, zero-shot cross-dataset transfer

TL;DR

This paper presents a decoupled multi-view panoptic segmentation framework demonstrating that a frozen Large View Synthesis Model (LVSM) trained solely on RGB photometric reconstruction inherently learns input-agnostic geometric correspondence, enabling direct propagation of 3-bit binary-encoded panoptic labels to unobserved novel viewpoints without explicit 3D reconstruction or task-specific fine-tuning (reaching 33.56 dB PSNR and 0.5949 mIoU on ScanNet).

Background & Motivation

Predicting panoptic segmentation from novel viewpoints requires an autonomous agent to assign both semantic category labels and distinct instance identities to every pixel in an unobserved view, given only a sparse set of unposed scene images. This capability forms an indispensable foundation for embodied agents anticipating scene layout prior to navigation and expanding training annotations from limited views. The prevailing paradigm over the past few years has strictly relied on a two-stage "reconstruct-then-lift" pipeline: first reconstructing an explicit 3D representation via Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS), and subsequently lifting noisy 2D segmentations into the 3D primitives through per-scene optimization. Even recent feed-forward advances (such as LSM and SIU3R) that bypass per-scene optimization still tie semantic features directly to pixel-aligned 3D Gaussians or dense point clouds.

The central tension of these explicit 3D pipelines is that segmentation quality is fundamentally bottlenecked by geometric reconstruction quality: errors, geometric distortions, or floaters in the reconstructed 3D representation propagate directly to the lifted semantic masks. Furthermore, joint optimization introduces severe gradient conflict between appearance synthesis and categorical discrimination, catastrophically degrading novel view rendering fidelity (often plummeting to 20–26 dB PSNR). Conversely, feed-forward large view synthesis models (LVSMs such as LVSM and Less3Depend) leverage self- and cross-attention mechanisms to establish implicit geometric relationships across view tokens, rendering novel views with state-of-the-art fidelity without explicit 3D inductive biases. However, all prior investigations have treated LVSMs strictly as appearance renderers, leaving completely unexplored whether their learned cross-view correspondence generalizes to non-photorealistic dense spatial signals.

Through gradient-based saliency analysis, the authors uncover that the cross-view attention within an unposed LVSM is driven primarily by relative geometric pose rather than RGB appearance features: when natural RGB images are replaced with non-photorealistic binary instance encodings, the network's attention continues to concentrate precisely on geometrically corresponding source regions. The core idea is to decouple multi-view panoptic segmentation from novel view rendering by repurposing a frozen, RGB-only pretrained large view synthesis model as a universal geometric propagation engine, directly transferring discrete binary-encoded source panoptic labels to target views without explicit 3D reconstruction while unlocking a rendering leap of over 7 dB.

Method

Overall Architecture

At inference time, the proposed system operates via two decoupled parallel pathways through the identical large view synthesis model: a novel-view RGB rendering pathway and a novel-view panoptic propagation pathway. Given \(N\) uncalibrated sparse source views, the pipeline first extracts cross-view-consistent panoptic segmentations using a shared query decoder. The resulting masks are then converted into compact 3-bit binary channel representations and fed into the frozen LVSM (built upon the unposed Less3Depend backbone). Finally, an inverse thresholded decoding step recovers clean, un-aliased panoptic segmentation maps on the novel viewpoint.

The end-to-end operational flow and component dependencies are depicted below:

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Source Images<br/>N unposed RGB views"] --> B["Shared Query Source Segmentation<br/>Joint cross-view Mask2Former decoding"]
    A --> C["Scene Latent Modeling<br/>DINOv2 encoding & Plücker estimation"]
    B --> D["Binary Channel Encoding & Decoding<br/>3-bit discrete quantization & LUT mapping"]
    C --> E["Frozen LVSM Cross-View Propagation<br/>Dual-stream feedforward & uncertainty filtering"]
    D --> E
    E --> F["Novel View Rendering Output<br/>High-fidelity RGB image (33.56 dB)"]
    E --> G["Novel View Panoptic Output<br/>Spatially consistent panoptic map (0.5949 mIoU)"]

Key Designs

1. Shared Query Source Segmentation: Enforcing Cross-View Instance Identity Consistency

The foundational prerequisite for cross-view label propagation is that identical physical objects across source views must possess strictly matching instance IDs; contradictory assignments on the same object inevitably trigger catastrophic attention ambiguity during propagation. Rather than relying on fragile post-hoc geometric feature matching that frequently fails under large baselines or partial occlusions, the method employs a shared query decoder \(\mathcal{D}\) built upon a frozen DINOv2 visual backbone equipped with a ViT-Adapter. The decoder introduces a set of shared, learnable object queries \(Q = \{q_j\}_{j=1}^{Q}\) that concurrently attend to feature representations across all \(N\) input views:

\[\{(m_j, c_j)\}_{j=1}^{Q} = \mathcal{D}(Q, I_1^s, \dots, I_N^s)\]

where \(m_j \in [0, 1]^{N \times H \times W}\) represents the spatial mask predicted by query \(q_j\) jointly across all \(N\) source views, and \(c_j \in \mathcal{C}\) denotes its semantic class. Because each query attends across all views simultaneously, it natively tracks and grounds a unique physical entity in 3D space, ensuring robust cross-view instance identity alignment without requiring hand-crafted matching heuristics. Crucially, this clean interface also permits seamless zero-shot modular swapping with modern off-the-shelf segmenters such as SAM2 or PanSt3R.

2. Binary Channel Encoding & Decoding: Discrete Immunity Against Boundary Interpolation Artifacts

Large view synthesis models natively output continuous floating-point responses via linear combinations in attention layers. If instance IDs or continuous multi-class logits are directly propagated through the network, the continuous spatial blending across object boundaries forces intermediate pixel values to slide across valid class intervals, spawning severe, unfilterable "phantom instances." To eliminate this vulnerability, the authors introduce an interpolation-resilient 3-bit binary channel encoding scheme. For each pixel \(p\), its assigned instance index \(I(p) \in \{0, \dots, K-1\}\) (supporting up to \(K \le 8\) active instances per single forward pass) is encoded into a 3-dimensional binary codeword:

\[b(p) = \mathcal{B}(I(p)) \in \{0, 1\}^3\]

accompanied by a lightweight look-up table \(\mathcal{L}\) storing the instance-to-class mapping. During novel view inference, the continuous model outputs \(\hat{b}^t(p) \in [0, 1]^3\). Along boundary transition bands where 0 and 1 blend linearly, intermediate values naturally converge toward \(0.5\)—the exact midpoint that is maximally distant in Euclidean space from both valid binary states \(\{0, 1\}\). By establishing a confidence decision threshold \(\tau = 0.2\):

\[\hat{I}^t(p) = \mathcal{B}^{-1}(\text{round}(\hat{b}^t(p))), \quad \hat{S}^t(p) = \mathcal{L}(\hat{I}^t(p))\]

any pixel satisfying \(|\hat{b}_{:,:,k}^t(p) - 0.5| \le \tau\) across any channel \(k\) is designated as an uncertain boundary pixel and safely left unassigned. This mechanism cleanly filters out transition artifacts without polluting foreground predictions. For complex scenes exceeding 8 active instances, a multi-pass propagation protocol is applied to encompass all objects.

3. Frozen LVSM Cross-View Propagation: Zero-Cost Reuse of Photometric Correspondence Priors

To dismantle the computational overhead and degradation tied to explicit 3D Gaussian representations, this design exploits the critical insight uncovered by gradient saliency analysis: the attention mapping inside an LVSM (Less3Depend) is steered by Plücker ray geometry rather than input appearance semantics. Consequently, for the novel-view panoptic synthesis path, the pipeline completely reuses the scene latent \(z\) and target view camera embedding \(\hat{e}_t\) derived from the RGB path:

\[z^{\mathrm{seg}} = \mathcal{E}(b_1^s, \dots, b_N^s, \hat{e}_t, z_0), \quad \hat{b}^t = \mathcal{R}(z^{\mathrm{seg}}, \hat{e}_t)\]

All weights in the scene encoder \(\mathcal{E}\) and render decoder \(\mathcal{R}\) remain strictly frozen, with zero gradients backpropagated from semantic annotations. This complete decoupling guarantees that the RGB rendering branch operates at its pristine synthesis quality, avoids accumulation of intermediate geometric reconstruction artifacts, and allows the framework to directly inherit future performance gains as more powerful foundation LVSM architectures emerge.

Loss & Training

The framework's training protocol is entirely segregated into two independent stages: - View Synthesis Model (Less3Depend Backbone): Trained exclusively on sparse multi-view images using photometric reconstruction objectives (pixel MSE, LPIPS perceptual loss, and unposed ray consistency losses), completely isolated from any semantic or instance supervision. - Source-View Segmentation (Shared Query Decoder \(\mathcal{D}\)): Supervised on annotated source views using the standard Mask2Former multi-task loss formulation combining binary cross-entropy, Dice loss, and categorical cross-entropy: $\(\mathcal{L}_{\mathrm{seg}} = \lambda_{\mathrm{cls}} \mathcal{L}_{\mathrm{CE}} + \lambda_{\mathrm{mask}} \mathcal{L}_{\mathrm{BCE}} + \lambda_{\mathrm{dice}} \mathcal{L}_{\mathrm{Dice}}\)$ - Training Budget: On ScanNet, pairs/quadruplets are sampled with visual overlap IoU within \([0.3, 0.8]\). Optimized across 8 NVIDIA RTX A6000 GPUs with a batch size of 64; convergence at 100 epochs requires merely ~3 hours. At test time, novel view inference executes in a single forward pass without any test-time optimization (TTO).

Key Experimental Results

Main Results

On the ScanNet indoor benchmark, the proposed method was evaluated against 2D segmentation upper bounds (Mask2Former, LSeg) and leading feed-forward 3D Gaussian-based joint NVS and scene understanding models (LSM, SIU3R). The quantitative results confirm unprecedented rendering fidelity alongside competitive panoptic parsing:

Method Paradigm Novel RGB PSNR (dB)↑ Novel RGB SSIM↑ Novel RGB LPIPS↓ Input View mIoU↑ Input View PQ↑ Novel View mIoU↑ Novel View PQ↑
Mask2Former (2022) 2D Upper Bound — — — 0.6186 0.5925 — —
LSeg (2022) 2D Open-Vocab — — — 0.3976 — — —
LSM (NeurIPS 2024) Joint 3DGS 20.96 0.7245 0.3176 0.2810 — 0.2707 —
SIU3R (NeurIPS 2025) Joint 3DGS 25.88 0.8220 0.1831 0.5899 0.6565 0.5894 0.6565
Ours Decoupled LVSM Prop. 33.56 0.9109 0.1149 0.6186 0.5949 0.5949 0.6092

Ablation Study on Design Alternatives

To examine the optimal integration topology between segmentation heads and large view synthesis models, the paper benchmarked four distinct architectural configurations:

Strategy Topological Flow NVS Weights PSNR (dB)↑ SSIM↑ LPIPS↓ Novel mIoU↑ Novel PQ↑ Key Mechanism & Failure Analysis
Joint Decoder Feat. Seg head on intermediate NVS layers Trainable 24.76 0.7502 0.2774 0.4146 0.4716 Severe multi-task gradient conflict degrades synthesis
Joint Decoder Feat. Seg head on intermediate NVS layers Frozen 33.56 0.9109 0.1149 0.2008 0.1816 Pure photometric features lack semantic discrimination
NVS \(\to\) Segmentation Render target RGB then apply 2D M2F Frozen 33.56 0.9109 0.1149 0.6239 0.5740 Uncorrelated per-view inference breaks instance tracking
Ours (Seg \(\to\) NVS) Source seg \(\to\) Binary propagation Frozen 33.56 0.9109 0.1149 0.5949 0.6092 Full rendering quality preserved + high instance consistency

Under challenging low-overlap regimes (depth-based IoU in \([0.01, 0.3]\)), SIU3R's PQ dropped sharply from 0.6565 to 0.5737, whereas the proposed method degraded far more gracefully from 0.6092 to 0.5726. Consequently, the PQ gap between SIU3R and Ours shrank from 0.0473 down to just 0.0011, while Ours maintained an overwhelming 5.14 dB PSNR margin (26.47 dB vs. 21.33 dB), proving that token-level attention correspondence extrapolates across occlusions more reliably than explicit Gaussian rasterization.

Modular Swapping & Cross-Dataset Transfer

When replacing the source-view segmenter with off-the-shelf PanSt3R and performing zero-shot evaluation on the unseen Replica dataset without any fine-tuning: - SIU3R: Severely overfit to ScanNet's geometry distribution, collapsing completely on Replica (PSNR 14.05 dB, SSIM 0.527, PQ 0.186, mIoU 0.171). - Ours (PanSt3R + Latent Pose): Maintained high generalization, achieving 23.52 dB PSNR, 0.394 PQ (+112% relative gain), and 0.454 mIoU (+165% relative gain).

Key Findings

  • Appearance-Agnostic Geometric Correspondence: Saliency gradient maps demonstrate that the attention layers of an RGB-pretrained LVSM track true 3D spatial correspondences regardless of whether the input is natural RGB or high-contrast binary masks.
  • Elimination of Multi-Task Interference: Decoupling the rendering backbone from semantic supervision completely averts negative gradient interference, preserving state-of-the-art 33.56 dB rendering fidelity.
  • Propagation Boundary Loss: The minor mIoU reduction from input views to novel target views (0.6186 \(\to\) 0.5949) is predominantly caused by threshold rejection around dis-occluded object silhouettes and sub-pixel boundary smoothing.

Highlights & Insights

  • Repurposing Foundation LVSMs for Scene Understanding: Successfully reframes large view synthesis models from mere neural image renderers into general-purpose geometric correspondence engines capable of propagating arbitrary view-independent dense features.
  • Robust Binary Channel Discretization: Elegant mathematical design where continuous boundary interpolation between \(\{0, 1\}\) naturally falls near \(0.5\), converting neural blending artifacts into safely filterable uncertainty zones rather than false-positive phantom classes.
  • Zero-Shot Modular Flexibility: Complete decoupling permits independent plug-and-play upgrades—swapping source segmenters (e.g., SAM2, PanSt3R) or underlying view synthesis foundations (Less3Depend, RayZer) without full system retraining.

Limitations & Future Work

  • Inference Compute Overhead: Every target viewpoint demands a full forward pass through the heavy transformer-based view synthesis model, creating latency bottlenecks when fast multi-view dense mapping is desired.
  • Constrained to View-Independent Signals: The formulation fundamentally relies on the invariance of semantic class and instance identity across viewpoints. View-dependent quantities (such as metric depth, surface normals, or specular highlights) cannot be directly propagated without introducing explicit camera coordinate transformations.
  • Multi-Pass Overhead for Dense Scenes: Because 3 binary channels accommodate up to 8 active instances in a single forward pass, scenes containing dozens of simultaneous instances necessitate multi-pass sequential propagation.
  • vs SIU3R / LSM (Feed-forward Joint 3DGS): SIU3R couples 3D Gaussian primitives with semantic attributes via joint rasterization. While achieving slightly higher in-domain PQ on dense scenes, its multi-task training compromises rendering quality (25.88 dB vs. 33.56 dB) and degrades severely under domain shifts; our decoupled approach offers dramatically higher synthesis fidelity and superior zero-shot cross-dataset transferability.
  • vs Panoptic Lifting / N2F2 (NeRF-based Optimization): NeRF lifting methods require dense posed captures and hours of per-scene optimization; this approach executes instantaneously in a single forward pass from as few as two unposed source images.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ First demonstration that RGB-only pretrained large view synthesis models inherently possess transferable geometric correspondence capable of propagating dense non-photorealistic panoptic labels.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarking covering novel view rendering, panoptic segmentation, extreme low-overlap stress testing, four architectural topologies, and zero-shot cross-dataset transfer.
  • Writing Quality: ⭐⭐⭐⭐⭐ Cohesive theoretical narrative tightly corroborated by gradient saliency attribution and empirical ablations.
  • Value: ⭐⭐⭐⭐⭐ Bridges the divide between generative view synthesis foundation models and 3D scene understanding, establishing an elegant, reconstruction-free paradigm for embodied AI perception.