Skip to content

Latent Fusion: Decoding Consolidated 3D Geometry from Feed-forward Geometry Transformer Latents

Conference: ECCV 2026
Paper: ECCV Official Link
Code: https://lorafib.github.io/fus3d/
Area: 3D Vision
Keywords: geometry transformer, Signed Distance Field (SDF), voxel query, feed-forward 3D reconstruction, multi-view feature fusion

TL;DR

This paper introduces Fus3D, which discards the conventional predict-then-fuse paradigm in feed-forward geometry transformers and instead leverages learned canonical voxel queries with interleaved cross/self-attention to decode dense Signed Distance Fields (SDF) directly from intermediate multi-view latents (VGGT), achieving pose-free, artifact-free, and complete 3D surface reconstruction.

Background & Motivation

Reconstructing high-fidelity 3D geometry from uncalibrated multi-view image collections or sparse video sequences remains a foundational pillar for spatial perception, scene interaction, and robotics. Recently, feed-forward multi-view geometry transformers (FFGT) such as DUSt3R, VGGT, and DA3 have driven rapid advancements in the field. Pretrained on massive datasets, these models acquire rich global and spatial geometry priors, enabling direct feed-forward estimation of per-image depth maps, point clouds, and camera parameters without relying on classical feature matching and explicit camera calibration.

However, existing feed-forward pipelines suffer from a fundamental architectural bottleneck when assembling a coherent 3D scene: they almost universally adhere to a "predict-then-fuse" paradigm, routing transformer representations through 2D prediction heads before merging per-view outputs via post-hoc point cloud registration, TSDF fusion, or Poisson reconstruction. This two-stage separation incurs severe dual failure modes. Under sparse-view regimes, the learned global multi-view prior is constrained locally within individual image planes, leaving unobserved regions underspecified and resulting in extensive holes and severe incompleteness. Conversely, in dense multi-view settings, minor inter-view prediction discrepancies accumulate across geometric integration space, introducing high-frequency noise that degrades surface reconstruction fidelity at scale.

The key entry point to resolving this dilemma lies in reconsidering the transformer's internal representations: the intermediate layers of FFGTs have already assembled an expressive, joint 3D world representation, which is discarded when forced into 2D per-view output heads. The core idea of this paper is to bypass per-view predictions and post-hoc fusion entirely, using position-conditioned canonical voxel queries to attend directly into the geometry transformer's intermediate latent space, lifting multi-view features into a dense 3D latent grid that regresses a continuous Signed Distance Field (SDF).

Method

Overall Architecture

Fus3D consists of a three-stage feed-forward pipeline: first, a pretrained multi-view geometry transformer backbone (VGGT) extracts multi-stage 2D joint geometric feature maps from unposed input views; second, a 3D extraction transformer (\(E\)) conditioned on learned canonical voxel queries alternates between 2D-to-3D cross-attention and 3D volumetric self-attention to progressively aggregate multi-view representations into a structured 3D latent grid; third, a lightweight 3D convolutional decoder head (\(H_{3D}\)) maps this latent grid directly into a dense SDF volume.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Unposed Multi-view Image Collection"] --> B["Multi-view Geometry Transformer Backbone<br/>Extract 2D Multi-stage Intermediate Features from VGGT"]
    C["3D-Position-Conditioned Canonical Embeddings<br/>16³ Latent Voxel Spatial Coordinate Projection"] --> D["Interleaved Cross- and Self-Attention<br/>2D-to-3D Lifting and Spatial Propagation"]
    B --> D
    D --> E["Dense Volumetric Latent Grid<br/>Consolidates Global Joint Geometry Prior"]
    E --> F["Lightweight Convolutional Decoder Head<br/>Direct Upsampling to 64³ Dense SDF Grid"]
    F --> G["Validity-Aware SDF Supervision<br/>Masking Non-watertight Meshes & Sign Ambiguity"]
    G --> H["Consolidated 3D Geometry Surface"]

Key Designs

1. 3D-Position-Conditioned Canonical Embeddings: Bridging the 2D-to-3D Dimensionality Gap Feed-forward geometry transformers produce flattened sequences of 2D patch tokens, whereas target geometry is an explicit spatial field. Directly mapping unstructured 2D latents into a 3D grid risks severe spatial misalignment and overfitting. Drawing inspiration from continuous neural spatial memory, Fus3D introduces scene-agnostic learned canonical embeddings \(z_{3D}\) (\(16^3\) spatial resolution with feature dimension \(d=2048\)). These embeddings are conditioned on their corresponding 3D normalized voxel coordinates within the spatial domain \(\Omega\) via linear projection. This design establishes an initial 3D memory state that acts as a projection-free spatial anchor, lifting features via attention without requiring explicit camera intrinsics or ray-based marching.

2. Interleaved Cross- and Self-Attention: Progressive Multi-Scale Feature Lifting To ingest representations across multiple layers of abstraction from the backbone, the 3D extraction transformer \(E\) employs a hierarchical interleaved architecture. The canonical voxel embeddings serve as Queries, while 2D feature maps \(z^{2D}_b\) extracted from the \(b\)-th intermediate stage of VGGT (\(B=4\)) provide Keys and Values. A 2D-to-3D cross-attention block allows each voxel to attend dynamically across all input views, absorbing relevant geometric evidence; this is immediately paired with a 3D global self-attention block that diffuses information across the volume \(\Omega\) for spatial consistency. This four-stage sequence is repeated twice (forming 8 transformer blocks in total), allowing the resulting latent volume \(\hat{z}_{3D}\) to encode object-level symmetry, structural completion, and category-level shape priors.

3. Lightweight Convolutional Decoder Head & Region-of-Interest Domain Querying: Dense Field Regression Given the updated \(16^3 \times 2048\) volumetric latent grid \(\hat{z}_{3D}\), Fus3D uses a fully convolutional upsampling decoder \(H_{3D}\) to expand the feature grid into a dense \(64^3\) SDF grid \(\hat{f}_{X_\Omega}\). Unlike coordinate-based MLPs that require millions of dense point queries, 3D convolutions exploit spatial inductive bias and preserve smooth topology while completing inference in under 3 seconds. Furthermore, by rescaling and shifting the normalized coordinate queries within \(\Omega\), the extended model variant Fus3D+ supports arbitrary bounding volume queries for flexible zooming and cropping of regions of interest.

4. Validity-Aware SDF Supervision: Graceful Handling of Non-Watertight Meshes Large-scale 3D asset datasets (such as Objaverse and DTU) frequently contain non-watertight meshes, internal face self-intersections, and unobserved cavities. In such cases, inside/outside definitions become ill-posed, and naive SDF losses induce destructive sign flips. Fus3D resolves this by computing two spatial masks over \(X_\Omega\): a validity mask \(M_{val}\) identifying regions where the target ground truth is well-defined, and an Eikonal mask \(M_{eik}\) computed by evaluating the Eikonal term on the ground-truth field itself. Voxel neighborhoods where the ground-truth gradient magnitude severely deviates from 1 are classified as unreliable and gracefully downgraded to an unsigned distance loss, effectively unlocking scalable supervision on imperfect real-world 3D assets.

Loss & Training

Training proceeds in two curriculum stages. The first stage supervises a low-resolution (\(16^3\)) intermediate prediction via a single linear decoder using the base SDF loss \(\mathcal{L}_{SDF}\) and camera loss \(\mathcal{L}_{cam}\). The second stage introduces the full convolutional decoder \(H_{3D}\) to regress the \(64^3\) SDF volume under the comprehensive objective:

\[\mathcal{L}_{tot} = \lambda_s \mathcal{L}_{SDF}(X_S) + \mathcal{L}_{SDF}(X_\Omega) + \lambda_\nabla \mathcal{L}_{\nabla}(X_\Omega) + \lambda_{eik} \mathcal{L}_{eik}(X_\Omega) + \lambda_{cam} \mathcal{L}_{cam}\]

Here \(X_S\) denotes dense surface samples near the boundary \(\partial S\), and \(X_\Omega\) denotes the regular grid centers. The objective incorporates an L1 distance loss, a finite-difference gradient regularizer \(\mathcal{L}_\nabla\), and the Eikonal penalty \(\mathcal{L}_{eik}\). The model is trained using AdamW with cosine learning rate scheduling and lightweight LoRA fine-tuning on the geometry backbone.

Key Experimental Results

Main Results

On the forward-facing DTU benchmark under a fully feed-forward, pose-free setup, Fus3D is compared against generalizable implicit SDF baselines (VolRecon and UFORecon, fed with VGGT-predicted poses). Comparisons are divided into favorable overlapping view sets (views 23, 24, 33) and unfavorable wide-baseline sets (views 1, 16, 36).

Method Camera Pose Source Favorable \(CD \downarrow\) Favorable \(D_{GT \to P} \downarrow\) Favorable \(D_{P \to GT} \downarrow\) Unfavorable \(CD \downarrow\) Unfavorable \(D_{GT \to P} \downarrow\) Unfavorable \(D_{P \to GT} \downarrow\)
VolRecon COLMAP (†) 2.905 2.331 3.478 5.781 3.811 7.751
UFORecon COLMAP (†) 2.771 2.101 3.440 3.275 2.252 4.298
VolRecon Predicted (Feed-forward) 3.678 3.369 3.987 8.878 5.869 11.887
UFORecon Predicted (Feed-forward) 3.415 2.989 3.841 4.942 3.737 6.146
Fus3D (Ours) Pose-free (End-to-end) 2.432 1.804 3.059 3.525 2.814 4.236

On 170 test scenes from Objaverse evaluated at 8 input views, Fus3D is benchmarked against fine-tuned VGGT coupled with post-hoc surface reconstruction baselines (TSDF fusion and Poisson surface reconstruction):

Method \(F_\epsilon \uparrow\) \(F_{0.5\epsilon} \uparrow\) \(CD \downarrow\) \(EMD \downarrow\) \(SDF_{MAE} \downarrow\)
Fus3D (Ours) 0.83 0.51 0.021 0.019 0.004
\(\text{VGGT}_{ft} + \text{TSDF}\) 0.80 0.43 0.022 0.014 0.023
\(\text{VGGT}_{ft} + \text{Poisson}\) 0.70 0.38 0.044 0.052 –
\(\text{VGGT (default)} + \text{TSDF}\) 0.35 0.12 0.080 0.082 0.136
\(\text{VGGT (default)} + \text{Poisson}\) 0.29 0.10 0.121 0.788 –

Ablation Study

Ablation analysis on the spatial query mechanism and flexible region-of-interest domain selection (Variable \(\Omega\) vs. Default \(\Omega\)):

Config Variable \(\Omega\) \(F_{0.5\epsilon} \uparrow\) Variable \(\Omega\) \(CD \downarrow\) Default \(\Omega\) \(F_{0.5\epsilon} \uparrow\) Default \(\Omega\) \(CD \downarrow\) Note
Fus3D+ (Full Spatial Query) 0.281 0.018 0.467 0.021 Learned canonical embeddings + 3D position encoding
Fus3D+ w/o Q (Linear Position Only) 0.274 0.021 0.466 0.022 Query replaced entirely by projected coordinates

Key Findings

  • In uncalibrated feed-forward evaluations on DTU, Fus3D improves Chamfer Distance by 28.7% over UFORecon in favorable view setups (2.432 vs. 3.415) and by 28.6% in unfavorable view setups (3.525 vs. 4.942), matching or exceeding UFORecon using offline COLMAP poses (3.275), demonstrating that volumetric latent lifting effectively offsets the absence of camera calibration.
  • Across input view scaling from 2 to 24 views on Objaverse, Fus3D exhibits excellent scaling behavior: under sparse 2-view settings, it achieves plausible shape completion over occluded areas; at 24 views, it avoids the error accumulation that plagues TSDF fusion, achieving an \(SDF_{MAE}\) of 0.004 compared to 0.023 for \(\text{VGGT}_{ft} + \text{TSDF}\).
  • Principal Component Analysis (PCA) on the extracted latent grid \(\hat{z}_{3D}\) demonstrates that the principal components smoothly follow the spatial zero-crossing of the SDF across varied object poses and align semantically with shared geometric parts (e.g., wheels, chassis, headlights) across object instances within the same category.

Highlights & Insights

  • Paradigm Shift in 3D Reconstruction: Identifies the inherent trade-off of the predict-then-fuse paradigm and pioneers direct extraction of continuous 3D implicit fields from the latent space of multi-view transformers.
  • Projection-Free 2D-to-3D Lifting: Employs position-conditioned canonical voxel queries with interleaved cross/self-attention, entirely bypassing explicit ray-sampling overhead and sensitivity to calibration errors.
  • Validity-Aware Decoupled Supervision: Uses an intrinsic Eikonal validity mask to gracefully degrade the loss to an unsigned distance variant over non-watertight surfaces, enabling scalable training on imperfect real-world assets.

Limitations & Future Work

  • Volumetric Resolution Constraints: Bounded by the cubic memory complexity of dense 3D grids, the current \(64^3\) resolution limits high-frequency surface detail, warranting future integration with sparse octrees or multi-scale upsampling.
  • Predefined Region of Interest: The canonical voxel grid assumes a known bounding volume relative to the reference view; scaling to unbounded large-scale indoor or outdoor environments remains an open challenge.
  • Unidirectional Feature Extraction: The extraction transformer currently operates as a read-only consumer of backbone features; closing the loop by providing 3D-aware feedback into the 2D backbone could enable persistent interactive spatial understanding.
  • vs. VGGT / DUSt3R: While VGGT and DUSt3R predict 2D depth maps and point clouds per view that require external TSDF or Poisson post-hoc fusion, Fus3D treats the transformer purely as a geometric feature extractor and directly outputs a unified, complete 3D surface.
  • vs. VolRecon / UFORecon: Both require accurate camera parameters to build cost volumes; Fus3D operates in a completely pose-free, feed-forward manner while delivering superior surface completeness and faster inference.
  • vs. LRM / MeshLRM: Generative 3D transformers rely on 3D-to-3D autoencoders trained on massive, clean 3D datasets; Fus3D shifts the primary burden of representation learning to scalable 2D feed-forward models, utilizing 3D supervision only for latent extraction.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ First framework to decode dense 3D Signed Distance Fields directly from intermediate multi-view geometry transformer latents via learned canonical voxel queries.
  • Experimental Thoroughness: ⭐⭐⭐⭐☆ Rigorous evaluation across DTU and Objaverse with sparse/dense scaling analyses, ablation on queries, and PCA geometric interpretability.
  • Writing Quality: ⭐⭐⭐⭐⭐ Lucid formulation of the predict-then-fuse failure modes paired with clean pipeline diagrams and mathematical rigor.
  • Value: ⭐⭐⭐⭐⭐ Sets a compelling precedent for end-to-end continuous 3D field extraction from large-scale multi-view foundation models.