Skip to content

SyncFix: Multi-View Consistent Diffusion Refinement of 3D Reconstructions

Conference: ECCV 2026
Paper: ECCV Official
Code: https://syncfix.github.io/
Area: 3D Vision
Keywords: 3D Gaussian Splatting / Latent Bridge Matching / Multi-View Consistency / Diffusion Refinement / Sparse-View Reconstruction

TL;DR

To eliminate cross-view geometric and semantic inconsistencies caused by independent 2D diffusion refinement in sparse-view 3D reconstruction, SyncFix reformulates multi-view repair as a joint latent bridge matching problem, using permutation-invariant cross-view attention to achieve pose-free synchronization and single-step deterministic inference with superior 3D consistency.

Background & Motivation

Under sparse-view input captures or off-trajectory novel viewpoints, neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) inevitably suffer from severe artifacts, including floaters, surface fragmentation, and blurred high-frequency textures. To suppress these degradations, recent works extensively leverage the expressive priors of 2D generative diffusion models (such as Nerfbusters and Difix3D+). These frameworks render novel views from corrupted 3D representations, apply 2D diffusion models to refine each view, and subsequently distill the denoised outputs back into the underlying 3D model.

However, marginal refinement methods like Difix3D+ fundamentally treat each rendered viewpoint as an isolated, independent 2D generation task. This independent paradigm violates the intrinsic physical constraint that all views are projections of the same underlying 3D scene. When encountering ambiguous structures or regions with large corruptions, single-view diffusion models tend to hallucinate mutually conflicting contents across viewpoints—for example, introducing a painting on a blank wall in one view while erasing it in another, or altering an object's geometry across viewpoints. When these contradictory signals are distilled back into the 3D model, optimization destabilizes, leading to severe appearance drift and structural degradation. While video diffusion-based approaches (such as 3DGS-Enhancer and GSFixer) mitigate inconsistencies via temporal attention, they remain constrained by fixed sequence lengths, heavy computational overhead, and an inability to handle wide-baseline, variable numbers of sparse viewpoints.

The angle of attack in this paper is to abandon the marginal generation assumption and directly internalize multi-view geometric and semantic constraints during the diffusion refinement process. Core idea: reformulate multi-view 3D repair as a joint latent bridge matching problem, defining deterministic probability flows across the product latent space and coupling views via permutation-invariant cross-view attention to achieve pose-free, single-step multi-view synchronization and consistent refinement.

Method

Overall Architecture

SyncFix takes as input a set of \(N\) rendered viewpoints with severe geometric corruptions \(X_D = \{x_D^{(1)}, \dots, x_D^{(N)}\}\) from a sparse 3DGS model, along with an optional clean reference viewpoint image \(x_{GT}^{ref}\). A frozen pretrained VAE encoder maps each view into latent space to construct joint distorted latents \(Z_D\) and clean target latents \(Z_{GT}\). Next, the framework establishes a continuous stochastic bridge path connecting the source and target distributions. A diffusion U-Net augmented with cross-view self-attention predicts a joint velocity field \(v_\theta\), enabling single-step deterministic ODE integration to synchronize all latents simultaneously. Finally, the predicted clean latents are decoded and supervised with joint image-space perceptual and structural objectives.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multi-View Corrupted Renderings Input<br/>X_D = {x_D^(1), ..., x_D^(N)}"] --> B["Multi-View Joint Latent Bridge Matching<br/>Construct deterministic bridge trajectory Z_t & velocity field"]
    B --> C["Permutation-Invariant Cross-View Attention Synchronization<br/>Global token-level self-attention for cross-view feature interaction"]
    C --> D["Image-Space Joint Perceptual & Structural Supervision<br/>Decoded multi-view supervision against hallucination drift"]
    D --> E["Single-Step Inference Output Consistent Refined Views<br/>X_GT* for stable 3D reconstruction distillation"]

Key Designs

1. Multi-View Joint Latent Bridge Matching: Shifting from Marginal Denoising to Joint Probability Flows

Standard single-view refinement methods assume that reconstruction artifacts resemble additive Gaussian noise, applying fixed-timestep noise injection and denoising. In reality, 3D reconstruction errors are spatially structured, view-dependent, and non-Gaussian. SyncFix discards the independent marginal mapping \(P(x_{GT}^{(i)} \mid x_D^{(i)})\) in favor of learning the joint conditional distribution \(P(X_{GT} \mid X_D)\) over the product latent space. Let \(Z_D = \{z_D^{(1)}, \dots, z_D^{(N)}\}\) and \(Z_{GT} = \{z_{GT}^{(1)}, \dots, z_{GT}^{(N)}\}\) denote the joint latent tensors. Latent bridge matching defines a continuous stochastic path for \(t \in [0, 1]\) directly transporting the degraded distribution to the clean target distribution:

\[Z_t = (1 - t) Z_D + t Z_{GT} + \sigma \sqrt{t(1 - t)} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I)\]

where \(\sigma\) modulates path stochasticity. The ideal velocity field \(v^*(Z_t, t) = Z_{GT} - Z_D\) directs the latents straight toward the clean manifold. The network \(v_\theta\) is trained via the flow matching objective:

\[\mathcal{L}_{\text{flow}} = \mathbb{E}_{t, Z_D, Z_{GT}} \|v_\theta(Z_t, t) - (Z_{GT} - Z_D)\|^2\]

By predicting the joint velocity field across paired viewpoints simultaneously, \(v_\theta\) couples the dynamics of all views, transforming the refinement process into optimal transport over the multi-view product manifold.

2. Permutation-Invariant Cross-View Attention Synchronization: Pose-Free Global Alignment in Latent Space

To facilitate cross-view communication without requiring camera poses, epipolar line constraints, or explicit depth estimation, SyncFix modifies the self-attention layers of the SDXL U-Net. Flattened spatial tokens from all views within a batch (along with the optional nearest clean reference view \(z_{GT}^{ref}\)) are concatenated along the token sequence dimension, allowing global multi-head self-attention while preserving within-view spatial positional encodings.

This architecture offers two decisive advantages: first, computing attention across concatenated tokens makes the operation strictly permutation-invariant. The network does not rely on a fixed camera ordering, allowing a model trained solely on view pairs (\(N=2\)) to generalize seamlessly to an arbitrary number of \(N\) views at inference time; second, because attention routing is content-driven, shared scene structures and high-frequency textures naturally align across views. Rich geometric cues from reference views are adaptively propagated to degraded regions across views, preventing view-specific hallucination discrepancies.

3. Image-Space Joint Perceptual & Structural Supervision: Multi-Scale Texture and Alignment Beyond Latents

Supervising the velocity field purely within the latent space (as in FlowR) often leads to blurry reconstructions and suboptimal high-frequency recovery. SyncFix keeps the VAE encoder and decoder frozen, decodes the single-step predicted latent \(Z_{GT}^* = Z_D + v_\theta(Z_D, 0)\) into image space \(\hat{X}\), and directly optimizes explicit multi-view objectives against ground-truth clean images \(X_{GT}\):

\[\mathcal{L} = \mathcal{L}_{\text{flow}} + \mathcal{L}_{\text{pixel}} + \lambda_{\text{lpips}} \mathcal{L}_{\text{LPIPS}} + \lambda_{\text{Gram}} \mathcal{L}_{\text{Gram}}\]

where \(\mathcal{L}_{\text{pixel}} = \|\hat{X} - X_{GT}\|_1\) enforces rigid spatial alignment, \(\mathcal{L}_{\text{LPIPS}}\) penalizes deep perceptual discrepancy, and the style Gram matrix loss \(\mathcal{L}_{\text{Gram}} = \|G(\phi(\hat{X})) - G(\phi(X_{GT}))\|_F^2\) matches texture correlations and material statistics. This comprehensive image-space supervision ensures crisp high frequencies and cross-view photometric harmony.

Loss & Training

SyncFix is built upon the SDXL foundation model. The VAE encoder and decoder are frozen while fine-tuning the cross-view attention-augmented diffusion U-Net. Training is conducted on 4 NVIDIA H100 GPUs for 100,000 steps with a batch size of 2 (each batch containing 2 degraded view pairs), using the Adam optimizer with a learning rate of \(4 \times 10^{-5}\). Loss weights are set to \(\lambda_{\text{lpips}} = 10\) and \(\lambda_{\text{Gram}} = 0.1\). At inference time, the model executes a single-step ODE solve from \(t=0\), completing refinement for an entire scene of 63 views in just 5.63 seconds.

Key Experimental Results

Main Results

To rigorously evaluate cross-view semantic and geometric agreement, the paper introduces the Cross-View Semantic Consistency (CVSC) metric: keypoint correspondences are detected across adjacent views using RaCo and LightGlue, geometrically verified via RANSAC estimation of the fundamental matrix to remove outliers, and evaluated by computing the cosine similarity between matched DINOv3 dense semantic features. The table below presents quantitative comparisons on DL3DV and NeRFBusters test sets († denotes without clean reference views):

Dataset Method PSNR ↑ LPIPS ↓ DreamSim ↓ FID ↓ CVSC ↑ Runtime (s) ↓
DL3DV 3DGS (Corrupted Baseline) 15.94 0.454 0.297 80.8 0.875 -
DL3DV Fixer (Single-View, No Ref) 15.73 0.466 0.257 45.5 0.822 15.8
DL3DV Difix3D+ † (No Ref) 16.05 0.368 0.171 26.2 0.819 30.6
DL3DV Difix3D+ (With Ref) 16.17 0.343 0.135 16.7 0.850 30.6
DL3DV FlowR (Flow Matching Multi-View) 16.32 0.369 0.138 19.7 0.832 109.8
DL3DV SyncFix † (Ours, No Ref) 16.45 0.334 0.129 22.5 0.862 5.63
DL3DV SyncFix (Ours, Full Model) 16.94 0.305 0.099 17.5 0.880 5.63
NeRFBusters 3DGS (Corrupted Baseline) 11.85 0.575 0.336 139.4 0.951 -
NeRFBusters Fixer 12.32 0.588 0.296 126.7 0.817 15.8
NeRFBusters Difix3D+ † (No Ref) 11.94 0.528 0.202 97.4 0.925 30.6
NeRFBusters Difix3D+ (With Ref) 12.05 0.496 0.174 81.1 0.931 30.6
NeRFBusters FlowR 13.09 0.560 0.231 86.0 0.916 109.8
NeRFBusters SyncFix † (Ours, No Ref) 14.07 0.517 0.183 86.2 0.937 5.63
NeRFBusters SyncFix (Ours, Full Model) 14.34 0.494 0.173 80.4 0.940 5.63

When generalized to feed-forward 3D Gaussian Splatting (AnySplat) on Mip-NeRF 360, SyncFix demonstrates robust zero-shot refinement capability:

Dataset Method PSNR ↑ (Grayed) LPIPS ↓ DreamSim ↓ CVSC ↑
Mip-NeRF 360 AnySplat (Feedforward Baseline) (13.58) 0.454 0.232 0.649
Mip-NeRF 360 Difix3D+ (13.41) 0.448 0.196 0.589
Mip-NeRF 360 SyncFix (Ours) (13.51) 0.430 0.189 0.686

Ablation Study

Systematic ablations on the DL3DV test set evaluate the impact of joint training, multi-view inference, and clean reference conditioning:

Config PSNR ↑ LPIPS ↓ DSSIM ↓ FID ↓ CVSC ↑ Note
3DGS Baseline 15.94 0.454 0.297 80.8 0.875 Raw sparse 3DGS renderings
Difix3D (train-in-loop) 16.14 0.445 0.285 76.6 0.853 Early marginal refinement
Difix3D+ (single-view train) 16.17 0.343 0.135 16.7 0.850 Default single-view baseline
Difix3D+ (multi-view inference) 16.08 0.370 0.163 29.28 0.827 Supplying multi-view at test time degrades quality
SyncFix (single-view, w/o ref) 16.26 0.347 0.142 22.2 0.821 Validates latent bridge matching formulation
SyncFix (multi-view, w/o ref) 16.45 0.334 0.129 22.5 0.862 Cross-view attention boosts CVSC by +0.041
SyncFix (multi-view, with ref) 16.94 0.305 0.099 17.5 0.880 Full model SOTA, 27% lower DreamSim

Key Findings

  • Joint multi-view training is irreplaceable: Simply feeding multiple views at test time into a model trained on single views (Difix3D+) harms performance (CVSC drops from 0.850 to 0.827), proving that multi-view consensus cannot be retrofitted post-hoc.
  • The raw 3DGS CVSC paradox: Unrefined 3DGS shows artificially high CVSC (0.875) because physical floaters and blur are geometrically continuous and appear identical across viewpoints despite poor visual quality. Independent refinement (Difix3D+) introduces disparate hallucinations that ruin consistency (CVSC drops to 0.819/0.850). SyncFix dramatically improves perceptual fidelity (DreamSim drops to 0.099) while maintaining top semantic consistency (CVSC 0.880).
  • Substantial inference speedup: SyncFix resolves the learned ODE in a single deterministic step, refining 63 views in 5.63 seconds—over \(5.4\times\) faster than Difix3D+ (30.6s) and nearly \(20\times\) faster than FlowR (109.8s).

Highlights & Insights

  • From Gaussian denoising to continuous optimal transport: Reformulating 3D refinement as Brownian bridge interpolation allows direct mapping from corrupted distributions to clean distributions, eliminating iterative sampling steps and ungrounded Gaussian assumptions.
  • Permutation invariance scales across view counts: Decoupling viewpoint order through concatenated global attention allows cheap view-pair (\(N=2\)) training while generalizing to arbitrary view counts at inference, where more viewpoints naturally provide tighter consensus.
  • Emergent pose-free geometric consistency: Without receiving explicit camera matrices, epipolar lines, or depth loss, cross-view attention autonomously discovers correspondences in latent space, proving that pretrained diffusion representations contain sufficient implicit 3D grounding.

Limitations & Future Work

  • Reliance on visual overlap: SyncFix assumes sufficient overlap between viewpoints to form reliable attention links. Under extremely wide baselines with nearly disjoint camera frustums, the synchronization signal diminishes, causing the model to revert toward independent hallucinations.
  • Plausible but unfaithful hallucination under catastrophic loss: When underlying 3DGS models suffer from massive structural absence, the generative prior synthesizes visually coherent geometries that may deviate from ground-truth reality.
  • Future directions: Integrating coarse relative camera pose conditioning or end-to-end differentiable reprojection loops into latent bridge matching to strengthen geometric rigidity under extreme sparse-view regimes.
  • vs Difix3D+: Difix3D+ uses single-step diffusion filtering independently per view and relies on 3DGS optimization to harmonize discrepancies, resulting in geometry drift and instability; SyncFix couples views directly inside the latent bridge matching process via cross-view attention.
  • vs FlowR: FlowR applies multi-view flow matching with latent-only velocity loss, resulting in blurred fine textures and slow convergence; SyncFix incorporates comprehensive image-space perceptual and Gram-matrix losses with single-step inference for crisper details and faster runtime.
  • vs 3DGS-Enhancer / GSFixer: Video diffusion methods enforce temporal attention along smooth camera trajectories, suffering from high computational cost and poor scaling to unordered wide-baseline sets; SyncFix uses permutation-invariant attention, natively handling arbitrary view configurations.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering integration of latent bridge matching and permutation-invariant multi-view attention for pose-free 3D refinement]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across DL3DV, NeRFBusters, and AnySplat, supported by the newly designed CVSC metric and thorough ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous formulation, and tight alignment between architectural diagrams and narrative]
  • Value: ⭐⭐⭐⭐⭐ [Resolves the critical bottleneck of conflicting multi-view hallucinations during 2D diffusion-to-3D distillation with real-time inference speeds]