SyncFix: Multi-View Consistent Diffusion Refinement of 3D Reconstructions¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://syncfix.github.io/
Area: 3D Vision
Keywords: 3D Gaussian Splatting / Latent Bridge Matching / Multi-View Consistency / Diffusion Refinement / Sparse-View Reconstruction
TL;DR¶
To eliminate cross-view geometric and semantic inconsistencies caused by independent 2D diffusion refinement in sparse-view 3D reconstruction, SyncFix reformulates multi-view repair as a joint latent bridge matching problem, using permutation-invariant cross-view attention to achieve pose-free synchronization and single-step deterministic inference with superior 3D consistency.
Background & Motivation¶
Under sparse-view input captures or off-trajectory novel viewpoints, neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) inevitably suffer from severe artifacts, including floaters, surface fragmentation, and blurred high-frequency textures. To suppress these degradations, recent works extensively leverage the expressive priors of 2D generative diffusion models (such as Nerfbusters and Difix3D+). These frameworks render novel views from corrupted 3D representations, apply 2D diffusion models to refine each view, and subsequently distill the denoised outputs back into the underlying 3D model.
However, marginal refinement methods like Difix3D+ fundamentally treat each rendered viewpoint as an isolated, independent 2D generation task. This independent paradigm violates the intrinsic physical constraint that all views are projections of the same underlying 3D scene. When encountering ambiguous structures or regions with large corruptions, single-view diffusion models tend to hallucinate mutually conflicting contents across viewpoints—for example, introducing a painting on a blank wall in one view while erasing it in another, or altering an object's geometry across viewpoints. When these contradictory signals are distilled back into the 3D model, optimization destabilizes, leading to severe appearance drift and structural degradation. While video diffusion-based approaches (such as 3DGS-Enhancer and GSFixer) mitigate inconsistencies via temporal attention, they remain constrained by fixed sequence lengths, heavy computational overhead, and an inability to handle wide-baseline, variable numbers of sparse viewpoints.
The angle of attack in this paper is to abandon the marginal generation assumption and directly internalize multi-view geometric and semantic constraints during the diffusion refinement process. Core idea: reformulate multi-view 3D repair as a joint latent bridge matching problem, defining deterministic probability flows across the product latent space and coupling views via permutation-invariant cross-view attention to achieve pose-free, single-step multi-view synchronization and consistent refinement.
Method¶
Overall Architecture¶
SyncFix takes as input a set of \(N\) rendered viewpoints with severe geometric corruptions \(X_D = \{x_D^{(1)}, \dots, x_D^{(N)}\}\) from a sparse 3DGS model, along with an optional clean reference viewpoint image \(x_{GT}^{ref}\). A frozen pretrained VAE encoder maps each view into latent space to construct joint distorted latents \(Z_D\) and clean target latents \(Z_{GT}\). Next, the framework establishes a continuous stochastic bridge path connecting the source and target distributions. A diffusion U-Net augmented with cross-view self-attention predicts a joint velocity field \(v_\theta\), enabling single-step deterministic ODE integration to synchronize all latents simultaneously. Finally, the predicted clean latents are decoded and supervised with joint image-space perceptual and structural objectives.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-View Corrupted Renderings Input<br/>X_D = {x_D^(1), ..., x_D^(N)}"] --> B["Multi-View Joint Latent Bridge Matching<br/>Construct deterministic bridge trajectory Z_t & velocity field"]
B --> C["Permutation-Invariant Cross-View Attention Synchronization<br/>Global token-level self-attention for cross-view feature interaction"]
C --> D["Image-Space Joint Perceptual & Structural Supervision<br/>Decoded multi-view supervision against hallucination drift"]
D --> E["Single-Step Inference Output Consistent Refined Views<br/>X_GT* for stable 3D reconstruction distillation"]
Key Designs¶
1. Multi-View Joint Latent Bridge Matching: Shifting from Marginal Denoising to Joint Probability Flows
Standard single-view refinement methods assume that reconstruction artifacts resemble additive Gaussian noise, applying fixed-timestep noise injection and denoising. In reality, 3D reconstruction errors are spatially structured, view-dependent, and non-Gaussian. SyncFix discards the independent marginal mapping \(P(x_{GT}^{(i)} \mid x_D^{(i)})\) in favor of learning the joint conditional distribution \(P(X_{GT} \mid X_D)\) over the product latent space. Let \(Z_D = \{z_D^{(1)}, \dots, z_D^{(N)}\}\) and \(Z_{GT} = \{z_{GT}^{(1)}, \dots, z_{GT}^{(N)}\}\) denote the joint latent tensors. Latent bridge matching defines a continuous stochastic path for \(t \in [0, 1]\) directly transporting the degraded distribution to the clean target distribution:
where \(\sigma\) modulates path stochasticity. The ideal velocity field \(v^*(Z_t, t) = Z_{GT} - Z_D\) directs the latents straight toward the clean manifold. The network \(v_\theta\) is trained via the flow matching objective:
By predicting the joint velocity field across paired viewpoints simultaneously, \(v_\theta\) couples the dynamics of all views, transforming the refinement process into optimal transport over the multi-view product manifold.
2. Permutation-Invariant Cross-View Attention Synchronization: Pose-Free Global Alignment in Latent Space
To facilitate cross-view communication without requiring camera poses, epipolar line constraints, or explicit depth estimation, SyncFix modifies the self-attention layers of the SDXL U-Net. Flattened spatial tokens from all views within a batch (along with the optional nearest clean reference view \(z_{GT}^{ref}\)) are concatenated along the token sequence dimension, allowing global multi-head self-attention while preserving within-view spatial positional encodings.
This architecture offers two decisive advantages: first, computing attention across concatenated tokens makes the operation strictly permutation-invariant. The network does not rely on a fixed camera ordering, allowing a model trained solely on view pairs (\(N=2\)) to generalize seamlessly to an arbitrary number of \(N\) views at inference time; second, because attention routing is content-driven, shared scene structures and high-frequency textures naturally align across views. Rich geometric cues from reference views are adaptively propagated to degraded regions across views, preventing view-specific hallucination discrepancies.
3. Image-Space Joint Perceptual & Structural Supervision: Multi-Scale Texture and Alignment Beyond Latents
Supervising the velocity field purely within the latent space (as in FlowR) often leads to blurry reconstructions and suboptimal high-frequency recovery. SyncFix keeps the VAE encoder and decoder frozen, decodes the single-step predicted latent \(Z_{GT}^* = Z_D + v_\theta(Z_D, 0)\) into image space \(\hat{X}\), and directly optimizes explicit multi-view objectives against ground-truth clean images \(X_{GT}\):
where \(\mathcal{L}_{\text{pixel}} = \|\hat{X} - X_{GT}\|_1\) enforces rigid spatial alignment, \(\mathcal{L}_{\text{LPIPS}}\) penalizes deep perceptual discrepancy, and the style Gram matrix loss \(\mathcal{L}_{\text{Gram}} = \|G(\phi(\hat{X})) - G(\phi(X_{GT}))\|_F^2\) matches texture correlations and material statistics. This comprehensive image-space supervision ensures crisp high frequencies and cross-view photometric harmony.
Loss & Training¶
SyncFix is built upon the SDXL foundation model. The VAE encoder and decoder are frozen while fine-tuning the cross-view attention-augmented diffusion U-Net. Training is conducted on 4 NVIDIA H100 GPUs for 100,000 steps with a batch size of 2 (each batch containing 2 degraded view pairs), using the Adam optimizer with a learning rate of \(4 \times 10^{-5}\). Loss weights are set to \(\lambda_{\text{lpips}} = 10\) and \(\lambda_{\text{Gram}} = 0.1\). At inference time, the model executes a single-step ODE solve from \(t=0\), completing refinement for an entire scene of 63 views in just 5.63 seconds.
Key Experimental Results¶
Main Results¶
To rigorously evaluate cross-view semantic and geometric agreement, the paper introduces the Cross-View Semantic Consistency (CVSC) metric: keypoint correspondences are detected across adjacent views using RaCo and LightGlue, geometrically verified via RANSAC estimation of the fundamental matrix to remove outliers, and evaluated by computing the cosine similarity between matched DINOv3 dense semantic features. The table below presents quantitative comparisons on DL3DV and NeRFBusters test sets († denotes without clean reference views):
| Dataset | Method | PSNR ↑ | LPIPS ↓ | DreamSim ↓ | FID ↓ | CVSC ↑ | Runtime (s) ↓ |
|---|---|---|---|---|---|---|---|
| DL3DV | 3DGS (Corrupted Baseline) | 15.94 | 0.454 | 0.297 | 80.8 | 0.875 | - |
| DL3DV | Fixer (Single-View, No Ref) | 15.73 | 0.466 | 0.257 | 45.5 | 0.822 | 15.8 |
| DL3DV | Difix3D+ † (No Ref) | 16.05 | 0.368 | 0.171 | 26.2 | 0.819 | 30.6 |
| DL3DV | Difix3D+ (With Ref) | 16.17 | 0.343 | 0.135 | 16.7 | 0.850 | 30.6 |
| DL3DV | FlowR (Flow Matching Multi-View) | 16.32 | 0.369 | 0.138 | 19.7 | 0.832 | 109.8 |
| DL3DV | SyncFix † (Ours, No Ref) | 16.45 | 0.334 | 0.129 | 22.5 | 0.862 | 5.63 |
| DL3DV | SyncFix (Ours, Full Model) | 16.94 | 0.305 | 0.099 | 17.5 | 0.880 | 5.63 |
| NeRFBusters | 3DGS (Corrupted Baseline) | 11.85 | 0.575 | 0.336 | 139.4 | 0.951 | - |
| NeRFBusters | Fixer | 12.32 | 0.588 | 0.296 | 126.7 | 0.817 | 15.8 |
| NeRFBusters | Difix3D+ † (No Ref) | 11.94 | 0.528 | 0.202 | 97.4 | 0.925 | 30.6 |
| NeRFBusters | Difix3D+ (With Ref) | 12.05 | 0.496 | 0.174 | 81.1 | 0.931 | 30.6 |
| NeRFBusters | FlowR | 13.09 | 0.560 | 0.231 | 86.0 | 0.916 | 109.8 |
| NeRFBusters | SyncFix † (Ours, No Ref) | 14.07 | 0.517 | 0.183 | 86.2 | 0.937 | 5.63 |
| NeRFBusters | SyncFix (Ours, Full Model) | 14.34 | 0.494 | 0.173 | 80.4 | 0.940 | 5.63 |
When generalized to feed-forward 3D Gaussian Splatting (AnySplat) on Mip-NeRF 360, SyncFix demonstrates robust zero-shot refinement capability:
| Dataset | Method | PSNR ↑ (Grayed) | LPIPS ↓ | DreamSim ↓ | CVSC ↑ |
|---|---|---|---|---|---|
| Mip-NeRF 360 | AnySplat (Feedforward Baseline) | (13.58) | 0.454 | 0.232 | 0.649 |
| Mip-NeRF 360 | Difix3D+ | (13.41) | 0.448 | 0.196 | 0.589 |
| Mip-NeRF 360 | SyncFix (Ours) | (13.51) | 0.430 | 0.189 | 0.686 |
Ablation Study¶
Systematic ablations on the DL3DV test set evaluate the impact of joint training, multi-view inference, and clean reference conditioning:
| Config | PSNR ↑ | LPIPS ↓ | DSSIM ↓ | FID ↓ | CVSC ↑ | Note |
|---|---|---|---|---|---|---|
| 3DGS Baseline | 15.94 | 0.454 | 0.297 | 80.8 | 0.875 | Raw sparse 3DGS renderings |
| Difix3D (train-in-loop) | 16.14 | 0.445 | 0.285 | 76.6 | 0.853 | Early marginal refinement |
| Difix3D+ (single-view train) | 16.17 | 0.343 | 0.135 | 16.7 | 0.850 | Default single-view baseline |
| Difix3D+ (multi-view inference) | 16.08 | 0.370 | 0.163 | 29.28 | 0.827 | Supplying multi-view at test time degrades quality |
| SyncFix (single-view, w/o ref) | 16.26 | 0.347 | 0.142 | 22.2 | 0.821 | Validates latent bridge matching formulation |
| SyncFix (multi-view, w/o ref) | 16.45 | 0.334 | 0.129 | 22.5 | 0.862 | Cross-view attention boosts CVSC by +0.041 |
| SyncFix (multi-view, with ref) | 16.94 | 0.305 | 0.099 | 17.5 | 0.880 | Full model SOTA, 27% lower DreamSim |
Key Findings¶
- Joint multi-view training is irreplaceable: Simply feeding multiple views at test time into a model trained on single views (Difix3D+) harms performance (CVSC drops from 0.850 to 0.827), proving that multi-view consensus cannot be retrofitted post-hoc.
- The raw 3DGS CVSC paradox: Unrefined 3DGS shows artificially high CVSC (0.875) because physical floaters and blur are geometrically continuous and appear identical across viewpoints despite poor visual quality. Independent refinement (Difix3D+) introduces disparate hallucinations that ruin consistency (CVSC drops to 0.819/0.850). SyncFix dramatically improves perceptual fidelity (DreamSim drops to 0.099) while maintaining top semantic consistency (CVSC 0.880).
- Substantial inference speedup: SyncFix resolves the learned ODE in a single deterministic step, refining 63 views in 5.63 seconds—over \(5.4\times\) faster than Difix3D+ (30.6s) and nearly \(20\times\) faster than FlowR (109.8s).
Highlights & Insights¶
- From Gaussian denoising to continuous optimal transport: Reformulating 3D refinement as Brownian bridge interpolation allows direct mapping from corrupted distributions to clean distributions, eliminating iterative sampling steps and ungrounded Gaussian assumptions.
- Permutation invariance scales across view counts: Decoupling viewpoint order through concatenated global attention allows cheap view-pair (\(N=2\)) training while generalizing to arbitrary view counts at inference, where more viewpoints naturally provide tighter consensus.
- Emergent pose-free geometric consistency: Without receiving explicit camera matrices, epipolar lines, or depth loss, cross-view attention autonomously discovers correspondences in latent space, proving that pretrained diffusion representations contain sufficient implicit 3D grounding.
Limitations & Future Work¶
- Reliance on visual overlap: SyncFix assumes sufficient overlap between viewpoints to form reliable attention links. Under extremely wide baselines with nearly disjoint camera frustums, the synchronization signal diminishes, causing the model to revert toward independent hallucinations.
- Plausible but unfaithful hallucination under catastrophic loss: When underlying 3DGS models suffer from massive structural absence, the generative prior synthesizes visually coherent geometries that may deviate from ground-truth reality.
- Future directions: Integrating coarse relative camera pose conditioning or end-to-end differentiable reprojection loops into latent bridge matching to strengthen geometric rigidity under extreme sparse-view regimes.
Related Work & Insights¶
- vs Difix3D+: Difix3D+ uses single-step diffusion filtering independently per view and relies on 3DGS optimization to harmonize discrepancies, resulting in geometry drift and instability; SyncFix couples views directly inside the latent bridge matching process via cross-view attention.
- vs FlowR: FlowR applies multi-view flow matching with latent-only velocity loss, resulting in blurred fine textures and slow convergence; SyncFix incorporates comprehensive image-space perceptual and Gram-matrix losses with single-step inference for crisper details and faster runtime.
- vs 3DGS-Enhancer / GSFixer: Video diffusion methods enforce temporal attention along smooth camera trajectories, suffering from high computational cost and poor scaling to unordered wide-baseline sets; SyncFix uses permutation-invariant attention, natively handling arbitrary view configurations.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneering integration of latent bridge matching and permutation-invariant multi-view attention for pose-free 3D refinement]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across DL3DV, NeRFBusters, and AnySplat, supported by the newly designed CVSC metric and thorough ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous formulation, and tight alignment between architectural diagrams and narrative]
- Value: ⭐⭐⭐⭐⭐ [Resolves the critical bottleneck of conflicting multi-view hallucinations during 2D diffusion-to-3D distillation with real-time inference speeds]