FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors¶
Conference: ECCV 2026
Paper: ECCV Official
Project Page: https://fix-anything.github.io
Area: 3D Vision
Keywords: novel view synthesis, video diffusion models, 3D reconstruction, generative prior refinement
TL;DR¶
FixAnything introduces a generalist rendering refinement framework that reformulates sparse-view artifacts across 3DGS, NeRF, meshes, and sparse point clouds as latent-space video-to-video translation, combining mask-aware conditioning with SfM-pose-driven Flow-DPO preference alignment to achieve photorealistic and 3D-consistent outputs via lightweight LoRA adaptation.
Background & Motivation¶
Recent years have witnessed remarkable advances in 3D scene reconstruction and novel view synthesis, yet a fundamental bottleneck persists: whenever input training views are sparse or novel camera positions deviate significantly from captured trajectories, every 3D representation exhibits characteristic artifacts. 3D Gaussian Splatting (3DGS) suffers from semi-transparent floaters scattered across empty space; Neural Radiance Fields (NeRF) hallucinate foggy, low-frequency clouds; reconstructed meshes exhibit severe geometric distortions and missing textures at occlusion boundaries; and raw sparse point clouds leave extensive visual voids. The prevailing response to these degradations has been developing specialist generative pipelines: designing custom 3D diffusion regularizers tailored to NeRF, training bespoke spatial-temporal video decoders for 3DGS, or engineering complex camera controllers conditioned on explicit depth. Such specialist strategies are cumbersome and hard to scale, necessitating complete architectural redesigns and massive paired retraining datasets whenever a new 3D representation arises.
Upon deeper inspection, although artifacts produced by different 3D representations appear visually disparate, they all share a critical underlying property: they deviate from the manifold of natural videos while faithfully retaining the camera trajectory and coarse scene layout. The core challenge in leveraging generative models for cleanup lies in the delicate trade-off between realism and consistency. Processing novel views frame-by-frame with 2D generative models introduces glaring temporal flickering and identity drift across viewpoints. Conversely, applying off-the-shelf video diffusion backbones without explicit multi-view geometric grounding often generates hallucinated scene components that morph and shift across frames, severely breaking downstream structure-from-motion (SfM) camera tracking.
This paper's angle of attack is to repurpose a single pretrained video foundation model (Wan2.1) to exploit its rich implicit multi-view and natural video priors directly. Core idea: formulate multi-representation rendering refinement as latent video-to-video translation conditioned on concatenated degraded sequences and trust masks, while distilling multi-view geometric consistency directly into the model weights using structure-from-motion camera pose accuracy as a preference reward in Flow-DPO.
Method¶
Overall Architecture¶
The overall FixAnything pipeline encompasses representation-agnostic latent conditioning followed by geometry-aware preference optimization. Given any 3D representation (3DGS, NeRF, triangular mesh, or COLMAP sparse point cloud), the system renders a degraded video sequence \(x \in \mathbb{R}^{T \times 3 \times H \times W}\) along the desired camera trajectory and constructs a 1D binary mask \(m \in \{0, 1\}^T\) indicating which frames correspond to pristine, captured training viewpoints. The video frames and broadcast mask are encoded through a frozen Wan2.1 VAE and concatenated with the Rectified Flow noisy latents along the channel dimension before entering the Wan2.1 Diffusion Transformer (DiT), which is adapted solely via a lightweight LoRA module. Following supervised finetuning (SFT), the model undergoes Flow-DPO preference alignment using COLMAP camera pose estimation accuracy as the reward signal, baking rigorous multi-view consistency into the network weights without incurring any test-time computational overhead.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Degraded 3D rendering video + Binary training view mask"] --> B["Latent-Space Channel Concatenation<br/>VAE encodes degraded latents, mask, and noisy states"]
B --> C["Mask-Aware Conditioning<br/>Anchors known training views and propagates clean context"]
C --> D["Pose-Accuracy Preference Optimization<br/>SfM camera pose AUC reward guides Flow-DPO alignment"]
D --> E["Output: Photorealistic, 3D-consistent refined video"]
Key Designs¶
1. Latent-Space Channel Concatenation: Unifying Heterogeneous Artifacts into Video Inpainting
Prior generative refinement approaches construct dedicated sub-networks or custom cross-attention layers for specific 3D data formats, adding architectural complexity and preventing transfer across representations. FixAnything recognizes that degraded renderings inherently provide a reliable camera trajectory and coarse spatial layout scaffold. Consequently, no structural modifications to the base video DiT are needed. The degraded rendering video \(x\) is mapped to a conditional latent \(z_{\text{cond}} = \mathcal{E}(x)\) using the frozen pretrained VAE encoder, while the clean target video \(y\) is mapped to \(z_0 = \mathcal{E}(y)\). Under the rectified flow formulation at timestep \(t \in [0, 1]\), the interpolated noisy latent is: $\(z_t = (1 - t) z_0 + t \epsilon, \quad \epsilon \sim \mathcal{N}(0, \mathbf{I})\)$ The degraded latent \(z_{\text{cond}}\), the noisy latent \(z_t\), and the spatially broadcast binary mask \(m\) are concatenated along the channel dimension to form the augmented latent \(\hat{z}_t = [z_t ; z_{\text{cond}} ; m]\). A lightweight LoRA adapter (rank 64, updating less than 1% of total parameters) is trained on top of the frozen DiT to predict the velocity field \(v = \epsilon - z_0\). This design completely bypasses the need for explicit SE(3) camera coordinate conditioning (such as Plücker ray embeddings) and enables effective cleanup learning using as few as 20 paired training videos.
2. Mask-Aware Conditioning: Distinguishing Frames to Trust from Frames to Fix
In sparse-view rendering trajectories, visual degradation is highly non-uniform across time: viewpoints close to captured training poses render cleanly, whereas distant views suffer severe breakdown. If a generative model is prompted to clean the entire video without discrimination, it struggles to identify pristine content and tends to hallucinate over valid regions, degrading fidelity. FixAnything incorporates a 1D per-frame binary mask \(m\) to make this trust-versus-fix distinction explicit: \(m_i = 1\) for frames captured at training poses (to be trusted and preserved) and \(m_i = 0\) for all intermediate novel views (to be cleaned). This mask serves two crucial roles: first, it prevents unwanted hallucination over known training observations, preserving fine details of known objects; second, the clean frames act as high-fidelity spatiotemporal anchors, allowing the model's self-attention mechanism to propagate real appearance, lighting, and geometric structures to adjacent corrupted frames rather than guessing textures from scratch. Removing this mask degrades rendering PSNR by 1.28 dB in controlled ablations.
3. Pose-Accuracy Preference Optimization: Injecting 3D Consistency at Zero Inference Cost
While the flow matching loss in supervised finetuning (SFT) yields sharp per-frame textures, the SFT model frequently produces subtle hallucinations—such as background structures that morph or drift across frames—which appear visually plausible in individual frames but violate multi-view geometry. Running classical structure-from-motion (SfM) pipelines on these raw outputs leads to tracking failure and erratic camera trajectories. FixAnything reframes multi-view consistency as a preference optimization problem. For 1,000 DL3DV training scenes, the model generates five candidate outputs per scene using different random seeds. COLMAP is executed on each candidate using SuperPoint keypoints and LightGlue matching to estimate recovered camera poses against ground-truth poses. Geometric consistency is quantified using Area Under the Curve at a 5-degree threshold (AUC@5°) across relative rotation and relative translation accuracy. Candidate pairs \((y_w, y_l)\) with an AUC gap \(\ge 0.2\) are formed into preference pairs, optimized via Flow-DPO: $\(\mathcal{L}_{\text{DPO}} = -\mathbb{E} \left[ \log \sigma \left( \frac{\beta}{2} (\Delta_l - \Delta_w) \right) \right]\)$ where \(\Delta_w = \|v_w - v_\theta(\hat{z}_t^w, t)\|^2 - \|v_w - v_{\text{ref}}(\hat{z}_t^w, t)\|^2\), \(\Delta_l\) is the dispreferred counterpart, and \(v_{\text{ref}}\) is the frozen SFT checkpoint. This objective directly aligns the velocity field to suppress cross-view geometric hallucinations, improving camera pose recovery AUC@5° by 7.2% while adding zero additional parameters or inference latency.
Loss & Training¶
The training pipeline consists of two sequential stages: 1. Stage I (Supervised Finetuning - SFT): Trained on 500 paired DL3DV-10K video trajectories rendered across four 3D representations. The model is first initialized at \(288 \times 512\) resolution and subsequently fine-tuned at \(480 \times 832\) resolution (\(T=61\) frames) for 3,000 steps on a single H100 GPU using standard flow matching loss \(\mathcal{L}_{\text{FM}} = \mathbb{E}_{\hat{z}_t, t} [ \|v - v_\theta(\hat{z}_t, t)\|^2 ]\). 2. Stage II (Flow-DPO Preference Optimization): Using the SFT checkpoint as reference \(v_{\text{ref}}\), the LoRA parameters are further fine-tuned for 2,000 iterations on the pose-accuracy preference dataset. At inference time, the model executes 50 ODE denoising steps by default. Trajectories longer than 61 frames are processed via overlapping temporal windows.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on 20 held-out scenes from the DL3DV-10K dataset across 3-view, 6-view, and 9-view sparse settings. FixAnything evaluates a single unified model across four distinct input representations (NeRF, 3DGS, triangular mesh, and COLMAP sparse point clouds) and is compared against sparse-view 3D baselines and specialist post-hoc enhancement methods:
| Category | Method / Input Representation | 3 Views PSNR↑ | 3 Views SSIM↑ | 3 Views LPIPS↓ | 6 Views PSNR↑ | 6 Views SSIM↑ | 6 Views LPIPS↓ | 9 Views PSNR↑ | 9 Views SSIM↑ | 9 Views LPIPS↓ |
|---|---|---|---|---|---|---|---|---|---|---|
| Sparse-view Reconstruction | 3DGS | 10.97 | 0.248 | 0.567 | 13.34 | 0.332 | 0.498 | 14.99 | 0.403 | 0.446 |
| Sparse-view Reconstruction | RegNeRF | 11.46 | 0.214 | 0.600 | 12.69 | 0.236 | 0.579 | 12.33 | 0.219 | 0.598 |
| Sparse-view Reconstruction | FreeNeRF | 10.91 | 0.211 | 0.595 | 12.13 | 0.230 | 0.576 | 12.85 | 0.241 | 0.573 |
| Sparse-view Reconstruction | DNGaussian | 11.10 | 0.273 | 0.579 | 12.67 | 0.329 | 0.547 | 13.44 | 0.365 | 0.539 |
| Sparse-view Reconstruction | FSGS | 12.22 | 0.296 | 0.535 | 13.73 | 0.429 | 0.540 | 15.52 | 0.468 | 0.416 |
| Post-hoc Enhancement | 3DGS-Enhancer | 14.33 | 0.424 | 0.464 | 16.94 | 0.565 | 0.356 | 18.50 | 0.630 | 0.305 |
| Post-hoc Enhancement | Xu et al. | 14.62 | 0.471 | 0.491 | 17.35 | 0.566 | 0.396 | 19.19 | 0.616 | 0.335 |
| Post-hoc Enhancement | Difix3D | 12.85 | 0.392 | 0.557 | 14.84 | 0.445 | 0.462 | 16.76 | 0.520 | 0.399 |
| Post-hoc Enhancement | Difix3D+ | 12.37 | 0.363 | 0.512 | 14.41 | 0.424 | 0.400 | 16.39 | 0.498 | 0.330 |
| FixAnything (Ours) | w/ NeRF rendering | 14.22 | 0.427 | 0.451 | 17.01 | 0.522 | 0.329 | 18.86 | 0.605 | 0.297 |
| FixAnything (Ours) | w/ 3DGS rendering | 15.18 | 0.452 | 0.408 | 17.65 | 0.561 | 0.289 | 19.76 | 0.632 | 0.269 |
| FixAnything (Ours) | w/ mesh rendering | 15.74 | 0.482 | 0.366 | 17.95 | 0.583 | 0.269 | 19.86 | 0.646 | 0.233 |
| FixAnything (Ours) | w/ sparse SfM points | 15.52 | 0.463 | 0.381 | 17.74 | 0.568 | 0.271 | 19.72 | 0.624 | 0.241 |
Ablation Study¶
1. Effect of Flow-DPO and Mask-Aware Conditioning
| Config | Evaluation Target | PSNR↑ | SSIM↑ | LPIPS↓ | AUC@5° (%)↑ | Note |
|---|---|---|---|---|---|---|
| SFT only | Supervised baseline | 17.51 | 0.554 | 0.296 | 61.12 | Sharp per-frame renderings but suffers from subtle cross-view geometric drift |
| SFT + Flow-DPO | Full model | 17.65 | 0.561 | 0.289 | 68.32 | Suppresses floating hallucinations; pose recovery AUC improves by +7.2% with zero extra test-time cost |
| No mask (\(m=0\)) | Unconditioned mask | 16.37 | 0.525 | 0.311 | - | Fails to distinguish clean frames; hallucinates over training views, dropping 1.28 dB PSNR |
| With mask | Mask-aware model | 17.65 | 0.561 | 0.289 | 68.32 | Preserves trusted views and anchors context propagation across novel frames |
2. Training Data Volume and Denoising Steps vs. Runtime
| Training Videos | PSNR↑ | SSIM↑ | LPIPS↓ | Denoising Steps | PSNR↑ | SSIM↑ | LPIPS↓ | Time per 61-frame clip (s) |
|---|---|---|---|---|---|---|---|---|
| 20 | 16.70 | 0.531 | 0.309 | 5 steps | 18.02 | 0.574 | 0.313 | 31 (10× speedup) |
| 50 | 17.20 | 0.548 | 0.297 | 10 steps | 17.91 | 0.570 | 0.296 | 62 |
| 100 | 17.45 | 0.556 | 0.292 | 25 steps | 17.75 | 0.564 | 0.289 | 155 |
| 500 (default) | 17.65 | 0.561 | 0.289 | 50 steps (default) | 17.65 | 0.561 | 0.289 | 309 |
Key Findings¶
- Sparse Point Clouds Perform on Par with Dense 3DGS: Renderings derived from sparse COLMAP keypoints with clean anchor frames achieve a 6-view PSNR of 17.74 dB, closely matching dense 3DGS inputs (17.65 dB) and mesh inputs (17.95 dB). This reveals that intermediate, computationally expensive 3D fitting stages (e.g. dense NeRF or 3DGS training) are largely redundant for novel view generation when coupled with a potent video foundation model; the coarse rendering primarily serves as a camera trajectory scaffold.
- Remarkable Data Efficiency: While prior enhancement techniques require 80,000 to 150,000 paired frames, FixAnything achieves competitive restoration with as few as 20 paired video clips, with performance plateauing beyond 100 clips. The base video model already encapsulates natural physical priors, meaning the LoRA adapter only needs to learn the conditioning mapping.
- Fast Low-Step Inference: Thanks to the rectified flow formulation, reducing denoising steps from 50 to 5 preserves visual fidelity (18.02 dB PSNR) while accelerating inference tenfold to 31 seconds per 61-frame clip on an H100 GPU.
Highlights & Insights¶
- From Specialist Pipelines to a Single Foundation Prior: Instead of engineering bespoke architectures for each emerging 3D format, FixAnything demonstrates that diverse representation artifacts can be unified as out-of-distribution deviations on the natural video manifold, allowing a single lightweight LoRA adapter to resolve all of them.
- SfM Tracking Accuracy as an RL Alignment Reward: Using COLMAP pose estimation AUC as the preference signal in Flow-DPO bypasses intractable differentiable rendering gradients while embedding strict multi-view epipolar constraints directly into the network weights.
- Training-Free Epistemic Uncertainty Estimation: Measuring the pixel-wise standard deviation across multiple random generation seeds effectively highlights occluded, unobserved, or hallucinated scene geometry (high-confidence pixels achieve 25.7 dB PSNR vs. 14.4 dB for low-confidence ones), providing a natural confidence map for robotic path planning or active view capture.
Limitations & Future Work¶
- Author-Acknowledged Limitations: The temporal context window is bounded by Wan2.1's native 61-frame capacity. Synthesizing extended flight paths requires chunked sliding-window processing, which can occasionally exhibit subtle inter-chunk lighting or texture inconsistencies.
- Additional Identified Limitations: In scenarios with extreme camera extrapolations (e.g. navigating entirely behind an unobserved structure without line-of-sight overlap with any training view), the model must rely entirely on generative hallucination, which may deviate from actual physical reality.
- Future Directions: The generated, 3D-consistent novel views can be cycled back into 3DGS or NeRF optimization pipelines as pseudo-ground-truth observations, establishing a self-supervised loop from sparse views to dense, artifact-free 3D representations.
Related Work & Insights¶
- vs. 3DGS-Enhancer / Difix3D+: Existing enhancement pipelines process views independently or rely on complex custom spatial-temporal decoders followed by distillation into 3DGS, often resulting in cross-view flickering; FixAnything generates globally coherent video sequences in a single pass.
- vs. ViewCrafter / GEN3C: These approaches require training explicit camera pose injection layers (e.g. via Plücker ray coordinates) requiring large-scale retraining; FixAnything utilizes the rendered video itself as an implicit camera trajectory guide, drastically simplifying training.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates multi-representation 3D cleanup as video manifold translation and uses SfM pose accuracy as a Flow-DPO reward.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 4 representations, multiple view densities, LoRA scale, data volume, and step-reduction ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Lucid problem formulation, clear architectural justification, and rigorous empirical validation.
- Value: ⭐⭐⭐⭐⭐ Sets a compelling precedent for replacing complex specialist 3D pipelines with generalist video foundation models.