∂DIBR: Differentiable Depth Image-based Rendering for Fast Novel View Synthesis¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/LISA-VR/ddibr
Area: 3D Vision
Keywords: Depth Image-Based Rendering (DIBR), Differentiable Rendering, Novel View Synthesis, Triangle Primitives, Real-Time Rendering
TL;DR¶
Reconstructs classical Depth Image-Based Rendering (DIBR) into an end-to-end differentiable framework using connected triangle meshes to preserve geometric consistency and topology, reaching 277 FPS on a consumer GPU.
Background & Motivation¶
Modern Novel View Synthesis (NVS) is driven by an inverse rendering paradigm where explicit or implicit primitives are supervised by photometric reconstruction losses, prominently featured in 3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF). However, these modern primitives—such as independent 3D Gaussian ellipsoids or volumetric densities—evolve without topological constraints. During gradient-based optimization, they often overfit view-dependent appearances and produce severe floating artifacts, disconnected geometries, and irregular surfaces. More critically, standard computer graphics (CG) hardware and APIs (e.g., Vulkan, OpenGL, and standard rasterization pipelines) are fundamentally engineered around triangle primitives and cannot natively handle Gaussian rasterization or ray marching efficiently, hindering practical deployment in VR/AR and gaming engines.
Classical Depth Image-Based Rendering (DIBR) is inherently compatible with traditional CG pipelines and maintains explicit mesh connectivity by triangulating input RGBD images and reprojecting them into target viewpoints. Nevertheless, classical DIBR faces two fundamental hurdles. First, traditional pipelines depend on rigid geometric heuristics and depth estimation where discrete visibility decisions, z-buffer tests, and disocclusion handling are non-differentiable, preventing end-to-end photometric optimization. Second, large parallax shifts induce stretched triangles across depth discontinuities, resulting in severe hole artifacts. Recent attempts to introduce triangle primitives into modern NVS (such as Triangle Splatting) resort to disconnected "triangle soups," thereby losing topological connectivity and global geometric coherence.
This paper addresses these limitations by revisiting connected triangle meshes within an end-to-end differentiable pipeline, systematically overcoming the non-differentiabilities of classical DIBR. Core idea: build a differentiable DIBR framework on shader-level automatic differentiation, softening disocclusion culling via logistic probability gating, smoothing rasterization visibility with multi-sampling, and guiding joint geometry-color optimization via coarse-to-fine scheduling and learning-rate decoupling to achieve strict geometric connectivity and 277 FPS real-time rendering.
Method¶
Overall Architecture¶
∂DIBR takes a sparse set of multi-view RGB images along with initialized depth maps and directly executes forward differentiable rendering and backward gradient propagation within a Slang/Vulkan-based graphics pipeline. The workflow consists of five consecutive stages: two-level viewpoint selection, mesh warping and reprojection, probabilistic quality estimation, multi-sampled smooth rasterization, and quality-weighted multi-view blending.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Multi-view RGB Images & Initial Depths"] --> B["Two-level View Selection & Triangle Mesh Warping"]
B --> C["Soft Quality Gating"]
C --> D["Multi-sampled Smooth Rasterization"]
D --> E["Quality-weighted Multi-view Blending"]
E --> F["Hierarchical Decoupled Inverse Rendering"]
F -->|Backpropagate Photometric & TV Regularization Gradients| A
For each novel frame, the system performs two-level view selection: PAM clustering first identifies a representative geometric subset \(\mathcal{S}\), followed by real-time ranking based on baseline distance, viewing angle, and projected camera centers to select \(|\mathcal{I}|\) active views. Each input view is triangulated into a flat 2D mesh, projected into 3D using its depth map, and reprojected onto the target image plane. Soft quality gating identifies and smoothly suppresses distorted triangles caused by parallax disocclusion. Multi-sampling smooths visibility transitions, and the valid warped layers are blended into the final target image, optimized end-to-end via photometric and anisotropic total variation losses.
Key Designs¶
1. Soft Quality Gating: Relaxing Hard Disocclusion Discarding into Differentiable Probabilities In classical DIBR, adjacent pixels across depth boundaries become widely separated when warped to novel viewpoints due to parallax, forming large, elongated triangles (disocclusion cracks). Conventional methods apply hard thresholding to discard these stretched triangles, introducing step discontinuities that completely break gradient flow. ∂DIBR reinterprets triangle quality through a probabilistic lens, relaxing the discrete rejection into a smooth logistic formulation. For each warped triangle \(T_{i \rightarrow t}\), geometric distortion metrics \(f_q(T_{i \rightarrow t})\) (e.g., edge stretching, surface dilation) act as decision boundaries: $\(Q_{i \rightarrow t} = \sigma\left(\sum_{q \in \mathcal{Q}} \omega_q f_q(T_{i \rightarrow t})^{\lambda_q}; k, x_0\right)\)$ where \(k\) denotes the logistic slope and \(x_0\) is the midpoint threshold. When the warped geometry preserves natural proportions, \(Q_{i \rightarrow t} \to 1\); when severe parallax stretching occurs, \(Q_{i \rightarrow t} \to 0\). This smooth formulation eliminates visual tearing while providing continuous, well-behaved gradients to pull or contract depth vertices back to accurate boundaries.
2. Multi-sampled Smooth Rasterization: Mitigating Visibility Discontinuities in Triangle Buffers Determining which triangle encompasses a target pixel and resolving Z-Buffer depth ordering are inherently discrete operations. Infinitesimal perturbations to vertex positions can cause a sampling point to abruptly switch between different triangles or depths, generating severe gradient variance and optimization instability. ∂DIBR mitigates this discrete jump by adopting sub-pixel multi-sampling with \(|N|=4\) sample points per pixel. Each sub-pixel independently performs triangle identification and depth testing, fetching color attributes via barycentric interpolation: $\(C_{i \rightarrow t}(\vec{p}_t) = \frac{1}{|N|} \sum_{n=1}^{|N|} C_i(\vec{p}_{i, n})\)$ During multi-view blending, the warped colors from all selected views \(\mathcal{I}\) are fused using transparency \(\alpha_{i \rightarrow t}\) and soft quality \(Q_{i \rightarrow t}\), incorporating a minimal epsilon offset \(\epsilon\) to smoothly blend in background color \(C_0\): $\(C_t = \frac{\sum_{i \in \mathcal{I}} \alpha_{i \rightarrow t} Q_{i \rightarrow t} C_{i \rightarrow t} + \epsilon C_0}{\sum_{i \in \mathcal{I}} \alpha_{i \rightarrow t} Q_{i \rightarrow t} + \epsilon}\)$ Distributing visibility tests over multiple sub-pixel locations replaces single-point step transitions with piecewise continuous functions, substantially stabilizing backward gradient propagation.
3. Hierarchical Decoupled Inverse Rendering: Coarse-to-Fine Scheduling and Learning-Rate Decoupling Jointly optimizing depth geometry and RGB colors from scratch is ill-posed and susceptible to local minima. On repetitive textures, small initial depth inaccuracies cause periodic aliasing shifts where gradients easily stall; moreover, because color updates possess excessive capacity, the optimizer tends to modify pixel colors rather than correct depths, resulting in degenerate "color-overfitted false geometry." ∂DIBR tackles this issue with a two-fold strategy: a coarse-to-fine resolution pyramid that suppresses high-frequency noise in initial stages to settle macro geometry, and a decoupled learning rate schedule. In early coarse stages, the color learning rate is set three orders of magnitude lower than depth (\(10^{-5}\) vs \(10^{-2}\)), forcing photometric errors to drive geometric depth alignment. As coarse structures stabilize, the color learning rate is gradually increased until fine-stage joint refinement adjusts residual view-dependent effects and photometric balance.
Loss & Training¶
The overall training objective combines photometric reconstruction loss \(\mathcal{L}_c\) with an anisotropic total variation regularization term \(\mathcal{L}_{ATV}\): $\(\mathcal{L} = \mathcal{L}_c + \gamma \mathcal{L}_{ATV} = \sum_{t} |C_t - \tilde{C}_t| + \gamma \sum_{i} (|\Delta_x D_i| + |\Delta_y D_i|)\)$ where \(\mathcal{L}_c\) computes the \(L_1\) color difference between rendered and target views, and \(\mathcal{L}_{ATV}\) enforces smoothness across neighboring depth pixels to eliminate spikes. The regularization coefficient \(\gamma\) is set to 4 in early stages and 12 in the final fine stage.
The model is optimized using Adam (\(\beta_1=0.9, \beta_2=0.999\)) for 40,000 iterations. The three-stage learning rates (color, depth) are: - Stage 1 (coarse): color \(10^{-5}\), depth \(10^{-2}\) - Stage 2 (medium): color \(5 \times 10^{-5}\), depth \(10^{-2}\) - Stage 3 (fine): color \(5 \times 10^{-3}\), depth \(5 \times 10^{-3}\)
Key Experimental Results¶
Main Results¶
∂DIBR was evaluated on two challenging forward-facing benchmarks, LLFF and Shiny, against representative modern novel view synthesis methods.
| Dataset | Metric | 3DGS (SIGGRAPH 23) | 2DGS (SIGGRAPH 24) | TS (CVPR 24) | ∂DIBR (Ours) |
|---|---|---|---|---|---|
| LLFF (Average) | PSNR (dB) ↑ | 24.73 | 23.88 | 21.91 | 21.81 |
| SSIM ↑ | 0.816 | 0.799 | 0.739 | 0.722 | |
| LPIPS ↓ | 0.114 | 0.150 | 0.197 | 0.202 | |
| Shiny (Average) | PSNR (dB) ↑ | 23.25 | 22.56 | 23.15 | 22.44 |
| SSIM ↑ | 0.788 | 0.778 | 0.766 | 0.735 | |
| LPIPS ↓ | 0.128 | 0.140 | 0.110 | 0.171 |
In terms of computational efficiency, memory consumption, and storage, ∂DIBR demonstrates remarkable practical advantages:
| Method | PSNR (dB) | Training Time (min) | GPU Memory (GB) | Storage Size (MB) | Framerate (FPS) ↑ |
|---|---|---|---|---|---|
| ∂DIBR (Lvl 1 - Coarse) | 20.08 | 8 | 3.4 | 6 | 483 |
| ∂DIBR (Lvl 2 - Medium) | 20.55 | +9 | 6.4 | 21 | 386 |
| ∂DIBR (Lvl 3 - Fine) | 21.81 | +16 | 17.46 | 87 | 277 |
| 3DGS | 24.73 | 8 | 4.2 | 246 | 168 |
| 2DGS | 23.88 | 10 | 4.2 | 146 | 114 |
| TS (Triangle Splatting) | 21.91 | 22 | 12.4 | 353 | 110 |
Ablation Study¶
The ablation experiments examine the impact of depth initialization choices, anisotropic total variation weight \(\gamma\), and view sampling budgets.
| Module / Parameter | Configuration | PSNR (dB) ↑ | SSIM ↑ | LPIPS ↓ | Note |
|---|---|---|---|---|---|
| Depth Initialization | None (mean depth) | 21.52 | 0.706 | 0.218 | Slower convergence (60k steps), recovers valid geometry |
| DepthAnything | 21.54 | 0.694 | 0.218 | Comparable to unguided initialization | |
| ZoeDepth (Default) | 21.81 | 0.722 | 0.202 | Best overall convergence and visual fidelity | |
| Depth Regularization \(\gamma\) | \(\gamma = 0.01\) | 22.38 | 0.702 | 0.226 | Severe high-frequency noise and depth spikes |
| \(\gamma = 0.1\) | 22.69 | 0.739 | 0.219 | Surfaces progressively regularized | |
| \(\gamma = 1.0\) | 22.94 | 0.779 | 0.182 | Markedly improved perceptual and geometric quality | |
| \(\gamma = 4.0\) (Default) | 22.69 | 0.782 | 0.178 | Best LPIPS, prevents overfitting to photometric noise |
Testing on representation subset size \(|\mathcal{S}|\) and per-frame view count \(|\mathcal{I}|\) indicates that \(|\mathcal{I}|=6\) offers the optimal balance between visual quality and rendering speed (22.69 dB PSNR at 190 FPS), while \(|\mathcal{I}|=3\) yields an ultra-fast 321 FPS with roughly 1 dB loss.
Key Findings¶
- 2× Rendering Acceleration: Running on an RTX 4090, ∂DIBR reaches 277 FPS, delivering a \(1.65\times\) speedup over 3DGS (168 FPS) and \(2.43\times\) over 2DGS (114 FPS), maintaining native compatibility with standard rasterization hardware.
- Compact Footprint: Storing multi-view RGBD representations requires only 87 MB (or 6 MB for coarse scale), substantially smaller than 3DGS (246 MB) and Triangle Splatting (353 MB).
- Superior Global Geometric Robustness: While Gaussian splatting methods capture sharper high-frequency edges, their depth maps reveal large patches of false geometry compensated by view-dependent spherical harmonics. In contrast, ∂DIBR enforces continuous mesh topology, producing globally consistent and physically plausible surface reconstructions.
Highlights & Insights¶
- Probabilistic Relaxation of Discrete Graphics Culling: Formulating triangle disocclusion culling as a continuous logistic gating function bridges non-differentiable graphics heuristics with gradient-based inverse rendering.
- Frequency Decoupling Against Geometric Degeneracy: Suppressing color learning rates by three orders of magnitude during early training effectively prevents appearance overfitting from compensating for geometric depth errors.
- Direct Integration with Standard Graphics Pipelines: Implemented via Slang/Vulkan shaders using native triangle primitives, ∂DIBR bypasses custom CUDA kernels and neural field decoders, presenting an ideal paradigm for low-latency VR/AR deployment.
Limitations & Future Work¶
- Visual Sharpness Deficit Compared to Splatting: Suffers up to a 2–3 dB PSNR drop on fine, semi-transparent, or thin structures (e.g., foliage, thin wires) where single-layer mesh topologies struggle to represent complex depth discontinuities.
- Constrained View-Dependent Appearance Modeling: Relies primarily on multi-view blending without explicit spherical harmonics or physically based reflectance modeling, limiting fidelity on non-Lambertian and mirror-like surfaces.
- Future Directions: Exploring Layered Depth Images (LDI) or neural material texture maps within the differentiable DIBR pipeline to capture complex occlusions and specular highlights while preserving high rendering framerates.
Related Work & Insights¶
- vs 3DGS / 2DGS: 3DGS models scenes using millions of independent Gaussian ellipsoids, achieving exquisite visual quality but lacking topological continuity and suffering from floating artifacts. 2DGS introduces planar constraints but maintains disconnected primitives. ∂DIBR adopts continuous triangle meshes, trading minor high-frequency sharpness for a \(2\times\) rendering speedup and robust global geometry.
- vs Triangle Splatting (TS): TS uses triangle primitives but treats them as an unstructured, floating triangle soup. ∂DIBR retains connected image-space grid topologies, preventing disjointed surface tears.
- vs Classical DIBR (e.g., Ravis, OpenDIBR): Classical DIBR serves strictly as forward image warping based on static depth priors; ∂DIBR turns the entire pipeline into a differentiable engine where depth maps are iteratively refined through multi-view photometric gradients.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Successfully formulates classical DIBR into an end-to-end differentiable shader architecture, offering elegant relaxations for disocclusion and visibility.
- Experimental Thoroughness: ⭐⭐⭐⭐☆ Comprehensive evaluations on LLFF and Shiny datasets covering rendering framerate, GPU memory, depth regularization, and view budgeting.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous narrative with clear mathematical derivations and lucid illustrations of non-differentiable approximations.
- Value: ⭐⭐⭐⭐⭐ Bridges modern inverse rendering optimization with legacy triangle graphics hardware, providing a highly practical path for real-time VR/AR applications.