LumiDepth: Stable Monocular Depth in Multi-Illumination Scenes¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://drive.google.com/drive/folders/1-3txE3KiFiscACxs1Xch8UO1Ax82VBNg
Area: 3D Vision
Keywords: Monocular Depth Estimation, Multi-Illumination Scenes, Diffusion Foundation Models, Label-Free Training, Frequency-Aware Consistency
TL;DR¶
Addressing geometric distortion and temporal inconsistency under real-world multi-light and shifting illumination conditions, LumiDepth presents a target-domain label-free framework that learns illumination-invariant monocular depth from multi-illumination RGB image groups via Disagreement-Calibrated Probabilistic Pseudo Supervision (DCPS) and Frequency-aware Consistency and Distillation (FaCD).
Background & Motivation¶
Monocular depth estimation (MDE) has seen substantial progress driven by depth foundation models (such as MiDaS, Depth Anything, Marigold, and Lotus) trained on massive RGB-D datasets. Nonetheless, nearly all these foundation models are optimized primarily on uniformly lit environments. In practical robotic manipulation workspaces with mixed lighting or nighttime autonomous driving with shifting headlights and street lamps, scenes frequently contain multiple spatially distributed light sources. These conditions create strong cast shadows, specular highlights, and drastic exposure shifts that distort surface visual cues. Because existing estimators tend to confuse illumination-induced contrast with underlying 3D geometry, they produce warped surfaces, phantom contours, or complete structural failures when lighting changes.
Directly overcoming this bottleneck with existing paradigms encounters two steep obstacles. First, real-world multi-illumination scenes with accurate ground truth depth are exceptionally difficult to capture, leading to a critical shortage of target-domain RGB-D pairs. Second, synthetically augmenting datasets using recent multi-light relighting diffusion models primarily optimizes for perceptual realism rather than geometric consistency, frequently introducing high-frequency visual artifacts that corrupt depth boundaries. Furthermore, temporal consistency methods used in video depth estimation rely on optical flow and motion cues under static lighting, which break down entirely in static scenes where lighting shifts abruptly without camera motion.
This paper tackles the challenge from a distinct angle: the underlying 3D geometry of a static scene remains invariant regardless of lighting perturbations. Hence, invariant geometric priors can be extracted directly from unlabeled groups of multi-illumination RGB images. Core idea: Disagreement-Calibrated Probabilistic Pseudo Supervision (DCPS) filters out unreliable outlier hypotheses while preserving supervision diversity, paired with Frequency-aware Consistency and Distillation (FaCD) that aligns global shapes in the low-frequency band and distills high-frequency structural edges bi-directionally without over-smoothing.
Method¶
Overall Architecture¶
LumiDepth is built upon latent diffusion depth models (specifically Marigold and Lotus backbones). During training, the input consists of standard uniformly lit synthetic RGB-D pairs for base geometry supervision, alongside unlabeled multi-illumination RGB image groups \(\mathcal{G} = \{\mathbf{x}_1, \dots, \mathbf{x}_N\}\) depicting identical static scenes under diverse lighting. The pipeline first uses the frozen-feature depth backbone to infer depth hypotheses across all illumination variants. The DCPS module measures cross-illumination disagreement to filter out unstable outliers and probabilistically samples pseudo labels. Then, during the conditional denoising training of the diffusion U-Net, FaCD decomposes predicted depth into low-frequency and high-frequency components: low-frequency consistency aligns global contours, while high-frequency stop-gradient distillation transfers crisp edges from confident views to corrupted views.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multi-Illumination RGB Images<br/>G = {x1, ..., xN}"] --> B["Disagreement-Calibrated Probabilistic Pseudo Supervision (DCPS)<br/>Cross-illumination disagreement + adaptive filtering + sampling"]
B --> C["Base Diffusion Denoiser<br/>Conditional U-Net training"]
C --> D["Frequency-aware Consistency and Distillation (FaCD)<br/>Low-frequency shape alignment + high-frequency edge distillation"]
D --> E["Illumination-Invariant Sharp Depth Map"]
Key Designs¶
1. Disagreement-Calibrated Probabilistic Pseudo Supervision: Adaptive filtering and diversity-preserving pseudo labeling
In a multi-illumination group \(\mathcal{G}\), different lighting setups degrade different surface regions of the model's predictions. Taking a naive element-wise mean or median across all candidate predictions pulls in degraded surface artifacts and produces blurred boundaries, while selecting a single "best" hypothesis with minimal overall variance can overfit to lighting-specific pseudo-features. DCPS first calculates the disagreement-based uncertainty for each prediction \(\mathbf{d}_i\) against all other illumination variants in the group: \(v_i = \frac{1}{N-1} \sum_{j \neq i} \|\mathbf{d}_i - \mathbf{d}_j\|_2^2\). To adapt to different scene complexities, an adaptive threshold \(\tau = \mathrm{median}(\{v_i\}_{i=1}^N)\) removes candidates with uncertainty above the median, retaining \(K \ge N/2\) reliable hypotheses \(\{\mathbf{d}_1, \dots, \mathbf{d}_K\}\) with corresponding uncertainty scores \(\{u_1, \dots, u_K\}\). These uncertainties are converted into normalized confidence weights via temperature-scaled softmax: $\(\rho_k = \frac{\exp(-u_k / T)}{\sum_{j=1}^K \exp(-u_j / T)}\)$ A candidate is then stochastically drawn from the categorical distribution \(k^* \sim \mathrm{Categorical}(\rho_1, \dots, \rho_K)\) to serve as the training pseudo target \(\mathbf{d}^{\mathrm{pseudo}} = \mathbf{d}_{k^*}\). This acts as an implicit ensemble over training iterations that guards against single-candidate bias while avoiding blur from explicit averaging.
2. Frequency-aware Consistency and Distillation: Decoupled global shape alignment and edge detail transfer
Enforcing strict full-band consistency across illumination variants (e.g., pulling latent codes or depth maps directly together via \(\ell_2\) distance) acts as an over-regularizer: because appearance variations cannot match perfectly, the optimization takes the path of least resistance by suppressing fine boundaries, resulting in overly smooth depth. FaCD addresses this by decoding predicted latents into depth space \(\hat{\mathbf{d}}_1 = \mathcal{D}(\hat{\mathbf{z}}_1)\) and \(\hat{\mathbf{d}}_2 = \mathcal{D}(\hat{\mathbf{z}}_2)\), applying a Gaussian low-pass filter \(\mathcal{LPF}(\cdot)\) to constrain global shapes: $\(\mathcal{L}_{\mathrm{LF}} = \big\|\mathcal{LPF}(\hat{\mathbf{d}}_1) - \mathcal{LPF}(\hat{\mathbf{d}}_2)\big\|_1\)$ To preserve structural sharpness, FaCD computes depth gradients to obtain edge magnitudes \(g_i = \|\nabla \hat{\mathbf{d}}_i\|\) and thresholded edge masks \(\mathbf{M}_i = \mathbb{I}(g_i > \tau_d)\). Rather than enforcing bidirectional pull which risks propagating corrupted edges from degraded lighting conditions, FaCD leverages the DCPS confidence weights \(\rho\) in an asymmetric stop-gradient distillation: $\(\mathcal{L}_{\mathrm{HF}} = \rho_2 \big\|\nabla \hat{\mathbf{d}}_1 - \mathrm{sg}[\nabla \hat{\mathbf{d}}_2]\big\|_{1, \mathbf{M}_2} + \rho_1 \big\|\nabla \hat{\mathbf{d}}_2 - \mathrm{sg}[\nabla \hat{\mathbf{d}}_1]\big\|_{1, \mathbf{M}_1}\)$ The more confident prediction acts as a teacher providing stronger supervision, allowing reliable high-frequency geometric edges to be transferred into views suffering from severe shadows or specular washouts.
Loss & Training¶
The overall training loss balances base synthetic RGB-D supervision, unlabeled pseudo supervision, and frequency-aware consistency: $\(\mathcal{L}_{\mathrm{total}} = \mathcal{L}_{\mathrm{GT}} + \lambda_1 \mathcal{L}_{\mathrm{pseudo}} + \lambda_2 \mathcal{L}_{\mathrm{LF}} + \lambda_3 \mathcal{L}_{\mathrm{HF}}\)$ The hyperparameters are fixed to \(\lambda_1 = 0.5\), \(\lambda_2 = 0.5\), and \(\lambda_3 = 0.1\) across all benchmarks. During training, the VAE encoder, decoder, and the initial two downsampling stages of the denoiser U-Net are frozen. The low-pass filter uses a Gaussian kernel of size 11 with \(\sigma = 10\), and the gradient edge threshold is set to \(\tau_d = 0.05\).
Key Experimental Results¶
Main Results¶
To evaluate cross-illumination geometric stability, the authors formulate the mean depth variation \(\mathrm{MeanDV} = \frac{1}{N}\sum_i v_i\) and maximum per-image variation \(\mathrm{MaxPIV} = \max_i v_i\). Zero-shot comparisons across the real-world multi-illumination benchmark ReMID and two nighttime driving benchmarks (NuScenes-Night and RobotCar-Night) are summarized below:
| Method | ReMID MeanDV↓ | ReMID MaxPIV↓ | ReMID AbsRel↓ | ReMID \(\delta_1\)↑ | NuScenes-N AbsRel↓ | NuScenes-N \(\delta_1\)↑ | RobotCar-N AbsRel↓ | RobotCar-N \(\delta_1\)↑ |
|---|---|---|---|---|---|---|---|---|
| DAv2 [65] | 0.051 | 0.076 | 0.119 | 0.857 | 0.262 | 0.707 | 0.242 | 0.514 |
| DAv3 [36] | 0.049 | 0.075 | 0.118 | 0.862 | 0.270 | 0.695 | 0.235 | 0.520 |
| GenPercept [62] | 0.052 | 0.088 | 0.111 | 0.880 | 0.256 | 0.584 | 0.228 | 0.663 |
| DA-AC [51] | 0.047 | 0.080 | 0.117 | 0.876 | 0.241 | 0.719 | 0.227 | 0.557 |
| Marigold [26] | 0.101 | 0.188 | 0.181 | 0.726 | 0.308 | 0.506 | 0.267 | 0.605 |
| LumiDepth (n-step) | 0.057 | 0.097 | 0.138 | 0.825 | 0.274 | 0.585 | 0.254 | 0.632 |
| Lotus [20] | 0.043 | 0.074 | 0.113 | 0.850 | 0.280 | 0.580 | 0.245 | 0.656 |
| Lotus [20] (st) | 0.048 | 0.085 | 0.120 | 0.835 | 0.306 | 0.554 | 0.289 | 0.632 |
| LumiDepth (1-step) | 0.030 | 0.055 | 0.085 | 0.914 | 0.223 | 0.636 | 0.217 | 0.676 |
Ablation Study¶
Ablations on pseudo-label selection and frequency decomposition validate each core module:
| Config / Variant | ReMID AbsRel↓ | ReMID \(\delta_1\)↑ | NuScenes-N AbsRel↓ | Note |
|---|---|---|---|---|
| DCPS Selection Strategy | ||||
| Average Hypothesis (Mean) | 0.092 | 0.909 | 0.265 | Mixes in distorted predictions, blurring boundaries |
| Geometric Median (Median) | 0.091 | 0.911 | 0.262 | Resilient to outliers but lacks diversity |
| Minimum Variance (Best) | 0.088 | 0.912 | 0.245 | Overfits to single lighting-specific artifact |
| DCPS Probabilistic (Full) | 0.085 | 0.914 | 0.223 | Adaptive pruning + confidence-weighted sampling |
| FaCD Consistency Scheme (n-step) | ||||
| Direct Latent Alignment (Latent) | 0.365 | 0.629 | 0.738 | Over-regularization causes geometric collapse |
| Direct Depth Alignment (Depth) | 0.274 | 0.692 | 0.520 | Frequency-agnostic constraint smooths fine edges |
| Low-Frequency Only (\(\mathcal{L}_{\mathrm{LF}}\)) | 0.145 | 0.818 | 0.282 | Preserves shape but lacks fine boundary guidance |
| High-Frequency Only (\(\mathcal{L}_{\mathrm{HF}}\)) | 0.170 | 0.766 | 0.289 | Lacks global scale anchoring, increasing absolute error |
| FaCD Full Module | 0.138 | 0.825 | 0.274 | Balances global consistency and sharp edge transfer |
Key Findings¶
- Naive self-training on unlabeled RGB multi-illumination data without geometric constraints (Lotus st) degrades stability and accuracy (ReMID MeanDV rises from 0.043 to 0.048; AbsRel rises from 0.113 to 0.120), highlighting that unfiltered pseudo supervision injects noise, whereas DCPS+FaCD produces clear improvements.
- Frequency decoupling is essential. Directly forcing latent vectors together degrades performance catastrophically (AbsRel climbs to 0.365), and undivided depth constraints cause blurry surfaces (AbsRel 0.274). Only isolating low-frequency alignment with directional gradient distillation achieves optimal results (AbsRel 0.138).
- The framework is model-agnostic: plugging DCPS and FaCD into the discriminative DAv2 backbone reduces its MeanDV from 0.051 to 0.036 on ReMID, proving broad architectural versatility beyond diffusion pipelines.
Highlights & Insights¶
- Cross-illumination disagreement as an uncertainty metric: By computing the variance across illumination variants of the same static scene, the method derives an effective self-supervised confidence score without ground truth, using median gating to maintain dynamic validity.
- Asymmetric stop-gradient distillation: Addressing error propagation between degraded and clear views, directional distillation weighted by confidence prevents noisy edges from contaminating reliable predictions.
- ReMID benchmark and stability metrics: Addresses the absence of real multi-illumination evaluation benchmarks in MDE, providing paired multi-light RGB-D captures alongside MeanDV and MaxPIV metrics for rigorous robustness benchmarking.
Limitations & Future Work¶
- Reliance on illumination diversity within groups: DCPS confidence estimation assumes that the input image groups contain sufficiently diverse, spatially distributed lighting variants; highly homogeneous or uniformly pitch-black sets fail to provide distinct disagreement signals.
- Training computational overhead: Processing multiple illumination views simultaneously during training to evaluate DCPS uncertainty and FaCD losses demands higher GPU memory and multi-view batching coordination.
- Future Directions: Extending the frequency-aware consistency paradigm to dynamic video environments with moving objects, and integrating physics-based illumination priors for extreme backlighting conditions.
Related Work & Insights¶
- vs Depth Anything Series (DA / DAv2 / DAv3): Depth Anything models depend on massive internet-scale data and discriminative teacher distillation, but still exhibit localized warping under intense non-uniform lighting; LumiDepth achieves superior illumination resilience without requiring target-domain depth labels.
- vs Marigold / Lotus: Diffusion-based MDE backbones offer fine edge synthesis but are susceptible to misinterpreting shadow boundaries as geometric depth jumps; LumiDepth remedies this vulnerability through frequency-decoupled distillation.
- vs Video Depth Consistency: Video approaches rely on camera ego-motion and optical flow temporal continuity, whereas LumiDepth explicitly targets zero-motion scenes experiencing severe lighting transitions.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Novel label-free formulation targeting multi-illumination depth; elegant frequency decoupling and probabilistic pseudo supervision]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Introduces real ReMID benchmark, testing across extreme indoor lighting, nighttime driving, and 18 RoboDepth corruption types]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically well-posed mechanisms, structured narrative and comprehensive ablation studies]
- Value: ⭐⭐⭐⭐⭐ [Provides an effective solution for robotics and autonomous driving safety in difficult illumination conditions without expensive target depth labeling]