Weather-Conditioned Depth Anything¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: 3D Vision
Keywords: monocular depth estimation, robust depth estimation, style-content disentanglement, style filter, parameter-efficient fine-tuning
TL;DR¶
To tackle catastrophic failures and feature entanglement in monocular depth foundation models under adverse weather conditions (rain, snow, fog, night), this paper presents DA-W (Weather-Conditioned Depth Anything), which disentangles weather style via a cross-domain contrastive Style Filter and modulates a frozen Depth Anything backbone through parameter-efficient, zero-initialized AdaLN-Zero layers without hurting clean-scene generalization.
Background & Motivation¶
Monocular depth estimation (MDE) serves as a corner-stone component in 3D perception for cost-effective autonomous driving, robotics, and mixed reality applications by providing geometric representations from single-view RGB streams. With the scaling of dataset curation and vision transformer architectures, recent foundation models such as the Depth Anything family achieve astonishing zero-shot generalization across heterogeneous domains. However, real-world deployment faces severe robustness bottlenecks under adverse weather conditions, including rain streaks, snowfall, dense fog scattering, and low-light nocturnal noise, which violate standard photometric and geometric assumptions, causing severe depth distortion and structural collapse.
Two dominant paradigms have attempted to bridge this robustness gap, but both encounter intrinsic tradeoffs. Two-stage pipelines deploy generalized restoration models (e.g., deraining, dehazing, or low-light enhancement) prior to depth estimation, yet these restoration models optimize for human perceptual fidelity rather than geometric consistency, frequently smoothing out high-frequency spatial cues or introducing hallucinated artifacts. On the other hand, adapting depth models directly to adverse weather typically relies heavily on synthetic corruptions due to the prohibitive acquisition cost of annotated dense ground truth in real degraded environments. As revealed by t-SNE probing of the [CLS] tokens from Depth Anything v2 encoders across real and synthetic weather scenes, existing representations fail to separate weather styles from underlying geometry, yielding a chaotic, entangled representation space where weather perturbations cannot be isolated from scene structures.
This paper's angle of attack is: rather than destructively re-training the general geometric foundation backbone, one can explicitly disentangle appearance style from structural content and inject pure degradation embeddings into the decoding stage. Core idea: train a cross-domain contrastive Style Filter that extracts content-independent, domain-aligned weather embeddings, and modulate the frozen Depth Anything foundation model using zero-initialized AdaLN-Zero layers at decoder stages, achieving robust all-weather adaptation without catastrophic forgetting of clean-scene generalization.
Method¶
Overall Architecture¶
The DepthAnything-W (DA-W) framework operates in two core modular stages: weather style extraction and condition-modulated depth regression. Given an arbitrary input RGB image (clean, synthetic, or real-world degraded), it is first fed into a dedicated Style Filter to compute multi-scale Gram matrices of low-level textural statistics, mapping the input into a compact 64-dimensional weather embedding \(\mathbf{w}\). Simultaneously, the input passes through the frozen DINOv2 vision transformer encoder to produce multi-scale geometric feature pyramids. Within the DPT decoder, lightweight AdaLN-Zero modulation layers condition intermediate features across four multi-scale fusion stages and one final pre-head layer (five sites in total). Supervised by a composite objective combining teacher distillation, synthetic-clean pair alignment, and real-weather augmentation consistency, the pipeline outputs robust, metric-faithful depth maps.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input degraded/clean image I"] --> B["Stage 1: Style Filter<br/>multi-scale Gram matrix & cross-domain contrastive extraction"]
A --> C["Frozen ViT Encoder<br/>extracts multi-scale pyramid features Fi"]
B --> D["Stage 2: Lightweight Weather Conditioning<br/>AdaLN-Zero scale/shift/gate prediction"]
C --> D
D --> E["Stage 3: Distillation & Alignment Supervision<br/>teacher guidance + paired alignment + augmentation consistency"]
E --> F["Robust Weather-Conditioned Depth Map"]
Key Designs¶
1. Style Filter: Disentangled and Domain-Aligned Weather Embedding Extraction The vulnerability of foundation depth models stems from the entanglement of weather degradations with intrinsic scene geometry, as well as the substantial domain gap between synthetic corruptions and real-world weather. To overcome this, the authors introduce a dedicated Style Filter \(F_\theta(\cdot)\) that transforms multi-scale Gram matrices into a 64-dimensional degradation-aware weather embedding \(\mathbf{w} \in \mathbb{R}^{64}\). Trained over a curated collection of 7 real-world weather datasets (15K images) and 4 synthetic degradation datasets (20K images), the model optimizes a cross-domain contrastive objective over all sample pairs \((a, b) \in \mathcal{P}\): $$ \mathcal{L}{\text{con}} = \sum}} \left[ \mathbb{I{ab} [m - d]+ + (1 - \mathbb{I} \right] $$ where }) d_{ab\(d_{ab}\) is the cosine similarity between embeddings \(\mathbf{w}_a\) and \(\mathbf{w}_b\), \(m\) is a positive distance margin, and \(\mathbb{I}_{ab}=1\) indicates that images \(a\) and \(b\) share the identical weather category regardless of domain origin. Crucially, treating synthetic and real samples of the same weather category as positive pairs forces the latent embedding space to discard semantic geometry while clustering purely by physical degradation attributes, bridging the synthetic-to-real gap.
2. Lightweight Weather Conditioning: Zero-Initialized AdaLN-Zero Decoder Modulation Full fine-tuning of depth foundation models under adverse conditions inevitably compromises their general-domain spatial representations, inducing catastrophic forgetting. DA-W keeps the entire ViT backbone frozen and introduces modulation exclusively into the DPT decoder across four multi-scale feature maps reduced to channel dimension \(C_i=64\), plus the final fused layer prior to depth regression. A compact MLP projects the 64-D weather vector \(\mathbf{w}\) into channel-wise scale \(\boldsymbol{\gamma}_i(\mathbf{w})\), shift \(\boldsymbol{\beta}_i(\mathbf{w})\), and a scalar gate \(g_i(\mathbf{w})\), updating feature maps via AdaLN-Zero: $$ F_i \leftarrow F_i + g_i(\mathbf{w}) \Big( \mathrm{LN}(F_i) \odot (1 + \boldsymbol{\gamma}_i(\mathbf{w})) + \boldsymbol{\beta}_i(\mathbf{w}) \Big) $$ where \(\mathrm{LN}(\cdot)\) represents channel-wise Layer Normalization without affine parameters. By zero-initializing the gating parameter \(g_i(\mathbf{w})\), the network identically matches the clean-domain output of Depth Anything at the onset of training. During optimization, it learns weather-specific residual shifts that counteract lighting dropouts or particle occlusions, preserving clean-scene zero-shot accuracy while empowering adverse-weather resilience.
3. Distillation & Alignment Supervision: Geometric Anchoring Without Ground-Truth Labels Acquiring dense, high-accuracy ground-truth depth maps in real adverse weather is often intractable. To circumvent this, the training framework coordinates clean images \(\mathcal{D}\), synthetic degraded images \(\mathcal{D}_s\), and unlabelled real weather images \(\mathcal{D}_r\) using affine-invariant disparity loss \(\ell(y, \hat{y})\): $$ \mathcal{L} = \mathcal{L}{\text{dist}} + \mathcal{L} $$ The distillation loss }} + \mathcal{L}_{\text{aug}\(\mathcal{L}_{\text{dist}}\) aligns the student model (ViT-S) to a frozen large teacher (ViT-L) on clean and real weather images, and anchors synthetic images \(x^s=\tau(x)\) directly to the teacher's prediction on the original clean image \(x\). Pairwise consistency \(\mathcal{L}_{\text{pair}}\) enforces prediction congruence between geometrically identical clean-synthetic pairs \((x, x^s)\), while \(\mathcal{L}_{\text{aug}}\) enforces geometric invariance on real weather samples subjected to color-jitter perturbations \(\tau_{\text{aug}}(x)\). This triad prevents unconstrained distortion and guarantees geometric fidelity.
Key Experimental Results¶
Main Results¶
Evaluation is performed across real-world adverse weather datasets (NuScenes-night, RobotCar-night, DrivingStereo rain/cloud/fog), synthetic corruptions (KITTI-C Dark, Snow, Fog, Motion), and clean benchmarks using a standard ViT-S backbone.
| Dataset / Benchmark | Metric | DepthAnything v2 | DepthAnything-AC | DA-W (Ours) | Gain vs. DA v2 |
|---|---|---|---|---|---|
| NuScenes-night (Real Night) | AbsRel ↓ / \(\delta_1\) ↑ | 0.200 / 0.725 | 0.198 / 0.727 | 0.194 / 0.737 | -0.006 / +0.012 |
| DS-rain (Real Rain) | AbsRel ↓ / \(\delta_1\) ↑ | 0.125 / 0.840 | 0.125 / 0.840 | 0.123 / 0.842 | -0.002 / +0.002 |
| DS-fog (Real Fog) | AbsRel ↓ / \(\delta_1\) ↑ | 0.103 / 0.890 | 0.103 / 0.889 | 0.101 / 0.896 | -0.002 / +0.006 |
| KITTI-C Dark (Synthetic Dark) | AbsRel ↓ / \(\delta_1\) ↑ | 0.130 / 0.832 | 0.130 / 0.834 | 0.126 / 0.837 | -0.004 / +0.005 |
| KITTI-C Snow (Synthetic Snow) | AbsRel ↓ / \(\delta_1\) ↑ | 0.115 / 0.872 | 0.114 / 0.873 | 0.107 / 0.884 | -0.008 / +0.012 |
| KITTI-C Fog (Synthetic Fog) | AbsRel ↓ / \(\delta_1\) ↑ | 0.097 / 0.905 | 0.096 / 0.906 | 0.093 / 0.910 | -0.004 / +0.005 |
| KITTI-C Motion (Synthetic Motion) | AbsRel ↓ / \(\delta_1\) ↑ | 0.127 / 0.840 | 0.126 / 0.841 | 0.125 / 0.841 | -0.002 / +0.001 |
| KITTI Clean (General Benchmark) | AbsRel ↓ / \(\delta_1\) ↑ | 0.083 / 0.934 | 0.083 / 0.934 | 0.080 / 0.937 | -0.003 / +0.003 |
On the compound mix-weather benchmark (11 composed weather conditions on KITTI Eigen split from RoboDepth), DA-W maintains consistent supremacy. Under the extreme combined corruption "Fog+Rain+Snow+Night", DA-W delivers \(\delta_1=0.630\) and \(\text{AbsRel}=0.236\), substantially surpassing DepthAnything v2 (\(0.562 / 0.268\)) and DepthAnything-AC (\(0.608 / 0.245\)).
Ablation Study¶
The impact of architectural fine-tuning schemes, real-world adverse training data inclusion, and AdaLN-Zero weather conditioning is summarized below:
| Encoder FT | Decoder FT | Real Weather Deg. | Weather Inject. | NS-night (AbsRel/\(\delta_1\)) | RC-night (AbsRel/\(\delta_1\)) | KITTI-Snow (AbsRel/\(\delta_1\)) | KITTI Clean (AbsRel/\(\delta_1\)) | Avg. rank ↓ |
|---|---|---|---|---|---|---|---|---|
| Full FT | Full FT | No | No | 0.194 / 0.727 | 0.245 / 0.493 | 0.109 / 0.888 | 0.084 / 0.932 | 2.89 |
| Frozen | Decoder FT | No | No | 0.190 / 0.741 | 0.248 / 0.495 | 0.113 / 0.877 | 0.081 / 0.936 | 2.68 |
| Frozen | Decoder FT | Yes | No | 0.195 / 0.731 | 0.243 / 0.507 | 0.112 / 0.878 | 0.081 / 0.936 | 2.29 |
| Frozen | Decoder FT | Yes | Yes (Full DA-W) | 0.194 / 0.737 | 0.239 / 0.513 | 0.107 / 0.884 | 0.080 / 0.937 | 2.14 |
In loss component ablations, removing the teacher distillation loss \(\mathcal{L}_{\text{dist}}\) leads to catastrophic degradation, with Global AbsRel escalating from 0.117 to 0.352, validating that distillation firmly grounds geometric scales. The pairwise alignment loss \(\mathcal{L}_{\text{pair}}\) further trims Global AbsRel from 0.120 to 0.117. An explicit RGB image restoration baseline (\(\lambda_{\text{rgb}}=0.1\)) achieves 0.118 Global AbsRel, demonstrating that task-oriented latent modulation outperforms naive image-level restoration.
Key Findings¶
- Frozen backbones yield superior trade-offs: Restricting optimization to the decoder while conditioning with weather vectors improves overall average rank from 2.89 (full tuning) to 2.14, fully mitigating clean-domain degradation while boosting adverse-weather performance.
- Substantial gains on snow and nocturnal scenes: KITTI-C Snow AbsRel drops by 7.0% relative margin (0.115 to 0.107), confirming that low-level scattering patterns extracted by the Style Filter effectively guide the decoder to mend occluded depth boundaries.
- Spatially adaptive feature modulation: Difference visualization of intermediate feature maps (\(|\Delta\text{Feat.}|\)) demonstrates that the AdaLN-Zero pathway naturally allocates stronger modulation intensity to severely degraded unlit patches in nighttime images and heavily scattered distant regions in foggy scenes without explicit spatial masks.
Highlights & Insights¶
- Latent condition modulation avoids perceptual reconstruction traps: Unlike two-stage pipelines that run heavy image deraining/dehazing models and risk distorting geometric cues, DA-W extracts a compact 64-D embedding to modulate depth feature spaces directly.
- Cross-domain positive pairing bridges synthetic-to-real gap: By treating real and synthetic samples sharing the same weather category as positive contrastive pairs, the Style Filter establishes a unified latent space where style distributions overlap seamlessly.
- Zero-initialized gating prevents catastrophic forgetting: Initializing modulation gates to zero guarantees that the network precisely matches the original foundation model at step zero, achieving seamless plug-and-play robustness without degrading clean benchmarks.
Limitations & Future Work¶
- Sensor-induced dynamic range shifts in night scenes: On benchmarks such as Oxford RobotCar-night characterized by extreme exposure variations, severe sensor motion blur, and non-standard camera response curves, the model slightly lags behind specialized fully-tuned baselines.
- Foundation model scale exploration: Controlled investigations in this paper concentrate on Depth Anything v2 Small (ViT-S); scaling the weather-conditioned interface across Base and Large variants requires further optimization calibration.
- Spatially heterogeneous multi-condition handling: Highly localized anomalies—such as wet road puddles reflecting glare alongside heavy snowfall—may benefit from spatially varying weather token maps beyond a single global embedding vector.
Related Work & Insights¶
- vs. Depth Anything v2 / v3: Standard foundation models rely purely on data scaling, leading to tangled weather and scene representations; DA-W explicitly decouples weather embeddings, outperforming them across 8 adverse benchmarks while preserving clean accuracy.
- vs. Depth Anything-AC / WeatherDepth: Prior adaptation methods either rely on progressive curriculum schemes or risk overriding clean-domain representations; DA-W freezes the encoder and employs zero-initialized AdaLN modulation, achieving clean parameter efficiency.
- vs. MWFormer / UniRestore: Two-stage restorative methods introduce perceptual artifacts and runtime overhead; DA-W conditions internal feature representations in an end-to-end framework, eliminating cascading error propagation.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Elegant disentanglement using a cross-domain contrastive Style Filter and zero-initialized AdaLN modulation for monocular depth models.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorously tested across 5 real adverse benchmarks, 4 synthetic corruption tracks, 11 composed mix-weather splits, and 5 clean datasets with extensive ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Highly coherent narrative, insightful feature visualizations, and sound theoretical framing.
- Value: ⭐⭐⭐⭐⭐ Establishes a practical, parameter-efficient paradigm for robustifying 3D foundation models in safety-critical autonomous systems.