Open-Weather Robust 3D Detection via Dual-Critic Diffusion Alignment¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: EventHosts PDF
Code: https://github.com/Mangonn/DCDA
Area: Autonomous Driving
Keywords: 3D Object Detection, Adverse Weather Robustness, Unseen Weather Generalization, LiDAR–4D Radar Fusion, Diffusion Alignment
TL;DR¶
Addressing LiDAR degradation under adverse weather and the closed-world limitations of existing fusion models, this paper introduces Dual-Critic Guided Diffusion Alignment (DCDA), which leverages 4D radar features to progressively restore degraded LiDAR representations toward a clean manifold without requiring paired clean-corrupted data or explicit weather labels, significantly boosting 3D detection generalization across unseen weather types and severities.
Background & Motivation¶
Adverse weather conditions—such as dense fog, heavy rain, and snowstorms—represent one of the most critical hurdles for reliable real-world autonomous driving. Particulate scattering and wet surfaces induce extensive backscatter clutter and severe beam attenuation in LiDAR point clouds, resulting in spatial sparsification and geometric distortion that precipitate dramatic drops in 3D object detection accuracy. Although millimeter-wave 4D radar offers high penetration capability and preserves complementary long-range spatial and Doppler velocity cues—prompting rapid progress in LiDAR–4D radar fusion architectures—the vast majority of existing detectors implicitly rely on a closed-world assumption: test-time environmental corruptions are assumed to mirror the weather distributions encountered during training.
However, in realistic open-world deployment, weather conditions continuously vary across both categorical types and continuous intensity levels (Open-Weather). Even within a single category such as precipitation or snowfall, variations in precipitation intensity produce substantially distinct LiDAR degradation profiles; models trained or overfitted on narrow, predefined corruption configurations exhibit severe generalization degradation when encountering novel or combined weather states.
The fundamental tension in bridging this open-weather generalization gap lies in how a perception model can restore reliable, high-fidelity geometric representations from severely distorted multi-modal signals without access to paired clean–degraded point cloud supervision or explicit domain/weather labels. The angle of attack pursued in this work is to decouple feature recovery from explicit weather modeling by reformulating restoration as a manifold alignment process. Core idea: condition a localized reverse diffusion process on weather-resilient 4D radar features, and constrain its denoising trajectory toward the clean-weather manifold using two frozen critics—a task-level detection critic for semantic preservation and an adversarial critic for holistic distributional alignment—without requiring paired data or explicit weather annotations.
Method¶
Overall Architecture¶
DCDA focuses on bird's-eye-view (BEV) feature space alignment across varying weather conditions. The overall pipeline integrates three collaborative components: a radar-conditioned diffusion aligner (\(\mathcal{A}_\theta\)), a frozen detection-guided critic (\(H\)), and a frozen weather adversarial critic (\(D\)). Taking degraded LiDAR BEV features \(F_L\) and synchronized 4D radar features \(F_R\) as inputs, the system does not generate content from pure white noise; instead, it initializes reverse diffusion from a lightly perturbed local observation \(F_T\), progressively refining corrupted LiDAR representations toward the clean manifold to output an aligned feature \(\tilde{F}_L\). During training, the dual critics enforce complementary semantic and distributional guidance; during inference, the adversarial critic is repurposed as an ultra-lightweight gating router, allowing clean frames to bypass diffusion entirely with zero overhead.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
InL["LiDAR BEV Features $F_L$"] --> Route{"Lightweight Weather Critic<br/>Evaluates Normal Confidence"}
InR["4D Radar Features $F_R$"] --> DiffAlign
Route -->|Normal Confidence $\ge \tau$ Bypass| OutFuse["Multimodal Fusion Features"]
Route -->|Adverse Weather $<\tau$ Activate| DiffAlign["Radar-Conditioned Diffusion Aligner<br/>Iterative Reverse Denoising"]
DiffAlign --> OutRefined["Refined LiDAR Features $\tilde{F}_L$"]
OutRefined --> OutFuse
OutRefined -.->|Gradients during training| TaskCritic["Detection-Guided Critic<br/>Enforces Task Semantics & Geometry"]
OutRefined -.->|Gradients during training| AdvCritic["Weather Adversarial Critic<br/>Enforces Global Distribution Alignment"]
OutFuse --> DetHead["Downstream 3D Detection Head"]
Key Designs¶
1. Radar-Conditioned Diffusion Alignment: Local Manifold Geometric Recovery
Directly mapping degraded features via single-step regressors or standard GANs often causes models to overfit training-specific artifacts or introduce hallucinated geometric structures. DCDA models feature recovery as a few-step conditional diffusion process. In the forward chain, the input LiDAR feature \(F_0 \equiv F_L\) is perturbed across \(T\) steps into \(F_T\). Unlike unconditional generative models that sample from arbitrary Gaussian noise, DCDA initiates its reverse trajectory from the perturbed observation \(F_T \sim q(F_T | F_0)\), restricting denoising to a local neighborhood around \(F_0\). This preserves global scene topology and suppresses generative hallucination while requiring only \(T=3\) steps for fast convergence.
At each reverse step, a conditional U-Net \(\mathcal{A}_\theta\) takes the channel-concatenated noisy feature \(F_t\) and radar feature \(F_R\) to directly predict the clean manifold estimate \(\hat{F}_0^{(t)} = \mathcal{A}_\theta([F_t; F_R], t)\). To mitigate negative transfer from radar sensor noise and low angular resolution while keeping features anchored, a self-reconstruction objective is enforced: $$ \mathcal{L}_{\text{diff}} = |\tilde{F}_L - F_L|_2^2 $$ This term functions as a content-preserving regularizer, leaving the pull toward the clean manifold entirely to the dual critics.
2. Dual-Critic Guidance Mechanism: Orthogonal Semantic and Distributional Constraints
Neither radar conditioning nor self-reconstruction alone can direct corrupted features toward the unobserved clean manifold. DCDA establishes two frozen critics as manifold guidance anchors. First, the Detection-Guided Critic reuses a 3D detection head \(H\) pretrained exclusively on Normal weather and kept frozen. Feeding the refined feature \(\tilde{F}_L\) through \(H\) yields the standard task loss \(\mathcal{L}_{\text{det}} = \mathcal{L}_{\text{cls}} + \mathcal{L}_{\text{reg}}\). Because \(H\) is tuned only to sharp, clear-weather targets, minimizing \(\mathcal{L}_{\text{det}}\) pulls \(\tilde{F}_L\) into the feature subspace where objects remain discriminative and spatially precise.
Second, the Weather Adversarial Critic enforces holistic representation alignment. During pretraining, discriminator \(\mathcal{D}\) learns to separate Normal from non-Normal LiDAR BEV representations via binary cross-entropy. Once frozen, \(\mathcal{D}\) defines an adversarial guidance objective for the refined representations: $$ \mathcal{L}{\text{adv}} = -\mathbb{E}_L)] $$ During optimization of }_L}[\log \mathcal{D}(\tilde{F\(\mathcal{A}_\theta\), gradients flow backward through both frozen heads to form \(\mathcal{L}_{\text{crit}} = \mathcal{L}_{\text{det}} + \mathcal{L}_{\text{adv}}\). The detection critic sharpens local instance boundaries, while the adversarial critic erases diffuse background degradation patterns and signal attenuation across the BEV plane.
3. Two-Stage Optimization and Lightweight Routed Inference: Balancing Convergence and Latency
Optimizing diffusion alignment alongside dual critics follows a disciplined two-stage curriculum. In Stage I (Normal Weather Prior Warm-up), \(\mathcal{A}_\theta\) is trained solely on clear-weather samples under \(\mathcal{L}_{\text{diff}} + \mathcal{L}_{\text{det}}\), establishing stable identity-preserving reconstruction and radar-guided feature routing. In Stage II (Adversarial Alignment), \(\mathcal{A}_\theta\) is fine-tuned over the full training distribution combining Normal and seen adverse weather: $$ \mathcal{L}{\text{total}} = \mathcal{L}}} + \lambda_1 \mathcal{L{\text{det}} + \lambda_2 \mathcal{L} $$ Throughout Stage II, the weight on }\(\mathcal{L}_{\text{diff}}\) is gradually annealed toward zero, shifting primary optimization momentum to the critic objectives.
During online deployment, running multi-step diffusion on every frame introduces undesirable computational latency. DCDA repurposes the pretrained discriminator \(\mathcal{D}\) as an inference router. By computing \(s = \mathcal{D}(F_L) \in [0, 1]\), frames where \(s \ge \tau\) are identified as clear weather and bypass diffusion entirely. Only corrupted frames with \(s < \tau\) trigger DCDA. Operating with 99.8% routing accuracy, this mechanism eliminates unnecessary computation on clear frames without sacrificing detection accuracy.
Key Experimental Results¶
Main Results¶
Evaluation is performed on K-Radar, covering seven real-world weather conditions: Normal, Overcast, Fog, Rain, Sleet, LightSnow, and HeavySnow. Under the open-weather protocol, models observe only Normal, Rain, and Sleet during training, and are evaluated on unseen conditions including Fog, HeavySnow, LightSnow, and Overcast (Sedan category, IoU = 0.5).
Table 1: Benchmark comparison on K-Radar type-open protocol (BEV AP / 3D AP, %)
| Method | Seen (Normal) | Unseen (Fog) | Unseen (HeavySnow) | Unseen (LightSnow) | Unseen (Overcast) | Unseen Macro Mean |
|---|---|---|---|---|---|---|
| RTNH | 35.64 / 21.69 | 72.43 / 10.36 | 31.63 / 11.12 | 61.93 / 9.79 | 64.96 / 28.57 | 57.74 / 14.96 |
| InterFusion | 67.65 / 42.88 | 62.05 / 34.26 | 24.36 / 11.52 | 56.40 / 28.96 | 71.22 / 40.51 | 53.51 / 28.81 |
| SpikingRTNH | 30.48 / 11.50 | 66.72 / 23.40 | 31.27 / 19.23 | 47.66 / 8.07 | 56.25 / 29.96 | 50.48 / 20.17 |
| V2X-R | 65.96 / 41.61 | 69.43 / 28.65 | 30.37 / 20.70 | 66.72 / 21.87 | 78.99 / 36.44 | 61.38 / 26.92 |
| L4DR | 68.33 / 44.71 | 81.32 / 36.73 | 31.74 / 21.88 | 74.81 / 29.60 | 80.42 / 49.85 | 67.07 / 34.52 |
| DCDA (Ours) | 68.39 / 49.92 | 88.81 / 41.91 | 35.94 / 26.29 | 77.07 / 35.09 | 81.76 / 65.70 | 70.90 / 42.25 |
Table 2: Real weather severity transfer benchmark (LightSnow \(\to\) HeavySnow, %)
| Method | Training Configuration | Clear Reference (Normal) | Extreme Blizzard Test (HeavySnow) | Relative Gain on HeavySnow |
|---|---|---|---|---|
| L4DR | Real Normal + LightSnow | 71.61 / 51.20 | 30.57 / 16.57 | Baseline |
| DCDA (Ours) | Real Normal + LightSnow | 68.39 / 50.01 | 37.92 / 25.61 | +7.35 BEV / +9.04 3D |
Ablation Study¶
Table 3: Ablation of guidance objectives on K-Radar type-open (macro mean, %)
| \(\mathcal{L}_{\text{diff}}\) | \(\mathcal{L}_{\text{det}}\) | \(\mathcal{L}_{\text{adv}}\) | Seen BEV AP | Seen 3D AP | Unseen BEV AP | Unseen 3D AP | Note |
|---|---|---|---|---|---|---|---|
| \(\checkmark\) | 61.76 | 43.12 | 66.23 | 39.11 | Self-reconstruction only | ||
| \(\checkmark\) | \(\checkmark\) | 66.68 | 44.95 | 70.47 | 41.50 | Adds detection-guided critic | |
| \(\checkmark\) | \(\checkmark\) | 63.97 | 44.06 | 68.22 | 40.99 | Adds weather adversarial critic | |
| \(\checkmark\) | \(\checkmark\) | \(\checkmark\) | 66.77 | 45.47 | 72.65 | 42.25 | Full dual-critic guided DCDA |
Key Findings¶
- Dual critics provide indispensable, complementary alignment: Relying solely on diffusion self-reconstruction (\(\mathcal{L}_{\text{diff}}\)) yields only 39.11% unseen 3D AP. Incorporating the detection critic adds +2.39% 3D AP by enforcing task semantics, while the adversarial critic adds +1.88% 3D AP via global distribution matching. Integrating both achieves 42.25% 3D AP, verifying their orthogonal benefits.
- Robustness in severe unseen conditions: In extreme HeavySnow where LiDAR reflections are severely degraded by backscatter noise, standard baselines degrade to roughly 20% 3D AP, whereas DCDA achieves 26.29% (+4.41% over L4DR). In the LightSnow-to-HeavySnow transfer experiment, DCDA improves 3D AP by +9.04 percentage points.
- Efficient inference via adaptive routing: Base inference without DCDA runs at 29.51 ms (33.89 FPS); executing 3-step diffusion on every frame increases latency to 70.86 ms (14.11 FPS). By activating DCDA only when the weather critic indicates corruption, latency drops to 59.59 ms (16.78 FPS) with zero degradation on clear days and 70.90% unseen BEV AP.
Highlights & Insights¶
- Localized few-step diffusion alignment: Rather than generating features from random noise, DCDA initiates reverse diffusion from the perturbed corrupted input, bounding the trajectory to a local manifold neighborhood and enabling high-fidelity restoration in just 3 steps.
- Dual-role discriminator as zero-cost inference router: The adversarial critic trained to match representations during training serves directly as a gating mechanism at test time, bypassing diffusion under clear conditions without requiring extra parameters or classification heads.
- Label-free and unaligned cross-weather generalization: The model operates without paired clean-degraded frames, physical simulation priors, or categorical weather labels, relying purely on self-supervised diffusion alignment anchored by frozen clean critics.
Limitations & Future Work¶
- Ceiling under near-total signal loss: When heavy fog or severe snow extinguishes virtually all returning LiDAR points, the sparse spatial resolution of 4D radar alone cannot fully reconstruct fine 3D bounding geometry, causing smaller relative gains in extreme settings.
- Synthetic severity proxy reliance: Due to the scarcity of dense severity annotations in real datasets, parts of the continuous intensity benchmark rely on physically simulated perturbations; validation on larger-scale continuous fleets remains desirable.
- Extension to vision-inclusive sensor suites: Expanding the dual-critic diffusion paradigm across camera-LiDAR-radar triplets could leverage rich optical semantics to better compensate for radar clutter in severe weather.
Related Work & Insights¶
- vs L4DR (ECCV 2024): L4DR utilizes heuristic radar denoising filters and gated cross-attention, which performs well on known weather types but lacks generative mechanisms to restore distorted LiDAR manifolds under unseen corruptions; DCDA surpasses L4DR by +7.73% 3D AP on unseen weather.
- vs V2X-R (CVPR 2024): V2X-R adopts diffusion denoising with radar conditioning but lacks explicit clean-manifold guidance, leaving feature trajectories vulnerable to semantic drift; DCDA resolves this through frozen detection and adversarial critics.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneers dual-critic manifold diffusion alignment for open-weather 3D detection without paired data]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Systematic evaluation across weather types, continuous severities, real transfers, and ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear structural organization, precise formulations, and consistent methodology narrative]
- Value: ⭐⭐⭐⭐⭐ [Directly tackles adverse-weather out-of-distribution perception, offering practical architectural insights for autonomous driving]