HVA-Fusion:Hierarchical Velocity-Aware 4D Radar-LiDAR Fusion for Robust 3D Object Detection¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/wujingxian1/HVA-Fusion
Area: Autonomous Driving
Keywords: 4D Imaging Radar, LiDAR-Radar Fusion, 3D Object Detection, Multi-Modal Fusion, Adverse Weather Robustness
TL;DR¶
Addressing the severe degradation of LiDAR in adverse weather and the sparsity and non-representative Doppler radial velocity of 4D radar, this paper proposes HVA-Fusion, a hierarchical velocity-aware fusion framework that estimates true motion velocity, performs pillar-level adaptive interaction, densifies features via RCS-guided Gaussian diffusion, and fuses BEV features using multi-scale bidirectional deformable attention gates.
Background & Motivation¶
In 3D object detection for autonomous driving and mobile robotics, LiDAR-based systems have long served as the primary perceptual modality due to their capability to reconstruct high-fidelity metric geometry. However, near-infrared laser pulses suffer catastrophic scattering and attenuation under adverse weather conditions such as fog, rain, snow, and dust, leading to severe point cloud degradation. In contrast, 4D millimeter-wave imaging radar operates at much longer wavelengths, exhibiting superior all-weather penetration while simultaneously providing micro-Doppler radial velocity and Radar Cross-Section (RCS) information. Nevertheless, hardware aperture constraints result in coarse angular resolution, pronounced sparsity, and pervasive multipath noise in raw 4D radar measurements.
Combining the complementary strengths of both modalities is a compelling avenue for robust all-weather perception, yet existing fusion paradigms face critical bottlenecks. First, prior approaches routinely treat Doppler velocity via shallow operations (e.g., taking absolute magnitudes or squaring), ignoring that radial velocity is merely a line-of-sight projection. When a dynamic object moves laterally across the sensor, its radial velocity drops near zero, rendering it indistinguishable from static background clutter. Second, LiDAR and 4D radar exhibit substantial discrepancies in point density, precision, and physical properties; existing early-fusion or single-feature-space methods rely heavily on rigid spatial alignment, failing to rectify calibration offsets and local sensor misalignments. Furthermore, 4D radar's extreme sparsity causes its features to be dominated by LiDAR in shared feature spaces, while temporal multi-frame accumulation introduces severe motion blur and spatial displacement in dynamic traffic scenes. Without dynamic gating, sensor degradation under severe weather or clutter inevitably triggers negative transfer.
This paper tackles these challenges by explicitly recovering true object velocity from physical motion priors and establishing a progressive two-stage fusion pipeline spanning local pillar adaptation to global deformable BEV interaction. Core idea: recover true motion velocities via point-wise heading prediction while filtering radar clutter, adaptively perform local cross-modal attention in pillar space guided by radar density, densify radar representations in BEV space via RCS- and range-aware Gaussian diffusion, and achieve robust, misalignment-resilient multi-modal fusion through a multi-scale bidirectional deformable attention gate.
Method¶
Overall Architecture¶
The input to HVA-Fusion consists of raw 4D radar point clouds and LiDAR point clouds. First, the Motion Velocity Estimation and Encoding (MVEE) module predicts point-wise motion headings to recover true motion velocities and filters background clutter. Next, points are mapped into pillars, where the Cross-Modal Progressive Adaptation (CPA) module performs local cross-modal propagation and density-aware neighborhood attention (Stage-I local fusion). To address radar feature sparsity prior to BEV projection, the RCS-Guided Gaussian Generation (RGG) module applies physics-informed adaptive Gaussian diffusion to densify radar representations. Finally, in BEV space, the Multi-Scale Bi-Directional Deformable Attention Gate (MS-BDDG) module (Stage-II global fusion) dynamically eliminates spatial misalignments and suppresses noise from degraded modalities via gating networks before feeding the fused features to the detection head for 3D bounding box regression and classification.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: 4D Radar Point Cloud + LiDAR Point Cloud"] --> B["Motion Velocity Estimation and Encoding (MVEE)<br/>Heading prediction for true velocity + foreground filtering"]
B --> C["Pillar Encoding & Feature Extraction"]
C --> D["Cross-Modal Progressive Adaptation (CPA)<br/>Density-aware adaptive k-NN + local cross-attention"]
D --> E["RCS-Guided Gaussian Generation (RGG)<br/>Physical & learnable scale fusion + range-aware Gaussian diffusion"]
E --> F["Multi-Scale Bi-Directional Deformable Attention Gate (MS-BDDG)<br/>Three-layer pyramid deformable interaction + adaptive gating"]
F --> G["Detection Head for 3D Bounding Boxes & Categories"]
Key Designs¶
1. Motion Velocity Estimation and Encoding (MVEE): Explicitly recovering true dynamic velocity from line-of-sight radial projections
Raw 4D radar point clouds provide relative radial velocity \(v_r\) and ego-compensated absolute radial velocity \(v_a\), which reflect only the velocity component projected along the sensor line-of-sight direction \(\beta = \arctan(y/x)\). When the true motion direction \(\alpha\) deviates from \(\beta\) (e.g., a pedestrian crossing laterally), \(v_a\) degrades toward zero, destroying dynamic discriminability. MVEE employs a PointNet++ architecture to estimate the point-wise motion direction \(\alpha\) (relative to the positive x-axis). During training, the ground-truth 3D bounding box orientation \(\alpha_{\text{heading}}\) supervises direction prediction via the direction loss \(L_{\text{dir}}\): $\(L_{\text{dir}} = \frac{1}{N} \sum_{i=1}^{N} (1 - \cos(\alpha_i - \alpha_{\text{heading}}))\)$ With the deviation angle \(\theta = \alpha - \beta\), the true motion velocity scalar \(v_{\text{pred}}\) is explicitly derived as: $\(v_{\text{pred}} = \begin{cases} \frac{v_a}{\cos\theta}, & \text{if } |\cos\theta| > 10^{-1} \\ 10 \cdot v_a, & \text{otherwise} \end{cases}\)$ This estimated true velocity is integrated into the point features as a dynamic descriptor. Simultaneously, a segmentation head predicts foreground probabilities to remove clutter points below a confidence threshold, substantially enhancing the visibility of vulnerable road users like pedestrians and cyclists against complex backgrounds.
2. Cross-Modal Progressive Adaptation (CPA): Pillar-level local adaptive neighborhood interaction guided by radar density
To overcome the sensitivity of rigid spatial concatenation to calibration inaccuracies and local sensor drift, CPA decomposes local integration into co-located propagation and density-guided adaptive neighborhood matching. Within identical pillar coordinates, the Multi-Modal Encoder (MME) calculates mutual distance statistics and cross-modal feature averages, producing enriched pillar features \(f^l\) and \(f^r\). Because radar points are distributed heterogeneously, CPA calculates the local radar pillar density within a search radius \(r_l\) centered at each LiDAR pillar. This density dynamically modulates the sampling size \(k_l\) for adaptive k-NN neighborhood search. LiDAR pillar features serve as Queries, while the selected neighboring radar features serve as Keys and Values in cross-attention: $\(\hat{f}_i^l = \sum_{j \in \mathcal{N}_i^r} \text{Softmax}\left(\frac{f_i^l (f_j^r)^T}{\sqrt{d}}\right) f_j^r\)$ Regions with dense radar support provide stronger foreground confidence to LiDAR representations, while sparse clutter regions receive constrained enhancement, flexibly bridging the semantic and geometric gaps between modalities.
3. RCS-Guided Gaussian Generation (RGG): Densifying radar representations via physical size priors and learnable range modulation
To prevent sparse radar features from being overwhelmed by dense LiDAR representations, RGG exploits the physical positive correlation between an object's physical dimensions and its Radar Cross-Section (RCS), diffusing discrete radar pillar features into continuous fields via isotropic Gaussian kernels. The diffusion scale \(\sigma_i\) is jointly determined by physical priors and neural refinement: a baseline scale \(\sigma_i^{\text{base}}\) provides a monotonically increasing mapping from RCS via a Sigmoid curve; a learnable refinement term \(\sigma_i^{\text{learn}}\) dynamically adapts to non-linear scene attributes via an MLP; and a normalized Euclidean distance factor \(d_i\) modulates scale according to range-dependent sparsity: $\(\sigma_i = \text{clip}\left(d_i \cdot \left(\gamma \sigma_i^{\text{base}} + (1 - \gamma) \sigma_i^{\text{learn}}\right), \sigma_{\min}, \sigma_{\max}\right)\)$ For each spatial location, features are aggregated across all overlapping radar pillars via normalized Gaussian weighting \(f^g = \frac{\sum w_i f_i^r}{\sum w_i}\), filtering weights below 0.1 and projecting them through a lightweight MLP into dense BEV radar features \(F^r\). This avoids temporal motion blur while preserving spatial integrity for distant targets.
4. Multi-Scale Bi-Directional Deformable Attention Gate (MS-BDDG): Global BEV fusion with alignment rectification and degradation gating
In BEV space, residual calibration inaccuracies and azimuthal noise can induce spatial misalignment between modalities. MS-BDDG constructs a three-layer feature pyramid (\(D \in \{1, 2, 3\}\)), where shallower layers capture geometric complementarity and deeper layers consolidate semantic consensus. Within each scale, deformable cross-attention dynamically establishes cross-modal correspondences by predicting learnable sampling offsets \(\Delta p_{hqk}\), alternately treating LiDAR and radar features as Queries and Values. To guard against negative transfer under severe weather (e.g., heavy rain blinding LiDAR or multipath clutter polluting radar), an adaptive gating network predicts channel-wise spatial modulation weights: $\(w^m = \text{Sigmoid}(\text{Conv}_{3\times3}(F_D^f)), \quad m \in \{l, r\}\)$ These weights modulate single-modality features via element-wise multiplication, suppressing noise from degraded sensors. Finally, multi-scale features are concatenated, enabling resilient global feature integration.
Loss & Training¶
The framework is trained end-to-end with a multi-task objective: $\(L_{\text{all}} = \beta_{\text{cls}} L_{\text{cls}} + \beta_{\text{loc}} L_{\text{loc}} + \beta_{\text{mask}} L_{\text{mask}} + \beta_{\text{dir}} L_{\text{dir}}\)$ with loss weights set to \(\beta_{\text{cls}} = 1.0, \beta_{\text{loc}} = 2.0, \beta_{\text{mask}} = 2.0, \beta_{\text{dir}} = 1.1\). Focal Loss is used for classification and foreground masking, while Smooth L1 Loss optimizes bounding box regression and direction estimation. Furthermore, a Random Modal Dropout strategy is applied during training: with probability \(P_{\text{drop}} = 0.2\), single-modality feature maps are randomly zeroed out, compelling the network to retain detection competence even under total sensor failure.
Key Experimental Results¶
Main Results¶
HVA-Fusion is evaluated across three major benchmark datasets: K-Radar, View-of-Delft (VoD), and TJ4DRadSet.
1. 3D Object Detection on the K-Radar Adverse Weather Dataset (Sedan class, IoU=0.5, AP in %)
| Method | Modality | AP3D (Total) | Normal | Fog | Rain | Sleet | Lightsnow | Heavysnow |
|---|---|---|---|---|---|---|---|---|
| RTNH | 4D Radar | 37.4 | 37.6 | 41.2 | 29.2 | 49.1 | 63.9 | 43.1 |
| PointPillars | LiDAR | 22.4 | 21.8 | 28.2 | 27.2 | 22.6 | 23.2 | 12.9 |
| RTNH | LiDAR | 37.8 | 39.8 | 59.8 | 28.2 | 31.4 | 50.7 | 24.6 |
| InterFusion | LiDAR+Radar | 17.5 | 15.3 | 47.6 | 12.9 | 9.33 | 56.8 | 25.7 |
| 3D-LRF | LiDAR+Radar | 45.2 | 45.3 | 51.8 | 38.3 | 23.4 | 60.2 | 36.9 |
| L4DR | LiDAR+Radar | 53.5 | 53.0 | 73.2 | 53.8 | 46.2 | 52.4 | 37.0 |
| HVA-Fusion (Ours) | LiDAR+Radar | 65.7 | 63.8 | 79.7 | 65.5 | 60.7 | 77.1 | 57.7 |
2. 3D Object Detection on the View-of-Delft (VoD) Validation Set (AP in %)
| Method | Modality | Entire Car | Entire Ped. | Entire Cyc. | Entire mAP | Driving Car | Driving Ped. | Driving Cyc. | Driving mAP |
|---|---|---|---|---|---|---|---|---|---|
| PointPillars | LiDAR | 65.55 | 55.71 | 72.96 | 64.74 | 81.10 | 67.92 | 88.96 | 79.33 |
| InterFusion | L+R | 67.50 | 63.21 | 78.79 | 69.83 | 88.11 | 74.80 | 87.50 | 83.47 |
| MutualForce | L+R | 71.67 | 66.26 | 77.35 | 71.76 | 92.31 | 76.79 | 89.97 | 86.36 |
| L4DR | L+R | 69.10 | 66.20 | 82.80 | 72.70 | 90.80 | 76.10 | 95.50 | 87.47 |
| SVEFusion | L+R | 71.83 | 67.85 | 83.97 | 74.55 | 90.91 | 79.37 | 96.31 | 88.86 |
| ELMAR | L+R | 76.41 | 69.34 | 78.91 | 74.89 | 91.74 | 81.35 | 93.01 | 88.70 |
| HVA-Fusion (Ours) | L+R | 70.12 | 72.46 | 83.96 | 75.51 | 90.91 | 78.03 | 96.51 | 88.48 |
Ablation Study¶
The individual and cumulative impacts of core components were ablated on the VoD dataset, using L4DR (direct BEV feature concatenation) as the baseline:
| Config | MVEE | CPA | RGG | MS-BDDG | Entire Area mAP (%) | Driving Area mAP (%) | Key Findings |
|---|---|---|---|---|---|---|---|
| 1 (Baseline) | 70.99 | 83.95 | L4DR baseline with concatenation | ||||
| 2 | ✓ | 73.36 (+2.37) | 85.35 (+1.40) | Motion velocity estimation significantly aids dynamic targets | |||
| 3 | ✓ | ✓ | 73.71 (+0.35) | 88.14 (+2.79) | Pillar-level adaptive interaction yields large driving area gain | ||
| 4 | ✓ | ✓ | ✓ | 74.11 (+0.40) | 88.24 (+0.10) | RCS-guided Gaussian diffusion relieves distant sparsity | |
| 5 | ✓ | ✓ | 74.02 | 88.01 | BEV alignment alone without local progressive adaptation | ||
| 6 | ✓ | ✓ | ✓ | 74.20 | 87.92 | Omitting CPA harms fine-grained pillar alignment | |
| 7 | ✓ | ✓ | ✓ | 74.77 | 88.43 | Omitting RGG leaves radar representations sparse | |
| 8 (Full model) | ✓ | ✓ | ✓ | ✓ | 75.51 (+4.52) | 88.48 (+4.53) | All modules cooperate to establish new state-of-the-art |
Key Findings¶
- Unrivaled Resilience Under Severe Weather: In the heavy snow scenario of K-Radar, LiDAR-only detection collapses to 12.9%~24.6% AP3D. HVA-Fusion achieves 57.7% AP3D, outperforming prior state-of-the-art L4DR (37.0%) by +20.7%, demonstrating the efficacy of adaptive gating against severe environmental degradation.
- Breakthrough in Vulnerable Road User Detection: In the VoD Entire Area evaluation, pedestrian AP reaches 72.46%, surpassing previous methods by 3%~6%. Explicitly decoupling true motion velocity resolves the long-standing limitation where lateral-moving objects produce near-zero radial Doppler signals.
- Hierarchical Synergy Between Pillar and BEV Stages: Ablation indicates that relying solely on BEV deformable attention without pillar-level CPA drops mAP to 74.20%, whereas omitting BEV attention drops it to 74.11%. Combining local progressive adaptation with global deformable interaction achieves 75.51%, validating the two-stage hierarchical paradigm.
- Inference Latency: Running on a single NVIDIA H20 GPU, HVA-Fusion requires 148 ms total latency (85 ms for baseline, 26 ms for MVEE, 5 ms for CPA, 4 ms for RGG, and 28 ms for MS-BDDG), maintaining manageable overhead for production autonomous systems.
Highlights & Insights¶
- Geometric Decoupling of Radial Doppler into True Velocity: Predicting object motion headings to formulate inverse cosine relations with the sensor's line-of-sight angle transforms ambiguous raw Doppler measurements into physically faithful dynamic motion vectors.
- Physics-Informed Gaussian Densification Without Motion Blur: Rather than relying on multi-frame accumulation that causes dynamic trailing artifacts, RGG modulates Gaussian kernels using physical RCS magnitudes and range priors, densifying radar representations within a single temporal frame.
- Two-Stage Multi-Scale Deformable Gating: Combining local density-aware neighborhood attention at the pillar level with deformable cross-attention at the BEV level systematically resolves multi-sensor spatial misalignments, while adaptive gates suppress corrupted features.
Limitations & Future Work¶
- Isotropic Kernels Blurring Asymmetric Targets: The isotropic Gaussian formulation in RGG tends to blur high-aspect-ratio boundaries (such as pedestrians in close proximity), capping gains in the near-range driving area; exploring anisotropic, orientation-aware Gaussian kernels represents an important next step.
- Challenging Micro-Doppler Signatures in Extreme Fog: Under simulated level-4 fog on VoD-Fog, cyclists exhibit intricate Doppler patterns from rotating wheels and metal frames. When LiDAR points are heavily attenuated, CPA's local matching can occasionally introduce unreliable associations; integrating Signal-to-Noise Ratio (SNR)-adaptive interaction weighting is an active area of future development.
Related Work & Insights¶
- vs L4DR (AAAI 2025): L4DR conducts basic pillar interactions and simple BEV concatenation. HVA-Fusion introduces MVEE velocity estimation, CPA adaptive neighborhood aggregation, RGG Gaussian densification, and MS-BDDG deformable gating, pushing K-Radar all-weather AP3D from 53.5% to 65.7%.
- vs InterFusion / 3D-LRF (CVPR 2024): Previous methods rely on raw radial Doppler scalars and single-feature-space interactions. HVA-Fusion establishes a cohesive hierarchy from point-level motion recovery to pillar-level adaptation and BEV-level deformable gating, setting a new standard for weather-robust multi-modal perception.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Directly estimating true dynamic velocity and formulating RCS-guided Gaussian densification effectively tackle core 4D radar bottlenecks]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Exhaustively validated across K-Radar adverse weather, VoD real and fog-simulated scenes, and TJ4DRadSet with clear ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Well-structured narrative with tight alignment between motivation, architectural design, and experimental evidence]
- Value: ⭐⭐⭐⭐⭐ [Delivers a robust, high-performance two-stage framework that significantly advances all-weather autonomous driving perception]