Skip to content

Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation

Conference: ECCV 2026
Paper: ECCV Official Link
Area: Segmentation
Keywords: Radar semantic segmentation, multi-view representation learning, learnable hypergraph, unbalanced optimal transport, structural consistency

TL;DR

To tackle high measurement sparsity, multi-path return fragmentation, and inter-view measurement density imbalance in multi-view radar semantic segmentation, HyperRadar introduces learnable hypergraph reasoning to capture higher-order dependencies among grouped returns and leverages Unbalanced Optimal Transport (UOT) for mass-relaxed, correspondence-free cross-view feature alignment, establishing new state-of-the-art benchmarks on CARRADA (63.8% mIoU) and RADIal (83.4% mIoU).

Background & Motivation

Millimeter-wave frequency-modulated continuous-wave (FMCW) radar operates reliably under adverse weather (such as rain, fog, and snow), severe glare, and extreme low-illumination conditions, making it an indispensable sensing modality for autonomous driving perception stacks. Nevertheless, compared to dense RGB imagery and precise geometric LiDAR point clouds, automotive radar returns are notoriously sparse, corrupted by speckle noise and multi-path ghost targets, and convey weak semantic signatures. Prior radar perception architectures predominantly convert the full 3D Range-Azimuth-Doppler (RAD) tensor into 2D projections (such as Range-Doppler [RD], Range-Azimuth [RA], and Angle-Doppler [AD] maps) and process them using standard convolutional networks (CNNs) or pairwise attention mechanisms. However, standard dense prediction pipelines inherently rely on the assumption of continuous, locally smooth, and informative visual patterns—an assumption that severely breaks down in sparse, scattered radar imagery.

The fundamental tension stems from a mismatch between physical radar signal formation and conventional neural interaction paradigms. When a physical object (e.g., an extended vehicle) interacts with radar waves, multiple electromagnetic reflections span across several range, azimuth, and Doppler cells. Conventional local convolutional kernels possess restricted receptive fields, while standard graph neural networks and attention modules are constrained to pairwise interactions, failing to explicitly model the higher-order group relational dependencies among dispersed radar returns originating from the same object. Furthermore, when radar signals are projected into heterogeneous 2D domains, different antenna apertures and physical integration mechanisms induce severe cross-view asymmetry. An extended vehicle producing a compact, high-energy 3-bin peak in the RD view smears across 15+ bins in the RA view. Traditional balanced optimal transport or rigid geometric projections enforce strict mass conservation, making them extremely fragile to view-dependent occlusion, clutter, and missing returns.

This paper tackles these challenges from two coupled perspectives: first, hypergraph formulation is deployed within each view to gather grouped reflections into learnable hyperedges, capturing region-level higher-order topology; second, Unbalanced Optimal Transport (UOT) is integrated to align cross-view latent representations without requiring explicit spatial correspondences, penalizing marginal distribution discrepancies via Kullback-Leibler divergences to gracefully absorb mass discrepancies. Core idea: HyperRadar establishes a unified framework that couples learnable hypergraph-based higher-order relational reasoning with unbalanced optimal transport cross-view alignment, fundamentally resolving the dual bottlenecks of fragmented radar returns and inter-view measurement density mismatch.

Method

Overall Architecture

HyperRadar takes as input three 2D projection tensors sliced from the 3D RAD representation over a temporal window \(T\): Range-Doppler (RD), Range-Azimuth (RA), and Angle-Doppler (AD). For each view, a dedicated backbone encoder extracts dense latent feature tokens. Each view-specific token set is then refined through a learnable hypergraph neural block that clusters fragmented radar peaks into non-local higher-order structures. Next, for each supervised prediction branch (RD and RA), an Unbalanced Optimal Transport (UOT) fusion module solves a regularized transport plan using an iterative Sinkhorn solver, transferring complementary features across views under relaxed mass conservation constraints. Finally, the fused tokens pass through cascaded adaptive channel attention gates to suppress residual radar clutter, and shallow decoders generate dense semantic segmentation masks supervised by multi-class segmentation losses and cross-view range consistency regularizations.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multi-View Radar Projections<br/>(RD / RA / AD 2D Tensors)"] --> B["View-Specific Encoders<br/>Extract Latent Feature Tokens"]
    B --> C["Learnable Hypergraph Feature Refinement<br/>Soft Incidence Matrix & Higher-Order Context"]
    C --> D["Unbalanced Optimal Transport Cross-View Alignment<br/>Sinkhorn Solver & Mass-Relaxed Feature Fusion"]
    D --> E["Adaptive Attention Gating & Decoding<br/>Cascaded Channel Filtering & Semantic Mask Output"]
    E --> F["Joint Consistency Regularization & Training<br/>Dice + CE + Shared Range-Axis Constraints"]

Key Designs

1. Learnable Hypergraph Feature Refinement: Modeling Non-Local Higher-Order Dependencies

Standard graph neural networks and pairwise self-attention connect pairs of nodes, whereas extended objects in radar scans produce clusters of correlated reflections across disparate range and Doppler bins. To explicitly capture multi-way dependencies, HyperRadar builds a hypergraph \(\mathcal{G}^{(v)} = (\mathcal{V}, \mathcal{E}^{(v)})\) for each view \(v \in \{\text{RD}, \text{RA}, \text{AD}\}\), comprising \(N_v\) latent tokens \(X^{(v)} \in \mathbb{R}^{N_v \times C}\) and \(M\) learnable hyperedges. Instead of relying on static spatial heuristics, the model computes a soft incidence matrix via a trainable projection matrix \(W_h \in \mathbb{R}^{C \times M}\):

\[H^{(v)} = \text{softmax}(X^{(v)} W_h) \in \mathbb{R}^{N_v \times M}\]

where the softmax normalizes along the hyperedge dimension, allowing each token to probabilistically distribute its membership across multiple semantic hyperedges. Hyperedge embeddings are formed by aggregating node features \(Z^{(v)} = (H^{(v)})^\top X^{(v)}\), transformed through a lightweight two-layer MLP \(\phi(\cdot)\), and projected back into token space to construct a residual update:

\[X_{\text{hg}}^{(v)} = X^{(v)} + \gamma H^{(v)} \phi\left((H^{(v)})^\top X^{(v)}\right)\]

where \(\gamma\) is a learnable scaling factor. This residual hypergraph mechanism enables discrete radar returns separated by noise or low-amplitude bins to exchange context through shared hyperedges, substantially boosting object-level structural coherence under extreme measurement sparsity.

2. Unbalanced Optimal Transport Cross-View Alignment: Correspondence-Free Mass-Relaxed Fusion

Radar measurement physics dictates significant energy distribution discrepancies across projection planes. While the RD projection integrates energy across the entire receiver antenna array to yield sharp, localized object peaks, the RA projection integrates over limited angular apertures, smearing target energy across broader azimuth bins. Classical balanced Optimal Transport enforces strict total mass preservation, causing severe distortion when aligning views with unequal return densities and clutter. HyperRadar addresses this by formulating cross-view feature alignment as an entropically regularized Unbalanced Optimal Transport (UOT) problem with marginal relaxation penalties:

\[\mathcal{T}^{(p,q)} = \arg\min_{T \ge 0} \langle T, C^{(p,q)} \rangle + \tau \text{KL}(T \mathbf{1} \,\|\, a) + \tau \text{KL}(T^\top \mathbf{1} \,\|\, b) + \varepsilon \sum_{i,j} T_{ij} (\log T_{ij} - 1)\]

where the ground cost matrix \(C_{ij}^{(p,q)} = \|x_i^{(p)} - x_j^{(q)}\|_2^2\) is evaluated between hypergraph-refined tokens, \(\tau\) controls marginal distribution divergence relaxation, \(\varepsilon\) represents the entropic regularization weight, and \(a, b\) are uniform reference distributions. Solved via 30 iterations of a generalized differentiable Sinkhorn algorithm, the optimal plan \(\mathcal{T}^{(p,q)}\) accommodates missing returns and spurious noise, transporting complementary features from view \(q\) to view \(p\) without requiring explicit spatial correspondence:

\[\widehat{X}^{(p \leftarrow q)} = \mathcal{T}^{(p,q)} X_{\text{hg}}^{(q)}, \qquad X_{\text{fused}}^{(p)} = X_{\text{hg}}^{(p)} + \sum_{q \in \mathcal{N}(p)} \beta_{p,q} \widehat{X}^{(p \leftarrow q)}\]

3. Adaptive Attention Gating and Cross-View Consistency Regularization: Noise Suppression and Geometric Anchoring

Even after UOT alignment, fused representations can retain background radar noise. HyperRadar passes the fused representations through an adaptive channel attention module prior to decoding. A global descriptor \(g^{(p)} \in \mathbb{R}^C\) is derived via spatial average pooling, and channel-wise modulation gates are generated by a lightweight network \(\psi(\cdot)\) to dynamically reweight features: \(\widetilde{X}^{(p)} = \psi(g^{(p)}) \odot X_{\text{fused}}^{(p)}\). In the final network configuration, 8 cascaded attention blocks are deployed to progressively enhance the signal-to-noise ratio of weak foreground targets.

To ensure geometric consistency across the RD and RA prediction branches along their shared physical range axis, the training objective combines cross-entropy and soft Dice losses with range-profile regularizers. By collapsing azimuth on RA and Doppler on RD (with a \(180^\circ\) coordinate rotation), the network derives 1D range confidence profiles \(Q^{(\text{RA})}\) and \(Q^{(\text{RD})}\), penalized via an \(L_2\) coherence loss \(\mathcal{L}_{\text{coh}}\) and a Huber alignment loss \(\mathcal{L}_{\text{mv}}\):

\[\mathcal{L}_{\text{coh}} = \| Q^{(\text{RD})} - Q^{(\text{RA})} \|_2^2, \qquad \mathcal{L}_{\text{mv}} = \text{Huber}\left(Q^{(\text{RD})}, Q^{(\text{RA})}\right)\]

This joint regularizer forces predictions across distinct radar projections to maintain rigorous physical alignment along the range dimension, suppressing spurious predictions in occluded views.

Key Experimental Results

Main Results

HyperRadar is evaluated on the multi-class, multi-view CARRADA benchmark and the high-resolution RADIal benchmark. On the CARRADA test split, HyperRadar sets new performance milestones across both RD and RA views, delivering remarkable gains on vulnerable road user classes characterized by sparse returns (pedestrians and cyclists).

View Method Bkg. IoU (%) Ped. IoU (%) Cycl. IoU (%) Car IoU (%) mIoU (%) mDice (%)
RD View U-Net (MICCAI'15) 99.7 51.1 33.4 37.7 55.4 68.0
RAMP-CNN (IEEE Sensors'20) 99.7 48.8 23.2 54.7 56.6 68.5
TMVA-Net (IEEE TIV'21) 99.7 52.6 29.0 53.4 58.7 70.9
PeakConv (IEEE TIV'23) - - - - 60.7 72.5
TransRadar (WACV'24) 99.7 56.7 30.2 61.7 62.1 73.4
HyperRadar (Ours) 99.7 59.2 32.4 63.8 63.8 75.2
RA View U-Net (MICCAI'15) 99.8 22.4 8.8 0.0 32.8 40.9
TMVA-Net (IEEE TIV'21) 99.8 26.0 8.6 30.7 41.3 51.0
PeakConv (IEEE TIV'23) - - - - 42.9 53.3
TransRadar (WACV'24) 99.9 29.9 6.5 35.3 42.9 52.6
HyperRadar (Ours) 99.9 31.8 8.4 37.5 44.4 54.3

On the single-view RADIal benchmark, incorporating HyperRadar as the backbone within the FFTRadNet framework achieves superior results across both object detection and semantic segmentation:

Backbone AP (%) ↑ AR (%) ↑ Range Error R (m) ↓ Angle Error A (°) ↓ mIoU (%) ↑
FFTRadNet 96.8 82.2 0.11 0.17 74.0
C-M DNN 96.9 83.5 - - 80.4
TransRadar 97.3 98.4 0.11 0.10 81.1
HyperRadar (Ours) 98.5 98.9 0.10 0.08 83.4

Ablation Study

The component ablation on CARRADA investigates the individual contributions of Hypergraph refinement (HG), Unbalanced Optimal Transport (UOT), and Adaptive Attention (Attn), alongside hyperparameter sensitivity analysis on marginal relaxation \(\tau\) and entropy weight \(\varepsilon\):

Config HG UOT Attn RD mIoU (%) RA mIoU (%) Avg mIoU (%) Note & Delta
Baseline - - ✓ 62.1 42.9 52.5 Attention-only baseline
w/ UOT only - ✓ - 61.7 42.2 52.0 Isolated UOT transport
w/ UOT + Attn - ✓ ✓ 62.4 43.1 52.8 Cross-view UOT alignment
w/ HG only ✓ - - 61.9 42.4 52.2 Isolated hypergraph module
w/ HG + Attn ✓ - ✓ 62.5 43.3 52.9 Intra-view HG refinement
w/ HG + UOT ✓ ✓ - 63.0 43.7 53.4 Without attention gating
Full Model ✓ ✓ ✓ 63.8 44.4 54.1 Full framework (+1.6% Avg)

In sensitivity evaluations, variations in \(\tau \in \{0.5, 0.8, 1.0\}\) and \(\varepsilon \in \{0.01, 0.05, 0.10\}\) yield stable average mIoUs within 53.6% to 54.1% without training divergence. In addition, increasing the hyperedge capacity beyond \(M=8\) leads to performance saturation, confirming robustness against hyperedge tuning.

Key Findings

  • Synergistic Orthogonal Gains: Hypergraph refinement alone (+0.4% Avg mIoU) captures intra-view reflection clusters, while UOT alignment alone (+0.3% Avg mIoU) addresses inter-view density mismatch. When combined, their synergy produces a notable +1.6% Avg mIoU gain, demonstrating the orthogonal benefits of higher-order intra-view modeling and mass-relaxed cross-view transport.
  • Substantial Recovery of Weak Returns: In the challenging RA projection, cyclist IoU improves from 6.5% (TransRadar) to 8.4% (a 29.2% relative gain), proving that the hypergraph structure successfully gathers dispersed, weak energy spikes that baseline convolutions discard as noise.
  • Reduced Computational Footprint with Higher Frame Rate: HyperRadar achieves higher throughput (10.55 FPS vs. 9.60 FPS for TransRadar) with fewer GFLOPs (1132.13 vs. 1138.57) and reduced peak VRAM consumption (3.18 GB vs. 3.25 GB), demonstrating that hypergraph relational modeling does not incur high computational overhead.

Highlights & Insights

  • Grounding Mathematical Transport in Radar Physics: Rather than applying generic neural modules, the authors ground their design in the physical scattering properties of automotive FMCW radar, using Unbalanced Optimal Transport to handle the natural mass expansion between RD and RA views.
  • Compact Soft Hyperedge Design: With only \(M=8\) trainable hyperedges and a soft incidence projection, the model circumvents the high memory footprint of dense relational graphs while achieving expressive higher-order spatial clustering.
  • Cross-Task Generalizability: The combination of hypergraph structural refinement and relaxed distribution matching readily transfers to other asymmetrical, sparse multimodal settings, such as asynchronous event camera fusion and infrared-radar cooperative sensing.

Limitations & Future Work

  • Information Loss in 2D Projections: While utilizing 2D RAD slices offers compelling inference efficiency, it discards cross-dimensional coupling present in the raw 3D tensor representation.
  • Constrained Cross-View Geometric Constraints: Current cross-view geometric supervision is restricted to the shared Range axis between RD and RA, lacking full inverse-projection constraints and multi-frame temporal bundle adjustment.
  • Expansion to Unified Multimodal Perception: Promising directions include investigating sparse hierarchical hypergraphs directly on 3D radar cubes, as well as extending UOT cross-view alignment to multi-modal perception pipelines combining radar, camera, and LiDAR for joint segmentation, detection, and tracking.
  • vs TMVA-Net (IEEE TIV 2021): TMVA-Net relies on standard 2D convolutions and pairwise multi-view fusion, remaining vulnerable to broken radar signatures and clutter. HyperRadar introduces higher-order hyperedge grouping and mass-relaxed UOT alignment, outperforming TMVA-Net by +5.1% RD and +3.1% RA mIoU on CARRADA.
  • vs TransRadar (WACV 2024): TransRadar employs directional self-attention to capture long-range dependencies, but pairwise attention degrades under severe cross-view density imbalances. HyperRadar leverages UOT to tolerate mass variation, beating TransRadar on CARRADA (+1.7% RD, +1.5% RA) and RADIal (+2.3% mIoU) while running at a faster frame rate (10.55 vs. 9.60 FPS).

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates radar reflection grouping as hypergraph message passing and addresses cross-view aperture imbalance via Unbalanced Optimal Transport.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarking across CARRADA and RADIal with per-class IoU/Dice, detailed ablation studies, sensitivity analyses, and runtime complexity profiling.
  • Writing Quality: ⭐⭐⭐⭐⭐ Deep technical analysis, clear physical motivation, rigorous mathematical formulation, and well-structured empirical validation.
  • Value: ⭐⭐⭐⭐⭐ Offers valuable insights into structured representation learning for sparse, noisy radar sensing and asymmetric multi-view perception.