Skip to content

HERO: Heterogeneous Evidential Robust Object-Level Collaborative Perception

Conference: ECCV 2026
Paper: ECCV Official
Area: Autonomous Driving
Keywords: Collaborative Perception, Object-Level Fusion, Evidential Deep Learning, Heterogeneous Perception, Pose & Latency Robustness

TL;DR

Addressing the heavy communication overhead of feature-level fusion and the vulnerability of standard late fusion under pose bias and communication latency in heterogeneous multi-agent systems, HERO introduces an evidential proposal protocol that decouples fusion into Beta-sum semantic evidence accumulation and max-evidence geometry selection, matching feature-level accuracy with only ~2 KB/frame bandwidth while demonstrating superior robustness under perturbations.

Background & Motivation

Collaborative perception in connected and automated driving enables multi-agent information aggregation across vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) links, dramatically extending visual range and eliminating occlusions. However, real-world deployment is inherently open and heterogeneous: collaborating agents deploy distinct sensing modalities (such as varying LiDAR configurations and multi-view cameras) alongside divergent detector architectures and training distributions. Mainstream feature-level fusion pipelines achieve strong 3D detection accuracy, but they fundamentally rely on compatible feature spaces and end-to-end multi-agent collaborative training. Moreover, exchanging dense BEV feature maps requires megabytes per frame (MB/frame), which easily overburdens realistic wireless communication channels.

In contrast, object-level late fusion directly shares sparse 3D bounding box proposals, naturally decoupling the fusion stage from local detector backbones and slashing bandwidth to kilobytes per frame (KB/frame)—making it ideal for open, vendor-agnostic deployments. Nonetheless, conventional late fusion protocols communicate only compact "3D box + scalar confidence" messages, which fail to provide adequate information under realistic system perturbations. Different agents exhibit disparate localization noise scales and unknown systematic pose biases, causing fixed IoU thresholds to trigger false merges or dropped associations; raw confidence scores cannot be compared across disparate model architectures; and transmission latency induces spatial displacement in stale proposals, meaning that naive box coordinate averaging corrupts well-localized ego estimations.

The foundational insight of this work is that the error statistics governing "whether an object exists (semantics)" and "where an object is precisely located (geometry)" are fundamentally asymmetric: multi-agent observations corroborate object existence additively, whereas geometric coordinates are prone to correlated source-level biases where coordinate averaging propagates systematic offsets. Core idea: equip each sparse proposal with evidential statistics (Beta distribution for semantics and Normal-Inverse-Gamma NIG distribution for geometry) and decouple fusion into summing semantic evidence across matched proposals while selecting geometry from the single most reliable source, unified with uncertainty-gated association, freshness decay, and training-free ego pose refinement.

Method

Overall Architecture

HERO is designed for heterogeneous, bandwidth-constrained V2X perception subject to real-world localization bias and asynchronous latency. The system comprises two main components: agent-side evidential detection heads and an ego-side multi-stage robust fusion engine. Locally, each heterogeneous agent independently executes its detector to predict 3D bounding boxes parameterized with Beta semantic evidence and NIG geometric uncertainty, selecting Top-K proposals via an evidence-aware joint score for transmission. Upon receiving proposals, the ego vehicle executes coordinate alignment and freshness evaluation, runs a two-pass detection-level pose refinement (DPR) anchored on ego proposals to cancel translation bias, clusters multi-source proposals using Mahalanobis distance with motion-inflated gating, and finally executes decoupled fusion: aggregating semantics via Beta-sum and selecting geometry from the single most reliable candidate.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Sparse Evidential Proposals<br/>(3D Box + Beta Semantics + NIG Geometry)"] --> B["Dual Gating & Mahalanobis Association<br/>Spatial Distance + Uncertainty-Aware Gating"]
    B --> C["Detection-Level Pose Refinement DPR<br/>Ego-Anchored Residual-Weighted Offset Correction"]
    C --> D["Freshness Decay & Motion Compensation<br/>Asymmetric Semantic/Geometric Decay + Variance Inflation"]
    D --> E["Decoupled Fusion: Beta-sum Semantics & Max-evidence Geometry<br/>Additive Existence Corroboration + Reliable Box Selection"]
    E --> F["Output: Final High-Accuracy 3D Detections"]

Key Designs

1. Evidential Detection Head and Proposal Sparsification: Explicit Dual-Uncertainty Parameterization

Standard detection heads output deterministic bounding box coordinates and scalar softmax probabilities, offering no mechanism to judge predictive reliability across heterogeneous networks. HERO leverages Evidential Deep Learning (EDL) to construct conjugate prior parameters for each output branch. For semantic objectness, the head outputs binary Beta distribution parameters \((\alpha^+, \alpha^-)\) with \(\alpha^\pm \ge 1\), yielding total evidence strength \(S = \alpha^+ + \alpha^-\), mean existence probability \(p = \alpha^+ / S\), and semantic uncertainty \(u = 2 / S\). For 3D bounding box regression, each dimension outputs Normal-Inverse-Gamma (NIG) parameters \((\gamma, \nu, \alpha_r, \beta_r)\), where pseudo-count \(\nu\) represents the effective observation count. Averaging across the 7 box dimensions produces a proxy for epistemic uncertainty: $\(u_{\text{epi}} = \frac{1}{\bar{\nu} + 1}\)$ Simultaneously, the predictive variances \(\sigma_{i,x}^2, \sigma_{i,y}^2\) from NIG parameterize BEV planar aleatoric uncertainty. To strictly satisfy low bandwidth limits (~2 KB/frame), local agents rank candidate proposals using the joint evidence-aware score \(s = (1 - u)p = \frac{(S-2)\alpha^+}{S^2}\), preserving only Top-K candidates after local NMS to guarantee high classification confidence coupled with strong supporting evidence.

2. Dual Gating and Mahalanobis-Guided Association: Adapting to Heterogeneous Noise Scales

Traditional late fusion treats cooperation as a deduplication problem by applying global NMS across all received boxes, discarding corroborating peer observations and suffering severe failures under localization noise. HERO transforms redundancy into an opportunity for evidence accumulation. Received proposals are transformed into the ego frame and evaluated via dual spatial and Mahalanobis gating: $\(d_{\text{mah}}^2(i, j) = \sum_{k \in \{x,y\}} \frac{(x_{i,k} - x_{j,k})^2}{\sigma_{i,k}^2 + \sigma_{j,k}^2 + \epsilon} \le \tau_{\text{mah}}, \quad \text{and } \|x_i - x_j\| \le \tau_{\text{abs}}\)$ where \(\epsilon\) denotes a numerical stabilizer. Mahalanobis gating naturally adapts to stochastic localization variance: proposals with higher regression uncertainty receive larger spatial tolerance, whereas high-precision predictions enforce tight boundaries to prevent false merges across adjacent targets. Clusters are formed greedily in descending order of evidence strength \(S\), strictly enforcing at most one proposal per agent per cluster to guarantee clean multi-source correspondence.

3. Freshness Decay and Motion-Variance Inflation: Asymmetric Spatiotemporal Drift Modeling

Transmission delays \(\Delta t\) introduce spatial discrepancy between stale proposals and current object states. Recognizing that semantic object category is relatively stable over short intervals while spatial geometry drifts rapidly, HERO models this temporal asymmetry via dual hyperbolic-style decay factors: $\(f_{\text{geo}}(\Delta t) = \frac{1}{1 + \left(\frac{v_{\text{max}} \cdot \Delta t}{d_{\text{ref}}}\right)^2}, \quad f_{\text{sem}}(\Delta t) = \sqrt{f_{\text{geo}}(\Delta t)}\)$ where \(v_{\text{max}}\) is the expected maximum velocity and \(d_{\text{ref}}\) is a reference object scale (e.g., standard vehicle width). Because \(f_{\text{sem}}\) decays noticeably slower than \(f_{\text{geo}}\), a 500 ms stale proposal reliably supports the existence of a vehicle without corrupting ego localization. Furthermore, across time intervals \(\Delta t_{ij}\), a saturated motion variance term \(\sigma_{\text{gate}}^2 = \min((v_{\text{max}}\Delta t_{ij})^2, \sigma_{\text{cap}}^2)\) is injected into the denominator of the Mahalanobis gate, preventing missed associations caused by target displacement while capping permissiveness to avoid over-association.

4. Detection-Level Pose Refinement (DPR): Training-Free Ego-Anchored Bias Correction

Inter-agent calibration drift and GPS inaccuracy introduce systemic planar translation bias across proposals transmitted by a given collaborator. HERO introduces Detection-Level Pose Refinement (DPR), which operates entirely during inference without ground truth annotations or joint network training. Using reliable ego proposals as geometric anchors, the ego vehicle computes the BEV residual \(r_{a,c} = x_c^{(e)} - x_c^{(a)}\) for each mutually matched cluster \(c\) involving collaborator \(a\). Weighting each match by reliability \(w_{a,c}\) derived from evidence strength, regression uncertainty, and message age yields the global translation correction: $\(\Delta t_a = \frac{\sum_c w_{a,c} r_{a,c}}{\sum_c w_{a,c}}\)$ Applying this estimated rigid translation to all proposals from collaborator \(a\) before a second, refined association pass rectifies coordinate displacement at negligible computational cost.

5. Decoupled Fusion Rule: Beta-Sum Semantic Accumulation and Max-Evidence Geometric Selection

This constitutes the foundational rule of HERO. Within each cluster \(C\), conventional pipelines frequently apply covariance-weighted box averaging. However, when collaborator proposals contain correlated directional bias, geometric averaging pulls the accurate ego prediction toward the erroneous direction, degrading high-IoU metrics like AP70. HERO resolves this via decoupled fusion: - Semantic Evidence Accumulation (Beta-sum): Leveraging conjugate prior addition, pseudo-counts from all cluster candidates are aggregated weighted by regression reliability \(r_i = \text{clip}(1 - u_{\text{epi},i}, r_{\text{min}}, 1)\) and semantic freshness \(f_{\text{sem}}(\Delta t_i)\): $\(\alpha^{*+} = 1 + \sum_{i \in C} r_i f_{\text{sem}}(\Delta t_i) (\alpha_i^+ - 1), \quad \alpha^{*-} = 1 + \sum_{i \in C} r_i f_{\text{sem}}(\Delta t_i) (\alpha_i^- - 1)\)$ yielding fused objectness \(p^* = \alpha^{*+} / S^*\) and fused confidence \(c^* = p^*(1 - u^*)\) with \(u^* = 2 / S^*\). - Geometric Bounding Box Selection (Max-evidence): Coordinate averaging is explicitly rejected in favor of selecting the single most reliable, fresh, and certain candidate box: $\(i^* = \arg\max_{i \in C} \left( S_i \cdot r_i \cdot f_{\text{geo}}(\Delta t_i) \right)\)$ By selecting geometry from the single most trustworthy source, the system is fundamentally immune to correlated pose bias contamination.

Loss & Training

Each agent network is trained independently in two stages without cross-agent communication or cooperative loss: 1. Warm-up Stage: Standard Focal Loss and Smooth-L1 Loss are used to establish stable feature representations and basic 3D detection capabilities. 2. Evidential Fine-Tuning Stage: Semantic heads are supervised with Beta-EDL loss (negative log marginal likelihood combined with a Kullback-Leibler divergence regularizer penalizing false evidence), while bounding box heads are trained with negative log-likelihood of the NIG distribution paired with evidence variance regularization.

Key Experimental Results

Main Results

Evaluation is conducted on the simulated OPV2V-H benchmark (combining PointPillars LiDAR [LP], SECOND LiDAR [LS], EfficientNet Camera [CE], and ResNet Camera [CR]) as well as the real-world DAIR-V2X vehicle-infrastructure dataset.

Table 1: Quantitative 3D detection performance on OPV2V-H across agent scaling

Method Fusion Paradigm Comm Volume (log2 B/f) LP + CE (AP50 / AP70) LP + CE + LS (AP50 / AP70) 4-Agent Full Set (AP50 / AP70)
F-Cooper Feature-level ~23.5 0.842 / 0.686 0.843 / 0.686 0.843 / 0.686
DiscoNet Feature-level ~23.5 0.865 / 0.716 0.870 / 0.720 0.869 / 0.720
AttFusion Feature-level ~23.5 0.864 / 0.696 0.867 / 0.697 0.868 / 0.694
CoBEVT Feature-level ~20.5 0.877 / 0.734 0.897 / 0.750 0.900 / 0.754
HEAL Feature-level (Hetero) ~21.0 0.875 / 0.782 0.922 / 0.853 0.923 / 0.855
STAMP Feature-level (Hetero) ~17.0 0.829 / 0.709 0.837 / 0.725 0.831 / 0.718
CodeFilling Codebook Compression ~22.0 0.882 / 0.751 0.916 / 0.786 0.916 / 0.787
GenComm Generative Comm ~14.0 0.892 / 0.733 0.924 / 0.774 0.925 / 0.775
HERO (Ours) Object-level ~11.0 (~2 KB) 0.881 / 0.788 0.922 / 0.840 0.922 / 0.841

Table 2: Performance of two-agent heterogeneous collaboration on DAIR-V2X (V2I) and OPV2V-H (V2V) in AP50

Method DAIR-V2X: LP + LS DAIR-V2X: LP + CE DAIR-V2X: LP + CR OPV2V-H: LP + LS OPV2V-H: LP + CE OPV2V-H: LP + CR
F-Cooper 0.545 0.627 0.611 0.868 0.844 0.828
DiscoNet 0.634 0.576 0.621 0.895 0.861 0.867
AttFusion 0.661 0.649 0.623 0.891 0.863 0.839
CoBEVT 0.692 0.638 0.660 0.934 0.860 0.865
HEAL 0.770 0.659 0.682 0.942 0.875 0.862
CodeFilling 0.706 0.671 0.662 0.926 0.882 0.877
GenComm 0.642 0.641 0.641 0.936 0.892 0.891
HERO (Ours) 0.823 0.744 0.753 0.933 0.881 0.880

Ablation Study

Table 3: Ablation of DPR and delay compensation on OPV2V-H under pose noise \(\sigma=0.6\) m/deg and latency 200 ms

Configuration DPR Correction Delay Compensation AP50 AP70 Key Takeaway
Baseline ✗ ✗ 0.737 0.661 Severely corrupted by spatial misalignment and stale proposals
w/ DPR ✓ ✗ 0.816 0.710 Translation bias corrected; AP50 surges by +7.9%
w/ Delay Compensation ✗ ✓ 0.825 0.736 Suppresses stale geometry; AP70 jumps significantly (+7.5%)
Full HERO Model ✓ ✓ 0.831 0.737 Combined synergy delivers peak robustness under severe noise

Table 4: Ablation of decoupled geometric and semantic fusion rules (AP50 / AP70)

Geometric Rule Clean Setting Noise 0.6 + Delay 200ms Semantic Rule Clean Setting Noise 0.6 + Delay 200ms
max-evidence (Ours) 0.902 / 0.810 0.831 / 0.737 beta-sum (+rel, Ours) 0.902 / 0.810 0.831 / 0.737
mean-box (Average) 0.890 / 0.783 0.824 / 0.661 beta-sum (w/o rel) 0.887 / 0.804 0.832 / 0.732
PoE (Product of Experts) 0.891 / 0.784 0.828 / 0.737 score-mean (+rel) 0.854 / 0.755 0.768 / 0.673
max-score (Top Confidence) 0.906 / 0.826 0.626 / 0.496 score-mean (w/o rel) 0.846 / 0.752 0.746 / 0.648

Key Findings

  1. Decoupled fusion is essential: Table 4 demonstrates that coordinate averaging (mean-box) causes AP70 to plummet to 0.661 under noise and latency (-7.6% drop). Similarly, reliance on uncalibrated confidence (max-score) collapses under perturbation (AP70 falls to 0.496). Decoupling into beta-sum existence aggregation and max-evidence single-source box selection provides unprecedented stability.
  2. Substantial gains in real-world V2I (DAIR-V2X): In real roadside-to-vehicle scenarios, substantial capability and perspective disparities lead to frequent source-selection conflicts (up to 72.6% in LP + CE). HERO filters out uncalibrated predictions via NIG reliability weights, outperforming all feature-level SOTA baselines (e.g., +8.5% AP50 over HEAL in LP + CE).
  3. Over 1000x communication reduction: Exchanging compact proposals requires only ~2 KB/frame (\(\log_2 \text{Bytes} \approx 11\)), compared to dense BEV feature exchange (\(\log_2 \text{Bytes} \approx 21-23\), or 2–8 MB/frame), saving over three orders of magnitude in transmission volume.

Highlights & Insights

  • "Aggregate Semantics, Select Geometry" Paradigm: Provides a profound physical and statistical insight: discrete semantic existence is corroborative and suitable for additive Bayesian updating, whereas continuous geometric coordinates carry correlated sensor and localization biases where spatial averaging induces negative transfer.
  • Training-Free Modular Interoperability: Eliminates the impractical requirement of joint cross-agent feature alignment training, allowing vehicles and infrastructure units from different manufacturers to interoperate seamlessly via standardized evidential proposal interfaces.
  • Unified Spatiotemporal Uncertainty Modeling: Gracefully handles transmission latency and spatial calibration error within a unified probabilistic framework by coupling hyperbolic freshness decay with motion-inflated Mahalanobis distance gating.

Limitations & Future Work

  • Representation Bottleneck and False Positives: Transmitting sparse 3D proposals strips away dense spatial context, leaving the ego vehicle unable to perform fine-grained background contextual verification. Under severe noise, this increases long-tail low-confidence false positives.
  • Ego-Anchor Dependence in DPR: DPR assumes the presence of overlapping mutual detections between the ego vehicle and collaborator. In sparse traffic or complete ego occlusion, translation bias estimation must rely on temporal filtering or be bypassed.
  • Future Directions: Integrating lightweight instance-level discriminative embeddings or imposing multi-frame temporal tracking consistency could suppress spurious false positives while maintaining kilobyte-level bandwidth.
  • vs HEAL (arXiv 2024): HEAL introduces an extensible feature-level cross-modal alignment pyramid, which requires pairwise collaborative training and high bandwidth. HERO operates at the object level, requires no collaborative training, and outperforms HEAL on DAIR-V2X.
  • vs OptiMatch (IV 2023) / VIPS (MobiCom 2022): While these methods also tackle perturbations at the object level, they rely on heuristic greedy matching or uncalibrated confidence scores. HERO incorporates Beta/NIG evidential statistics, outperforming OptiMatch by +18.4% AP50 under severe pose noise (\(\sigma=0.4\)) without complex multi-frame tracking graphs.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ (Pioneers the decoupled evidential fusion paradigm for heterogeneous collaborative perception)
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Thorough validation across OPV2V-H and DAIR-V2X with rigorous noise/delay stress tests and in-depth ablations)
  • Writing Quality: ⭐⭐⭐⭐⭐ (Crisply structured, mathematically elegant, and physically grounded)
  • Value: ⭐⭐⭐⭐⭐ (Directly addresses the trifecta of heterogeneity, bandwidth limitations, and calibration drift in real V2X systems)