DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Autonomous Driving
Keywords: Dual-Horizon Cooperation, Latent-Space Reasoning, Vehicle-Infrastructure Cooperation, Vision-Language Models, V2X Communication
TL;DR¶
Proposes DH-VLM, an asymmetric vehicle-infrastructure cooperative latent reasoning framework where an infrastructure large model aggregates multi-layer hidden states into global cognitive guidance, conditionally refining on-board ego planning via an infrastructure-driven latent evolution mechanism to cut communication costs by 57.3% while significantly enhancing long-horizon safety and planning accuracy.
Background & Motivation¶
End-to-end autonomous driving directly maps raw sensory observations to vehicle trajectory control commands, successfully mitigating error accumulation and stage-to-stage information bottlenecks inherent in classical modular pipelines. However, traditional end-to-end models rely heavily on low-level bird's-eye-view grid features or discrete query vectors, leaving them vulnerable in complex, interactive, and long-tail traffic scenarios that demand high-level semantic reasoning and causal inference. Incorporating vision-language models (VLMs) equips autonomous systems with strong zero-shot generalization and rich commonsense knowledge priors. Nevertheless, isolated single-vehicle systems are fundamentally constrained by on-board sensor range, field-of-view limits, and visual occlusions, while deploying high-capacity VLMs directly on vehicles incurs prohibitive compute and memory overhead that conflicts with strict real-time constraints.
Vehicle-to-Everything (V2X) cooperative autonomous driving provides an intuitive pathway to bypass single-vehicle sensory limitations by sharing views from external agents such as roadside infrastructure. Yet existing cooperation paradigms face a severe trilemma between bandwidth, compute, and reasoning robustness: early-fusion approaches that transmit raw images or point clouds saturate wireless channel capacity; intermediate-fusion frameworks based on perceptual queries or BEV features lack long-horizon semantic reasoning and suffer from perception-planning misalignment; late-stage natural language cooperation (e.g., LangCoop) requires vehicles to host massive LLMs to parse text and is extraordinarily brittle against upstream noise, where flawed text advice causes catastrophic planning failures.
Addressing this fundamental tension across on-board computational budgets, communication constraints, and reasoning robustness, this paper adopts a dual-system cognitive perspective to decouple global long-horizon scene comprehension from fast local trajectory execution. Core idea: construct an asymmetric dual-horizon cooperative latent reasoning framework, where a high-capacity roadside infrastructure aggregates deep hidden states to provide continuous semantic priors, and a lightweight on-board model refines local planning through an infrastructure-driven latent evolution mechanism, achieving robust, noise-resilient, and low-bandwidth cooperative driving while preserving ego decision autonomy.
Method¶
Overall Architecture¶
DH-VLM decouples vehicle-infrastructure driving into an asymmetric architecture comprising an infrastructure-side Global-Reasoning Horizon and an ego-vehicle Local-Planning Horizon. The roadside system deploys a high-capacity Qwen2.5-7B model, fusing multi-agent sensory queries in a shared 3D coordinate frame and employing a Hierarchical Latent Aggregation (HLA) module to extract compact cognitive guidance \(\mathbf{z}_t^{\text{inf}}\) from intermediate Transformer layers. The vehicle side hosts a lightweight Qwen2.5-0.5B model executing an adaptive two-stage forward reasoning pipeline. The vehicle first runs a base forward pass to generate initial latent representations; upon receiving infrastructure guidance over asynchronous V2X channels, the Infrastructure-Driven Latent Evolution (IDLE) module performs cross-domain feature alignment and cross-attention fusion before a second forward pass recursively refines trajectory predictions. If communication packets are lost or delayed, the vehicle gracefully falls back to its standalone pass without interruption.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Multimodal Cooperative Inputs<br/>Ego Local Vision + Infrastructure Global View + Ego State"] --> FVE["Fusion Vision Encoder<br/>Unified 3D Query Interaction & Alignment"]
FVE --> HLA["Hierarchical Latent Aggregation (HLA)<br/>Intermediate Layer Attention Pooling & Temporal Smoothing"]
HLA --> Comm{"Asynchronous Sparse V2X Channel<br/>5:1 Transmission Ratio / Packet Loss & Latency"}
Comm -->|Guidance Received| IDLE["Infrastructure-Driven Latent Evolution (IDLE)<br/>MLP Alignment & Two-Stage Latent Refinement"]
Comm -->|Interrupted/Timeout| Fallback["Single-Stage Ego Autonomous Fallback<br/>On-board Qwen2.5-0.5B Direct Planning"]
IDLE --> Plan["Cooperative Safe Trajectory<br/>Avoiding Occluded Risks & Interaction Conflicts"]
Fallback --> Plan
QA["Cooperation-Oriented Counterfactual QA<br/>Scene Topology + Ego Grounding + Trajectory Simulation"] -.->|Three-Stage Supervision| HLA
QA -.->|Safety Sensitivity| IDLE
Key Designs¶
1. Hierarchical Latent Aggregation (HLA): Multi-layer Representation Distillation and Temporal Smoothing
To overcome the dilemma between error-amplifying discrete text exchange and bandwidth-choking dense feature maps, the infrastructure must compress global reasoning into a compact yet semantically rich format. Transformer representations across depth reveal that final layers are heavily skewed toward discrete token vocabulary classification, shedding spatial geometry and rich continuous relations; conversely, intermediate hidden layers preserve structured spatio-temporal topology and causal interaction dynamics. HLA selects a subset of intermediate layers \(\mathcal{S} \subset \{1, \dots, L-1\}\) (experimentally optimal at 8 layers) and applies a learnable query vector \(\mathbf{q}\) for token-wise attention pooling within each layer:
The pooled representations are then combined across layers using learnable layer-wise weights \(w_l\): \(\tilde{\mathbf{z}}_t^{\text{inf}} = \sum_{l \in \mathcal{S}} \alpha_l \text{Agg}(\mathbf{h}_l^{(t)})\), where \(\alpha_l = \frac{\exp(w_l)}{\sum_k \exp(w_k)}\). To eliminate high-frequency jitter caused by sensory fluctuations and asynchronous processing, the system smooths representations across a sliding temporal window \(\mathcal{T}\) (window length \(k=5\)) via \(\mathbf{z}_t^{\text{inf}} = \mathcal{G}\left(\{\tilde{\mathbf{z}}_{t-\tau}^{\text{inf}}\}_{\tau \in \mathcal{T}}\right)\), yielding a temporally stable and compact semantic vector representing global scene comprehension.
2. Infrastructure-Driven Latent Evolution (IDLE): Heterogeneous Alignment and Recursive Latent Refinement
Infrastructure and vehicle models possess substantial discrepancies in parameter capacity (7B vs. 0.5B) and latent feature distributions. IDLE resolves this heterogeneity while safeguarding ego autonomy against misleading external signals. At timestamp \(t\), the ego vehicle performs an initial forward pass through its lightweight VLM over local sensory tokens and textual driving prompts, producing base semantic tokens \(\tilde{Q}_t^{\text{veh}}\). Upon receiving infrastructure guidance \(\mathbf{z}_t^{\text{inf}}\), a lightweight MLP adapter maps it into the vehicle's latent dimension: \(\bar{\mathbf{z}}_t^{\text{inf}} = \text{MLP}(\mathbf{z}_t^{\text{inf}})\). Taking the ego latent representation \(\mathbf{z}_t^{\text{veh}} = \text{MLP}(\tilde{Q}_t^{\text{veh}})\) as Query, and the aligned infrastructure features as Key and Value, cross-attention integrates the guidance:
Injecting external guidance as a continuous soft conditioning prior allows the vehicle to absorb vital non-line-of-sight hazard cues (e.g., pedestrians occluded by leading trucks) while filtering out infrastructure hallucinations. The vehicle then triggers a second recursive forward pass \(\hat{\mathbf{z}}_t^{\text{veh}} = \mathcal{F}_{\text{VLM}}^{\text{2nd}}(\tilde{\mathbf{z}}_t^{\text{veh}}, \mathcal{C}_t)\), seamlessly weaving the cooperative prior into internal reasoning dynamics to decode accurate future waypoints. Under complete packet loss or transmission timeouts, the model simply uses its initial pass output, preserving autonomous driving continuity.
3. Cooperation-Oriented Counterfactual QA Dataset: Ego-Centric Intent and Risk Probing
Conventional autonomous driving QA datasets concentrate on general, passive scene descriptions and fail to encourage infrastructure models to produce proactive, ego-customized safety recommendations. To bridge this gap, this work constructs a cooperation-oriented dataset containing over 80k question-answer pairs built upon the real-world DAIR-V2X benchmark. The annotations span two hierarchical tiers: Fundamental Scene Understanding supervises roadside models on global topological relations, road boundary constraints, and dynamic object trajectories; Ego-Personalized Comprehension projects ego vehicle coordinates onto infrastructure sensor perspectives and annotates mutual visual occlusions. Roadside models then engage in counterfactual simulation using Qwen2.5-VL-32B, evaluating hypothetical trajectories (e.g., aggressive overtaking or risky lane cut-ins) to assess what collisions or conflicts would emerge. Supervised with this counterfactual reasoning, intermediate hidden representations become deeply imbued with safety-critical awareness, generating potent latent guidance for downstream vehicle planning.
Loss & Training¶
The framework is optimized via a modular three-stage curriculum: - Stage 1 (Infrastructure Semantic Pre-alignment): Freezes the vision encoder backbone and fine-tunes roadside Qwen2.5-7B on multimodal visual inputs and the cooperative QA dataset with a learning rate of \(1 \times 10^{-5}\), establishing comprehensive global scene reasoning and ego-centric counterfactual evaluation capabilities; - Stage 2 (Ego Standalone Driving Baseline): Trains the on-board Qwen2.5-0.5B model using ego sensory streams with a learning rate of \(1 \times 10^{-4}\), guaranteeing reliable autonomous navigation and emergency failsafe capabilities under zero-communication conditions; - Stage 3 (Joint Latent Cooperation Fine-tuning): Freezes both language model backbones and trains the attention pooling parameters in HLA, the alignment MLP adapter, and the cross-attention interaction layers in IDLE. All optimization uses AdamW with a weight decay of 0.03. Communication frequency operates at an asynchronous 5:1 ratio, matching the disparate inference speeds of 7B and 0.5B models.
Key Experimental Results¶
Main Results¶
On the DAIR-V2X real-world benchmark, DH-VLM is compared against single-agent baselines (VAD, UniAD, SparseDrive, OpenDriveVLA) and state-of-the-art cooperative driving methods (V2VNet, CooperNaut, UniV2X, UniMM-V2X, V2X-VLM, LangCoop) in trajectory displacement error (L2 Error) and Collision Rate (CR).
| Paradigm | Method | 1s L2↓ | 3s L2↓ | 5s L2↓ | Avg. L2 (m)↓ | 1s CR↓ | 3s CR↓ | 5s CR↓ | Avg. CR (%)↓ |
|---|---|---|---|---|---|---|---|---|---|
| Single-Agent | VAD | 1.65 | 3.80 | 5.78 | 3.72 | 0.86 | 1.28 | 1.95 | 1.35 |
| Single-Agent | UniAD | 1.26 | 3.06 | 5.22 | 3.15 | 0.88 | 1.32 | 1.45 | 1.17 |
| Single-Agent | SparseDrive | 1.02 | 2.37 | 3.87 | 2.39 | 0.46 | 1.28 | 1.58 | 1.18 |
| Single-Agent | OpenDriveVLA | 0.64 | 2.69 | 5.57 | 2.89 | 0.15 | 0.59 | 0.44 | 0.53 |
| Cooperative | V2VNet | 1.96 | 3.41 | 4.73 | 3.28 | 0.74 | 1.03 | 1.58 | 1.09 |
| Cooperative | CooperNaut | 2.69 | 5.50 | 7.94 | 5.24 | 1.18 | 1.76 | 1.69 | 1.58 |
| Cooperative | UniV2X | 1.45 | 3.04 | 5.24 | 3.18 | 0.15 | 0.44 | 0.74 | 0.41 |
| Cooperative | UniMM-V2X | 0.78 | 2.05 | 5.00 | 2.60 | 0.05 | 0.15 | 1.18 | 0.42 |
| Cooperative | V2X-VLM | 0.49 | 2.89 | 5.34 | 2.87 | 0.05 | 0.15 | 0.68 | 0.26 |
| Cooperative | LangCoop | 0.30 | 1.96 | 4.58 | 2.19 | 0.05 | 0.44 | 1.02 | 0.51 |
| Cooperative | DH-VLM (Ours) | 0.26 | 1.76 | 3.72 | 1.87 | 0.05 | 0.11 | 0.44 | 0.19 |
For deployment feasibility and computational resource demands, on-board parameter count, runtime GPU memory footprint, and transmission overhead are benchmarked:
| Method | On-board #Params (M)↓ | GPU Memory (GB)↓ | Transmission Cost (BPS)↓ |
|---|---|---|---|
| V2VNet | 262.1 | 6.09 | \(8.19 \times 10^7\) |
| UniV2X | 264.2 | 6.11 | \(8.09 \times 10^5\) |
| UniMM-V2X | 267.8 | 6.14 | \(9.32 \times 10^5\) |
| V2X-VLM | 776.5 | 10.92 | \(1.24 \times 10^7\) |
| LangCoop | 7384.5 | 17.74 | \(\sim 4 \times 10^3\) |
| DH-VLM (Ours) | 640.6 | 4.54 | \(3.45 \times 10^5\) |
Ablation Study¶
To assess the resilience of guidance injection modalities (discrete text vs. continuous latent representations) against upstream infrastructure perception errors, performance was evaluated across three infrastructure guidance quality tiers (•: accurate \(<2\)m, ◦: medium \(2\sim 4\)m, △: noisy \(>4\)m):
| Guidance Modality | Infrastructure Quality | 1s L2↓ | 3s L2↓ | 5s L2↓ | Avg. L2 (m)↓ | 1s CR↓ | 3s CR↓ | 5s CR↓ | Avg. CR (%)↓ |
|---|---|---|---|---|---|---|---|---|---|
| Standalone (No V2X) | - | 0.64 | 2.69 | 5.57 | 2.89 | 0.15 | 0.59 | 0.44 | 0.53 |
| Discrete Text Guidance | • (High) | 0.19 | 1.59 | 3.98 | 1.83 | 0.05 | 0.05 | 0.45 | 0.17 |
| Discrete Text Guidance | ◦ (Medium) | 0.44 | 2.31 | 4.98 | 2.48 | 0.15 | 0.34 | 1.27 | 0.54 |
| Discrete Text Guidance | △ (Noisy) | 2.26 | 5.63 | 9.15 | 5.66 | 0.43 | 1.84 | 4.99 | 2.38 |
| Latent Fusion (Ours) | • (High) | 0.26 | 1.76 | 3.72 | 1.87 | 0.05 | 0.11 | 0.44 | 0.19 |
| Latent Fusion (Ours) | ◦ (Medium) | 0.24 | 1.82 | 3.99 | 1.98 | 0.05 | 0.24 | 0.58 | 0.28 |
| Latent Fusion (Ours) | △ (Noisy) | 0.26 | 2.07 | 4.50 | 2.19 | 0.15 | 0.57 | 1.02 | 0.56 |
Additional architectural ablations on communication frequency and module design demonstrate: - Layer Aggregation Depth: Under a 5:1 ratio, aggregating 8 intermediate layers attains 1.87m average L2 and 0.19% CR. Using solely the final layer (Fused Layers = 1) causes average L2 error to degrade to 2.06m and collision rate to jump to 0.32%, confirming that top-layer representations lose rich spatial semantics; - Ablation of HLA and IDLE: Replacing HLA with a plain MLP increases average L2 error by 23.5% (to 2.31m) and doubles collision rate to 0.39%; replacing IDLE with basic MLP concatenation increases L2 error by 32.6% (to 2.48m) and collision rate by 78.9% (to 0.34%); omitting both leads to 2.59m L2 error and 0.47% CR.
Key Findings¶
- Latent Cooperation Couples Precision with Extreme Noise Resilience: When roadside perception experiences severe tracking failures or misclassifications (Tier △), text-based methods blindly follow distorted textual recommendations, causing average L2 error to explode to 5.66m (+95.8% worse than standalone baseline) and collision rate to hit 2.38% (+349.1% deterioration). In sharp contrast, DH-VLM treats latent features as soft attention conditioning; the vehicle selectively filters out noisy cues, restricting average L2 error to 2.19m and collision rate to 0.56%, preserving baseline driving safety.
- Intermediate Layers Form the Semantic Sweet Spot: Transmitting only the final hidden layer induces substantial performance drops because representations near the output head collapse into discrete linguistic token distributions; the intermediate 8 layers optimally retain spatial geometry and interactive causal priors.
- Drastic Reduction in Memory and Bandwidth Footprints: Relative to V2X-VLM, transmission bandwidth drops from \(1.24 \times 10^7\) BPS to \(3.45 \times 10^5\) BPS (a 97.2% reduction) and GPU memory is reduced from 10.92 GB to 4.54 GB (a 58.4% saving); compared to query-based UniV2X, communication is reduced by 57.3%, making on-board deployment highly practical on resource-constrained vehicle hardware.
Highlights & Insights¶
- Asymmetric Dual-Horizon Cognitive Hierarchy: Transcends conventional symmetrical cooperation assumptions by assigning long-horizon causal reasoning to an infrastructure-based 7B model and fast reactive planning to an on-board 0.5B model, matching computational complexity to operational tempo.
- Continuous Latent Conditioning Over Discrete Text Bottlenecks: Avoids the lossy discretization of converting continuous scene dynamics into natural language tokens and re-parsing them back into coordinates, maintaining end-to-end differentiability and robust soft conditioning.
- Two-Stage Recursive Inference with Graceful Fallback: By structuring vehicle inference into an independent initial pass followed by conditional refinement, the system naturally absorbs severe network latency (>500ms) and 100% packet loss without safety compromises.
Limitations & Future Work¶
- Open-Loop Evaluation Constraints: Limited by the benchmark collection protocols of real-world datasets like DAIR-V2X, planning trajectories are evaluated in an offline open-loop replay setting, lacking interactive closed-loop dynamic traffic simulation.
- Absence of End-to-End Vision-Language-Action (VLA) Control: The framework currently predicts discrete future waypoint sequences rather than generating direct low-level throttle, brake, and steering torque commands, leaving continuous VLA control for future exploration.
Related Work & Insights¶
- vs. UniV2X / UniMM-V2X: These frameworks exchange dense BEV features or sparse perception queries, operating purely at geometric and perceptual levels without high-level causal common sense; DH-VLM introduces foundation model reasoning while maintaining lower bandwidth and achieving superior collision avoidance.
- vs. LangCoop: LangCoop pioneers vehicle-to-vehicle natural language cooperation but demands multi-billion-parameter LLMs on-board and suffers catastrophic failure when discrete text contains upstream hallucinations; DH-VLM uses continuous latent alignment to ensure noise resilience and lightweight execution.
- vs. V2X-VLM: V2X-VLM shares raw sensor observations or early vision features, driving transmission requirements beyond \(10^7\) BPS; DH-VLM employs a 5:1 sparse transmission ratio of compact latent vectors, making it realistic for real-world cellular V2X deployment.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First systematic exploration and validation of latent-level feature fusion in VLM-powered vehicle-infrastructure cooperative autonomous driving.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations covering trajectory accuracy, memory and bandwidth measurements, adversarial guidance noise, latency/packet loss degradation, and counterfactual visualizations.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulations, clear problem framing, and tight logical progression from motivation to empirical validation.
- Value: ⭐⭐⭐⭐⭐ Offers a practical and elegant paradigm resolving the fundamental triangle contradiction among compute, bandwidth, and safety in cooperative autonomous systems.