Skip to content

Teaching Vision-Language-Action Models What to See and Where to Look

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/ShivaTeam/DriveTeach-VLA
Area: Robotics & Embodied AI
Keywords: Vision-Language-Action model, Autonomous Driving, Driving-aware Vision Distillation, 2D Trajectory-Guided Prompting, Reinforcement Learning Alignment

TL;DR

DriveTeach-VLA introduces Driving-aware Vision Distillation (DVD) and 2D Trajectory-Guided Prompts (2D-TGP) to decouple and ground spatial attention on critical traffic objects and feasible driving corridors, achieving 90.4 PDMS on NAVSIM and setting new state-of-the-art trajectory prediction accuracy on nuScenes.

Background & Motivation

Autonomous driving systems are undergoing a paradigm shift from traditional modular pipelines toward unified Vision-Language-Action (VLA) architectures powered by Multimodal Large Language Models (MLLMs). Conventional modular stacks decompose driving into perception, motion prediction, and planning modules in sequence; although structurally interpretable and controllable, they suffer from compounded cascading errors and limited cross-task reasoning capabilities. To overcome these bottlenecks, recent VLA models directly map camera images and ego-vehicle states to future trajectory waypoints or discretized action tokens, utilizing visual question answering (VQA) and chain-of-thought (CoT) supervision to inject general driving knowledge.

However, existing AD-VLA frameworks exhibit a fundamental structural disconnect between perceptual representations and spatial planning. Current pretraining and instruction-tuning corpora are predominantly text-centric, optimizing the network to answer semantic queries, describe traffic scenes, or classify high-level intents rather than perform action-grounded spatial reasoning. As a consequence, visual encoders learn to verbally name what is present rather than internalize the physical geometry and spatial dependencies required for reliable trajectory synthesis. Qualitative and quantitative attention analyses reveal that models trained solely on textual traffic supervision display diffuse, ungrounded cross-attention during trajectory decoding, failing to focus on critical obstacles, vulnerable road users, or drivable bounds.

To bridge this perception-planning divide, the learning process must explicitly teach the model what to see and where to look before demanding safe trajectories. Core idea: decouple spatial guidance from trajectory generation via a dual-model architecture comprising a TGP-Prompter and a TGP-Planner, employing Driving-aware Vision Distillation (DVD) on bounding-box-augmented images to anchor visual representations to critical traffic objects, projecting BEV trajectories into 2D image coordinates (2D-TGP) to delineate feasible drivable corridors, and aligning planning policy with human driving preferences under GRPO reinforcement learning.

Method

Overall Architecture

DriveTeach-VLA adopts a decoupled dual-model architecture comprising a TGP-Prompter and a TGP-Planner, both built upon the Qwen2.5-VL-3B foundation model. The system operates through a three-stage vision-guided pipeline: identifying what to see via DVD pretraining, determining where to look via 2D-TGP guided supervised fine-tuning, and optimizing how to act via TGP-conditioned GRPO policy optimization. At inference, given a front-view camera frame, ego-vehicle state, and navigation instruction, the TGP-Prompter first generates 2D trajectory coordinates on the image plane as spatial prompts. The TGP-Planner then consumes these 2D prompts alongside visual features to perform structured 4-step CoT reasoning, autoregressively decoding the optimal 4-second future trajectory in BEV coordinates.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input: Front-view Image C + Ego State S + Task Instruction L"] --> D1["Driving-aware Vision Distillation (DVD)<br/>Grounding DINO bounding box augmentation + Block-wise ViT self-distillation"]
    D1 --> D2["2D Trajectory-Guided Prompts (2D-TGP)<br/>Pinhole camera projection + Prompter outputs 2D coordinate guidance"]
    D2 --> D3["Decoupled Dual-Model & Two-Stage Policy Alignment<br/>Planner conditions on 2D-TGP for 4-step CoT + GRPO preference optimization"]
    D3 --> Out["Output: Optimal 4-second future BEV trajectory"]

Key Designs

1. Driving-aware Vision Distillation (DVD): Self-distillation for Critical Traffic Priors

Standard vision transformers in MLLMs are pretrained on general web-scale imagery and lack driving-domain sensitivity toward traffic participants, road boundaries, and topological structures. Traditional methods rely on text-based VQA to transfer domain knowledge, but linguistic loss formulations fail to tightly constrain spatial feature maps. DVD circumvents textual intermediaries by establishing a self-distillation objective directly within the visual latent space. Grounding DINO is employed as an open-vocabulary detector to identify 11 critical traffic categories (cars, trucks, buses, trailers, construction vehicles, pedestrians, motorcycles, bicycles, barriers, traffic elements, and traffic lights), overlaying bounding boxes onto raw image \(C\) to create augmented image \(C_{\text{bbox}}\).

A student ViT receives the raw image \(C\), while a teacher ViT receives the augmented image \(C_{\text{bbox}}\). Rather than patch-level matching—where sparse critical objects would be overwhelmed by vast background regions—DVD partitions feature maps into \(K\) non-overlapping rectangular blocks \(\mathcal{B}_k\) (e.g., a \(2 \times 4\) grid), computing block-averaged representations aligned via Smooth-L1 loss:

\[\bar{v}_k^t = \frac{1}{|\mathcal{B}_k|}\sum_{i \in \mathcal{B}_k} v_i^t, \quad \bar{v}_k^s = \frac{1}{|\mathcal{B}_k|}\sum_{i \in \mathcal{B}_k} v_i^s\]
\[\mathcal{L}_{\text{distill}} = \frac{1}{K}\sum_{k=1}^K \text{Smooth-L1}(\bar{v}_k^t, \bar{v}_k^s)\]

The student ViT parameters are optimized via backpropagation, while teacher ViT weights are updated via an Exponential Moving Average (EMA, momentum 0.996), internalizing salient traffic cues directly into the visual backbone.

2. 2D Trajectory-Guided Prompts (2D-TGP): Projected Corridors for Feasible Driving Regions

Directly predicting 3D or BEV coordinate offsets from 2D images poses severe depth ambiguity and geometric reasoning challenges for autoregressive language decoders. Conversely, MLLMs possess strong native 2D visual grounding and pixel-coordinate localization priors. To bridge this gap, 2D Trajectory-Guided Prompts (2D-TGP) project 3D/BEV expert trajectories \(\mathcal{T}_w = \{(x_t, y_t, \psi_t)\}_{t=1}^T\) onto the 2D image plane using camera intrinsic matrix \(K\), extrinsic matrix \([R|t]\), and pinhole projection function \(\pi(\cdot)\):

\[(x_t^I, y_t^I) = \pi(K, [R|t], (x_t, y_t))\]

Under the flat-ground assumption (\(z=0\)), this transformation is invertible, allowing exact recovery of BEV positions for metric verification. The projected coordinate sequence \(P_I = [(x_1^I, y_1^I), \dots, (x_T^I, y_T^I)]\) explicitly delineates the vehicle's future drivable path in pixel space. The TGP-Prompter's language decoder is trained via autoregressive cross-entropy \(\mathcal{L}_{\text{TGP}}\) to predict these normalized 2D image coordinates, serving as high-fidelity visual-spatial prompts.

3. Decoupled Dual-Model Architecture & Policy Alignment: From Prompt Conditioning to Preference Optimization

Forcing a single model to concurrently predict 2D guidance coordinates and 3D BEV trajectory waypoints leads to severe instruction confusion and gradient interference. DriveTeach-VLA decouples these roles: TGP-Planner is initialized with the pretrained weights of TGP-Prompter, inheriting traffic perceptual priors while dedicating its capacity to motion planning. Conditioned on predicted 2D prompt coordinates \(\hat{P}_I\), TGP-Planner executes a structured 4-step CoT reasoning sequence: ① critical object detection, ② natural language trajectory explanation, ③ discrete meta-behavior selection, and ④ 4-second future BEV trajectory waypoint decoding.

During supervised training, teacher forcing supplies ground-truth 2D-TGP \(P_I\) to stabilize spatial condition learning. Subsequently, Group Relative Policy Optimization (GRPO) aligns planned trajectories with human driving preferences, guided by non-reactive closed-loop PDM-Score (PDMS) rewards:

\[\text{PDMS} = \prod_{c \in \mathcal{C}} c \times \frac{\sum_{m \in \mathcal{M}} w_m \cdot m}{\sum_{m \in \mathcal{M}} w_m}\]

where multiplicative constraints \(\mathcal{C} = \{\text{NC}, \text{DAC}\}\) enforce no-at-fault collisions and drivable area compliance, and weighted metrics \(\mathcal{M} = \{\text{EP}, \text{TTC}, \text{C}\}\) measure ego progress, time to collision, and comfort with weights \(\{5, 5, 2\}\). At test time, Prompter-predicted \(\hat{P}_I\) incurs a negligible performance gap (only 0.4 PDMS), proving the robustness of the decoupled pipeline.

Loss & Training

DriveTeach-VLA is trained across three distinct stages: 1. DVD Pretraining (1 epoch): Jointly minimizes feature distillation loss and 2D-TGP coordinate cross-entropy, \(\mathcal{L}_{\text{DVD}} = \mathcal{L}_{\text{distill}} + \lambda_{\text{TGP}} \mathcal{L}_{\text{TGP}}\), with balance hyperparameter \(\lambda_{\text{TGP}} = 0.1\), AdamW optimizer (learning rate \(4 \times 10^{-5}\), weight decay 0.05), cosine decay schedule, and 0.10 warmup ratio. 2. CoT-SFT (6 epochs): Supervised fine-tuning on 8 H100 GPUs with batch size 16, utilizing step-by-step reasoning annotations generated by Qwen2.5-VL-72B. 3. GRPO Reinforcement Learning: 180 update steps with group size 8 across sampled trajectory rollouts, updating Planner policy parameters against PDMS rewards to reinforce driving compliance and safety margins.

Key Experimental Results

Main Results

On the closed-loop non-reactive autonomous driving benchmark NAVSIM (navtest split), DriveTeach-VLA demonstrates superior planning performance (PDMS) compared with state-of-the-art end-to-end and VLA baselines:

Method Pub. Base VLM NC (%) ↑ DAC (%) ↑ TTC (%) ↑ C (%) ↑ EP (%) ↑ PDMS ↑
Human Expert - - 100.0 100.0 100.0 99.9 87.5 94.8
UniAD CVPR'23 - 97.8 91.9 92.9 100.0 78.8 83.4
TransFuser TPAMI'22 - 97.7 92.8 92.8 100.0 79.2 84.0
Hydra-MDP CVPRW'24 - 98.3 96.0 94.6 100.0 78.7 86.5
DiffusionDrive CVPR'25 - 98.2 96.2 94.7 100.0 82.2 88.1
WoTE ICCV'25 - 98.5 96.8 94.4 99.9 81.9 88.3
ASSCG arXiv'26 - 98.2 98.3 94.8 100.0 87.5 91.4
ReCogDrive ICLR'26 InternVL3-8B 98.2 97.8 95.2 99.8 83.5 89.6
ImagiDrive-S ICRA'26 InternVL2.5-4B 98.1 96.2 94.5 100.0 80.5 86.9
AutoVLA NeurIPS'25 Qwen2.5-VL-3B 98.4 95.6 98.0 99.9 81.9 89.1
CuriousVLA CVPR'26 Qwen2.5-VL-3B 97.7 95.9 97.2 98.2 89.2 88.9
DriveTeach-VLA (Ours) ECCV'26 Qwen2.5-VL-3B 98.5 96.9 97.9 98.2 88.5 90.4

On the open-loop nuScenes trajectory prediction benchmark, DriveTeach-VLA sets new state-of-the-art records across both standard protocol metrics: - ST-P3 Evaluation Protocol: Achieves average L2 trajectory error of 0.30 m (outperforming AutoVLA's 0.48 m and Impromptu VLA's 0.33 m) and a collision rate of 0.12%. - UniAD Evaluation Protocol: Achieves L2 error of 0.60 m (substantially lower than UniAD's 1.03 m and AutoVLA's 0.86 m) with a collision rate of 0.31%.

Ablation Study

Incremental ablation on NAVSIM isolates the contribution of each design component and training stage († indicates BEV trajectory recovered from Prompter's 2D-TGP via inverse camera projection):

Config & Training Stage Model DAC (%) ↑ EP (%) ↑ TTC (%) ↑ PDMS ↑ Note
Qwen2.5-VL-3B Baseline Planner 93.2 85.8 97.3 84.8 Direct fine-tuning without driving-aware pretraining
+ VQA Pretraining + CoT-SFT Planner 94.0 87.2 96.7 86.4 Conventional text-centric pretraining setup
+ DVD Pretraining + CoT-SFT Prompter † 94.9 87.5 96.7 87.3 Inverse projection validates DVD perceptual gains
+ DVD + CoT + GRPO Prompter † 95.5 90.2 96.5 88.2 RL alignment applied directly to Prompter
+ DVD Pretraining + CoT-SFT Planner 94.4 86.3 97.8 87.1 Planner initialized with DVD but without 2D-TGP
+ DVD + 2D-TGP + CoT-SFT Planner 95.8 87.2 96.6 88.2 2D path prompt significantly boosts drivable area compliance
+ DVD + CoT + GRPO Planner 95.8 90.6 96.7 89.1 GRPO optimization without 2D-TGP conditioning
+ DVD + 2D-TGP + CoT + GRPO (Full) Planner 96.9 88.5 97.9 90.4 Full dual-model vision-grounded pipeline

Key Findings

  • DVD Outperforms Conventional VQA Pretraining: Substituting text-centric VQA with vision-level bounding-box distillation significantly enhances attention on traffic objects. The Attention Mass (AM) score increases from \(2.28 \times 10^{-2}\) (baseline) to \(3.19 \times 10^{-2}\), yielding a direct gain of 0.7–0.9 PDMS.
  • 2D-TGP Eliminates Drivable Area Violations: Providing 2D image coordinates as explicit prompts boosts DAC compliance from 94.4% to 95.8% (under SFT) and to 96.9% (under GRPO), confirming that image-space corridors effectively mitigate geometric disorientation in MLLM action decoding.
  • Decoupled Architecture Prevents Instruction Interference: Consolidating Prompter and Planner into a single multi-turn dialogue model drops PDMS to 87.0—lower than the decoupled 88.2 and even trailing the Prompter-only inverse projection baseline (87.3). This proves the necessity of isolating guidance generation from trajectory planning.
  • Token Efficiency Offsets Dual-Model Latency: While the dual-model pipeline requires two sequential forward passes (3.03s latency on H100), visual spatial grounding drastically reduces the necessity for verbose textual CoT tokens. On an NVIDIA L20 GPU, DriveTeach-VLA runs in 3.18s, outperforming AutoVLA's text waypoint mode (7.65s) and discrete action token mode (3.95s).

Highlights & Insights

  • Non-Textual Domain Distillation: By drawing detection bounding boxes directly onto input images and employing Siamese ViT self-distillation, the vision backbone internalizes spatial traffic priors without relying on noisy language intermediaries.
  • Ground-Plane Projective Invertibility: Exploiting flat-world geometry (\(z=0\)) transforms complex 3D BEV trajectory estimation into a grounded 2D pixel coordinate prompt, seamlessly unlocking MLLMs' innate visual referencing capabilities.
  • Teacher-Forced Robust Training: Employing ground-truth 2D-TGP during offline training ensures optimal gradient signals; during test-time rollout, the planner remains remarkably resilient to moderate Prompter inaccuracies (incurring merely a 0.4 PDMS oracle drop).

Limitations & Future Work

  • Reliance on Upstream Open-Vocabulary Detectors: DVD pretraining assumes accurate detections from Grounding DINO. Under extreme perturbations (e.g., 40% bounding box drop/jitter), distilled priors degrade and impair planning performance (dropping to 85.3 PDMS).
  • Planar Ground Assumption: 2D-TGP formulation relies on monocular pinhole projection and the flat-surface assumption (\(z=0\)), which faces challenges in multi-camera configurations, complex multi-level interchanges, and steep terrains.
  • Memory Footprint from Dual-Model Deployment: Maintaining two 3B models in GPU memory increases consumption to 17.2 GiB (compared to ~8.2 GiB for single models). Future work could explore weight-sharing architectures via adapter branches or lightweight prompt heads.
  • vs AutoVLA: AutoVLA relies heavily on long textual CoT reasoning and discrete action bins to compensate for ungrounded visual representations, leading to high token overhead and scattered attention; DriveTeach-VLA provides explicit spatial anchors via DVD and 2D-TGP, delivering superior planning accuracy (90.4 vs 89.1 PDMS) with lower token latency.
  • vs ReCogDrive: ReCogDrive employs a much larger 8B parameter vision backbone (InternVL3) coupled with an external diffusion motion planner; DriveTeach-VLA demonstrates that a compact 3B MLLM with vision-guided spatial prompts can outperform complex hybrid architectures.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Elegant decoupling of driving-aware visual distillation and 2D trajectory prompts.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across NAVSIM and nuScenes, paired with detailed attention quantification, noise sensitivity tests, and runtime benchmarking.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear conceptual motivation, crisp mathematical formulation, and tight visual-textual consistency.
  • Value: ⭐⭐⭐⭐⭐ Establishes a practical, generalizable blueprint for spatial grounding in autonomous driving and embodied VLA models.