DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution¶
Conference: NeurIPS2026
arXiv: 2609.31814
Code: https://github.com/PerfectXu88/DriveHierarchy
Area: Autonomous Driving
Keywords: vision-language models, hierarchical capability diagnosis, open-loop evaluation, closed-loop simulation, supervised fine-tuning
TL;DR¶
DriveHierarchy organizes vision-language driving capabilities into perceptual grounding, contextual memory, mental reasoning, and closed-loop execution, diagnoses 15 models through unified open-loop tasks and 100 interactive simulation scenarios, and shows on Qwen3-VL-8B that repairing selected open-loop weaknesses can raise the closed-loop composite score from 0.902 to 8.021, without establishing real-world driving safety.
Background & Motivation¶
A vision-language model (VLM) may recognize vehicles, understand navigation instructions, describe traffic situations, and even produce driving actions, but these accomplishments are not interchangeable. Answering driving questions correctly does not establish timely steering during continuous vehicle motion. Conversely, a collision during a closed-loop route does not identify whether the failure arose from distance estimation, cross-view deduplication, future-behavior prediction, or the model-to-control interface. Existing open-loop datasets and closed-loop driving benchmarks serve useful purposes but rarely place both within an interpretable capability space.
The gap addressed here is an evaluation structure, not an additional planning module. Inspired by hierarchical accounts of cognition, the authors separately measure current-scene perception, information integration across views and time, future-oriented inference, and execution during environmental interaction. Public driving datasets already provide object identities, geometry, temporal information, and question-answer annotations, while CARLA and SUMO enable controlled interactions. Together, these resources make it possible to construct a research instrument covering both local capability profiles and overall driving outcomes.
This hierarchy is initially a task-organization hypothesis, not an established neural computation chain. Cross-model correlations examine whether the capabilities are related without being redundant, and targeted supervised fine-tuning tests whether weaknesses offer useful intervention points. Core Idea: connect open-loop understanding with closed-loop execution through a four-rank capability profile, retaining diagnostic information for individual abilities and testing its practical utility through controlled fine-tuning.
Method¶
Overall Architecture¶
DriveHierarchy is a benchmark and diagnostic protocol, not a new four-stage neural network. R1โR3 are separately executed open-loop tasks: the model receives a single image, synchronized multi-view images, or a short clip and answers a question. R4 connects the same model to interactive simulation, where online observations drive action outputs. The four ranks do not require inference to explicitly call R1, then R2 and R3, before generating R4.
The design has four elements: define capability ranks through different inputs and supervision; preserve comparable profiles through task-specific scoring and frozen weights; measure closed-loop execution through a shared control interface and a multiplicative metric; and select supervised fine-tuning targets from the resulting profiles. These are evaluation and analysis relationships rather than neural-module data flow, so the rank outline is not depicted as a model pipeline.
Two dataset sizes must be distinguished. The integrated corpus contains 84,279 frames and 76,798 question-answer pairs. All reported R1โR3 results use the refined version containing 31,945 frames and 14,000 question-answer pairs. R4 uses 100 separately constructed simulation scenarios; the full corpus size should not be presented as the actual test size underlying the main results.
Key Designs¶
1. Hierarchical capability tasks: distinguish driving problems through inspectable inputs and labels
R1, Perceptual Grounding, concerns reliable evidence in the current scene and contains 8 tasks: existence judgment, counting, state identification, nearest-object distance estimation, referred-object distance estimation, distance-bucket counting, language-guided localization, and situation description. The first three derive targets from object-category and state annotations. Distance tasks use egoโobject geometry and calibrated coordinates; localization produces a bounding box. Description tasks require an answer grounded in the current traffic situation rather than a plausible driving narrative based on general knowledge.
R2, Contextual Memory, contains 4 tasks. Its key requirement is unavoidable integration, not merely more images. Multi-view memory presents six synchronized camera views and asks for the deduplicated number of surrounding targets, retaining samples that genuinely require cross-view integration. Temporal counting aggregates unique objects across a fixed short clip; temporal state recognition tracks whether the same object changes state; spatial relation questions ask about relative front, rear, left, and right positions. Instance deduplication, synchronized views, and continuity checks prevent simply counting each image and adding the results from substituting for integration.
R3, Mental Reasoning, contains 2 tasks: predict a target's future action from the current context, and recover the temporal order of candidate observations. The former derives multiple-choice targets from future-aware annotations; the latter checks sequence consistency. โSequence planningโ here is image ordering, not unrestricted driving-trajectory generation. A high R3 score therefore does not establish comprehensive traffic-game reasoning or long-horizon planning.
Sources include NuScenes, nuPlan, NAVSIM, DriveLM, DriveBench, LingoQA, and DRAMA. The refined set applies visibility-projection checks, category-vocabulary normalization, scene quotas, invalid-box removal, instance-ID deduplication, state-transition balancing, and randomized choices. Released manifests preserve source identifiers and answer labels for auditing source proportions and answer distributions. These controls improve traceability but do not automatically remove public-data biases or contamination from model pretraining.
2. Task-specific scoring and frozen weights: distinguish abilities without repeatedly counting shared signal
Different tasks cannot all use answer-string accuracy. Binary and multiple-choice questions use normalized exact matching. Localization receives 100 times bounding-box IoU. Situation description receives 100 times the semantic-correctness probability from the official LingoQA text-classification judge; its 0.5 pass threshold is diagnostic only and does not enter the benchmark score. Sequence ordering receives partial credit according to the proportion of correctly ordered pairs.
Numeric tasks parse a valid number and use error decay rather than all-or-nothing credit. Counting error is the absolute prediction error divided by the larger of the target count and 1. Distance error is the absolute log ratio of predicted and target distances after adding 0.1. Both use the following bounded scoring function:
Counting uses \(\tau=0.05\) and \(\sigma=(0.20-0.05)/\ln 2\), giving half credit at normalized error 0.20. Distance uses \(\tau=0\) and \(\sigma=1.0\), giving half credit at log-ratio error \(\ln 2\). Small errors can thus receive partial credit while scores remain within 0โ100. Unparseable answers receive 0 and are logged. Samples with missing inputs, malformed records, or failed generations are instead logged and re-run when evaluation resumes; these cases should not be conflated.
Each task averages its sample scores, and each rank takes the arithmetic mean of its task scores. The headline open-loop score in Table 1 separately uses a fixed weighted sum of 14 task scores, not a simple average of three rank means. The authors first assign R1, R2, and R3 budgets of 0.3, 0.3, and 0.4, then distribute each budget according to subtask redundancy in an initial calibration pool:
Here \(\rho_{ij}\) is the Spearman correlation between subtasks, and \(B_{r(i)}\) is the corresponding rank budget. This expression combines Appendix F's redundancy load, inverse uniqueness, and within-rank normalization steps: within-rank positive correlations have penalty strength 1.0, cross-rank correlations have strength 0.35, and squared correlations are summed. Negative correlations are clipped to zero rather than rewarded as extra uniqueness. Weights are estimated once during the initial calibration phase after the first benchmark-testing round and then frozen, not recalculated for each new evaluated model.
This composite reduces repeated rewards from similar counting and distance tasks, but the 0.3/0.3/0.4 budgets and calibration model pool remain design choices rather than uniquely objective driving-capability weights. Task profiles remain essential: models with similar averages can differ substantially in localization and temporal memory.
3. Shared closed-loop interface and multiplicative metric: test understanding during continuous interaction
R4, Closed-loop Execution, uses CARLAโSUMO co-simulation on a real-road layout. CARLA provides the ego vehicle, visual observations, and vehicle actuation; SUMO controls background traffic; an interactive editor configures actors, trajectories, and traffic conditions. The 100 scenarios cover 10 families: pedestrian encounters, obstacle avoidance, turning, intersections, T-intersections, traffic disturbances, sudden braking, merging, yielding, and roundabouts. Real-road topology does not make this a real-road experiment.
At every decision step, all models receive a front-view image, a bird's-eye-view route image, and structured context including speed, remaining distance, navigation command, and short-horizon route points in ego coordinates. The default interface requests steering intent and target speed, not a long explanation or an independently executed full trajectory. A shared actuation bridge limits steering range and per-update change, then converts target-speed error into throttle or braking. Although it is not a separately trained controller for each model, it remains a common implementation condition underlying the closed-loop results.
Action execution runs at 4 Hz with a simulation step of 0.05 s, front-view resolution of 800ร450, and 1 retained dialogue turn. Steering changes by at most 0.12 per update. An unparseable action retains the previous valid control and records a failure, and the last control is held between model queries. CARLA restarts for every scenario; arrival, collision, persistent off-road behavior, blocking, or timeout terminates a run. R4 therefore measures observation understanding, action-format reliability, and continuous control together, rather than question answering alone.
The final score multiplies route progress, safety, and efficiency:
Route scores lie within 0โ100, while safety and efficiency lie within 0โ1. Arrival yields a route score of 100; otherwise, the score follows the completed fraction of the reference route. Progress updates only when projected lateral deviation does not exceed 6 m, and the default goal radius is 8 m. Safety multiplies event penalties and the fraction of time within the valid route region: the event multiplier is 0 for pedestrian collisions, 0.2 for vehicle collisions, 0.3 for static-object collisions, and 0.7 for scenario or blocked timeout. Other terminal-event penalties are implementation-defined; their values are not inferred here.
Efficiency compares actual distance with reference distance and actual duration with reference duration, allowing grace factors of 1.10 and 1.50. Multiplication prevents either unsafe completion or safe but prolonged inactivity from earning a high score through one favorable component. Released logs also retain route, safety, efficiency, parsing, and termination information. Diagnosis should inspect these components rather than infer a unique failure cause from their product.
4. Benchmark-guided supervised fine-tuning: select interventions from deficits without training on closed-loop test scenarios
Using Qwen3-VL-8B-Instruct, the authors compare each task with the best evaluated task score and identify deficits. Selected targets are nearest distance, referred distance, and localization in R1; multi-view memory and spatial relations in R2; and outcome prediction in R3. Single-task or grouped variants are trained, then changes in non-target tasks and R4 are examined. ALL denotes the jointly selected weak tasks, not all 14 task types and not additional closed-loop driving supervision.
The informative evidence lies outside the supervision scope. Target-task gains merely show that the model learned that supervision; non-target changes probe capability coupling, while R4, excluded from training, provides closed-loop transfer evidence. This remains an intervention case study on one base model. Cross-model correlations do not establish a causal hierarchy, and a single fine-tuning study does not establish that all models will benefit.
Loss & Training¶
Fine-tuning uses SWIFT's LoRA-based supervised fine-tuning pipeline, with the vision tower and aligner frozen and LoRA adaptation on linear layers. The paper explicitly excludes R4 scenarios from every training run. Each training dataset reserves 5% for validation. LoRA rank is 8, scaling factor is 32, learning rate is \(1\times10^{-4}\), warmup ratio is 0.05, and training uses bfloat16.
Per-device batch size is 1, gradient accumulation is 8, maximum sequence length is 16384, and the visual-token limit is 1024 per sample, using 4รA800. The paper does not introduce a new loss function here, so no specialized closed-loop objective is added. A fixed training recipe helps compare supervision groups, but training-set sizes, task mixtures, and overlap with evaluation sources still require manifest-level auditing.
Key Experimental Results¶
Main Results¶
The following representative models are drawn from Tables 1 and 2. Open-loop scores use frozen aggregation on the refined R1โR3 set; closed-loop scores use the routeโsafetyโefficiency composite over 100 scenarios. Although both occupy a 0โ100 scale, they measure different tasks and their magnitudes should not be compared directly.
| Model | Open-loop composite | R4 composite | R4 rank |
|---|---|---|---|
| Qwen3-VL(8B) | 45.09 | 0.902 | 9 |
| MiniCPM-V-4_5(9B) | 45.50 | 5.405 | 3 |
| Gemma-3-it(27B) | 49.53 | 12.619 | 1 |
| Qwen3-VL(32B) | 49.64 | 3.452 | 5 |
| Qwen2.5-VL(72B) | 54.47 | 9.366 | 2 |
| DA-DriveLM(4B) | 38.31 | 0.017 | 15 |
| ReasonDrive(7B) | 43.13 | 4.201 | 4 |
The open-loop leader is not necessarily the closed-loop leader: Qwen2.5-VL(72B) has the highest open-loop composite, whereas Gemma-3-it(27B) has a higher R4 score. MiniCPM has an open-loop score close to Qwen3-VL(8B) but a substantially different closed-loop result. Driving-specialized models also do not uniformly outperform generalist models, so successful domain adaptation cannot be inferred from one strong question-answering capability.
Ablation Study¶
The following supervision-group interventions correspond to Tables 3 and 4, not removal ablations of a new neural architecture. Open-loop changes retain Table 3's reported values, and R4 absolute scores retain Table 4's values.
| Qwen3-VL(8B) configuration | Open-loop composite change | R4 composite | R4 gain over baseline |
|---|---|---|---|
| Baseline | 0 | 0.902 | 0 |
| Nearest distance + referred distance + localization (R1 group) | +2.14 | 4.943 | +4.041 |
| Multi-view memory + spatial relations (R2 group) | +0.42 | 5.171 | +4.269 |
| Outcome prediction (R3_1) | +7.12 | 4.595 | +3.693 |
| ALL: joint training on all selected weak tasks | +20.88 | 8.021 | +7.119 |
Supervising referred-object distance alone improves nearest-object distance by +35.78, ordinary object counting by +10.91, and temporal counting by +3.03, demonstrating transfer across some numeric and memory tasks. Transfer is not uniformly positive: outcome-prediction supervision reduces multi-view memory by โ9.05, and the R2 combination reduces outcome prediction by โ19.90. Negative transfer should therefore remain visible rather than interpreting the authors' โno systematic collapseโ claim as an absence of any degradation.
Two source discrepancies require explicit boundaries. Table 1 reports Qwen3-VL(8B)'s open-loop composite as 45.09, whereas Table 3's Baseline* is 45.10; these are not silently harmonized. In Table 3, the three-task R1 combination improves localization by only +0.01, compared with +79.45 from localization-only supervision and +79.74 from ALL. This warrants verification given that the combination also supervises localization; +0.01 is not silently replaced with a larger value.
Key Findings¶
- Across 15 models, the Spearman correlations of R1, R2, and R3 with R4 are 0.664, 0.596, and 0.418. These support an association between open-loop capabilities and closed-loop performance, not causal claims that R1 causes R4 or that reasoning is unimportant.
- R1โR2 correlation is 0.843, while R1โR3 and R2โR3 correlations are 0.346 and 0.425. Nearest and referred distance tasks correlate at 0.961. The hierarchy retains shared and distinct variation, but its structure depends on the current tasks and evaluated models.
- The R2 combination gains only +0.42 in the open-loop composite yet achieves the highest single-group R4 score, 5.171. Outcome prediction gains +7.12 open loop but reaches 4.595 on R4. Composite-score improvement is not a reliable substitute for closed-loop transfer benefit.
- Even the strongest original model reaches only 12.619 on R4. Turning is particularly difficult, with that model scoring 0.212 in the turning family. No repeated-run confidence intervals are provided, so small score differences do not establish stable rankings.
Highlights & Insights¶
- Evaluation shifts from identifying the highest score to identifying strengths, weaknesses, and possible interventions. Combining task profiles with controlled fine-tuning gives more diagnostic value than a new leaderboard alone.
- Redundancy correction has a clear scope: reduce repeated measurement of similar abilities while freezing weights for stable subsequent comparison. The transferable principle is calibration followed by a fixed protocol, not copying these driving-rank budgets unchanged.
- A shared actuation bridge avoids different downstream controllers for each model while retaining parsing failures and execution constraints. Closed-loop capability is a property of the model, prompt, interface, and execution environment together.
Limitations & Future Work¶
- The authors explicitly acknowledge that simulated closed-loop performance is not a direct proxy for real-world deployment. A real-road layout, planned hardware integration, and controlled scenarios do not constitute real-world driving safety certification.
- R3 includes only outcome prediction and image ordering, under-covering long-horizon planning, counterfactual traffic interaction, and complex intent inference. Cognitive-rank terminology should not obscure the narrowness of the measured tasks.
- Single deterministic runs do not quantify variation from simulation, service latency, or parsing reliability. Repeated trials should report uncertainty in route, safety, efficiency, and failure events.
- Weights depend on the initial calibration pool, positive-correlation clipping, and rank budgets. Ranking stability should be tested across model pools and uniform-weight alternatives, alongside audits of training contamination, scene-level splits, and cross-dataset duplication.
- Fine-tuning is tested on one base model, and joint supervision contains a questionable localization result. Stronger conclusions require verification, additional model families, equal-data training, and random-target selection controls.
Related Work & Insights¶
- vs DriveLM and LingoQA: these provide driving-language questions and supervision resources; DriveHierarchy reorganizes multi-source tasks and adds closed-loop execution for capability analysis. Its value lies in the protocol, not in claiming that all integrated raw data were newly collected.
- vs Bench2Drive and Bench2Drive-VL: these emphasize closed-loop driving or VLM driving evaluation; DriveHierarchy emphasizes links between hierarchical open-loop profiles and closed-loop outcomes, adapting parts of Bench2Drive's scoring implementation. It is not a matched-environment replacement comparison against every earlier benchmark.
- Research direction: equal-data interventions across several base models could compare predicted and observed closed-loop improvements, testing whether deficit-guided supervision outperforms random task selection. This paper suggests that direction but does not establish such generality.
- Resources and provenance: this note is based on the arXiv v1 full text, methods, experiments, and Appendices AโI. The abstract provides the GitHub link above, while Appendix B separately lists the anonymous-review address
anonymous.4open.science/r/DriveHierarchy-2FD5; neither was verified during this offline writing task. Upstream datasets, checkpoints, and scoring code retain their own licenses, and project release does not authorize unrestricted redistribution of all raw assets.
Rating¶
- Novelty: 4/5 โ integrates hierarchical open-loop profiles, closed-loop execution, and intervention evidence within one diagnostic protocol; the main innovation is evaluation organization.
- Experimental Thoroughness: 3/5 โ covers 15 models and 100 scenarios but lacks repeated trials, interventions across base models, and extensive weighting sensitivity analysis.
- Writing Quality: 3/5 โ methods and appendices provide substantial implementation detail, but baseline rounding and joint-localization results need clarification.
- Value: 4/5 โ useful for identifying driving-VLM capability deficits; a research evaluation instrument, not proof of qualification for real-road deployment.