Skip to content

๐Ÿš— Autonomous Driving

๐Ÿง  NeurIPS2026 ยท 5 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (159) ยท ๐Ÿ“ท CVPR2026 (157) ยท ๐Ÿ”ฌ ICLR2026 (50) ยท ๐Ÿงช ICML2026 (8) ยท ๐Ÿค– AAAI2026 (56) ยท ๐Ÿง  NeurIPS2025 (47)

AutoExpert: Automating 3D LiDAR Annotation from Expert-Crafted Guidelines

AutoExpert turns textual expert rules and a few 2D examples into an adapted detector, then uses a vision-language model to supply instance-specific size and orientation priors for fixed-size multi-hypothesis search, achieving 25.4 mAP3D on AutoExpert-nuScenes without target-task 3D training annotations.

DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution

DriveHierarchy organizes vision-language driving capabilities into perceptual grounding, contextual memory, mental reasoning, and closed-loop execution, diagnoses 15 models through unified open-loop tasks and 100 interactive simulation scenarios, and shows on Qwen3-VL-8B that repairing selected open-loop weaknesses can raise the closed-loop composite score from 0.902 to 8.021, without establishing real-world driving safety.

Driving Video Retrieval for Complex Queries with Structured Grounding

STRIVE-D calibrates motion rules against a weak-label ranking objective on separate driving videos, reuses or adapts those rules per query, and fuses them with visual and lexical retrieval, raising DrivingDojo Acc@1 from the main table's strongest dense baseline of 14.6% to 26.8%; it does not simply ask an LLM to inspect videos and determine true geometry.

Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

MaDE learns inverse dynamics and physical residuals from state transitions without control labels, then corrects predicted trajectories in control space so that outputs satisfy its learned discrete dynamics map by construction while reducing inequality violations on a best-effort basis; known-model residuals fall substantially on inD, but position errors increase, without establishing true-physics correctness or safety.

Progressive Risk Estimation for Accident Anticipation

PRE-ACT uses continuous distance-to-event progress supervision and within-video clip ranking to improve traffic risk prediction from past-only sliding windows, raising CAP mAUC from TOP's 0.429 to 0.481 under FPR โ‰ค 0.1 and introducing a Separation Score to examine false-alarm tendencies across complete risk curves.