Skip to content

AutoExpert: Automating 3D LiDAR Annotation from Expert-Crafted Guidelines

Conference: NeurIPS 2026
arXiv: 2506.02914
Area: Autonomous Driving
Keywords: expert-crafted annotation guidelines, multimodal few-shot learning, 3D LiDAR annotation, instance-specific geometric priors, multi-hypothesis testing

TL;DR

AutoExpert turns textual expert rules and a few 2D examples into an adapted detector, then uses a vision-language model to supply instance-specific size and orientation priors for fixed-size multi-hypothesis search, achieving 25.4 mAP3D on AutoExpert-nuScenes without target-task 3D training annotations.

Background & Motivation

Annotating autonomous-driving data involves more than drawing a box around every visible car. Expert guidelines determine category boundaries, box contents, and special cases: nuScenes requires a bicycle box to include its rider when present, while groups of bicycles may require collective annotation. An ordinary open-vocabulary detector learns common categories from Internet data; prompting it with bicycle or police-officer does not ensure compliance with these rules. Recognizing more categories and following domain-specific annotation standards are different problems, and a larger general-purpose model does not automatically solve the latter.

Human annotators can learn rules from text and a few images, then perform a task in another modality. Machines face two transfers: supervision consists of 2D images and text, whereas the required output consists of 3D cuboids in LiDAR point clouds; example images often annotate only the category being explained, leaving other visible objects unlabeled. Conventional detector training would incorrectly treat these objects as negatives. Directly clustering points selected by a 2D box also risks including fences, backgrounds visible through windows, and points seen through bicycle wheels. Inadequate semantic supervision and contaminated geometric evidence arise together.

The paper formalizes this task as AutoExpert and builds benchmarks using authentic guidelines defining 18 nuScenes classes and 25 PandaSet classes. To avoid potential copyright issues with the guidelines' original Internet images, the authors select 4โ€“8 images per class from official training data to simulate iconic guideline examples; their 2D boxes and accompanying text form the few-shot training material. Validation simulates ongoing expert quality control, rather than establishing a fully unattended system. Core Idea: adapt expert rules into 2D detection, then use VLM-derived instance geometry to replace unconstrained fitting of contaminated point clouds with fixed-size, constrained position and orientation search.

Method

Overall Architecture

The final method is called auto3D. Inputs include expert text, few-shot 2D boxes, unlabeled multiview RGB images and LiDAR sweeps, known camera intrinsics and extrinsics, and sensor poses; outputs are class-labeled 3D cuboids with confidence scores. Preparation refines category prompts and fine-tunes GroundingDINO. Annotation performs 2D detection and SAM segmentation, associates LiDAR points, generates cuboids, and improves scores using temporal information.

The cross-modal bridge is sensor calibration, not a directly trained LiDAR foundation model: a 2D box defines a frustum, and the segmentation mask further removes points projecting outside the foreground. A VLM examines the object image to estimate dimensions and visible faces, while MHT checks candidate placements against actual points and 2D projections. Text specifies what should be annotated, visual priors describe the approximate geometry of this instance, and LiDAR constrains its physical location.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    G["Guidelines + few-shot 2D boxes"] -.->|Training supervision and validation selection| A["Guideline-adapted detection"]
    I["Unlabeled RGB images"] --> A
    A --> B["Offline sweep association"]
    P["LiDAR sweeps + calibration"] --> B
    B --> C["Instance-prior-constrained search"]
    A -->|Object image and category size prior| C
    C --> D["Geometric-temporal scoring"]
    D --> O["Class-labeled and scored 3D cuboids"]

The dashed edge denotes preparation-stage training and model selection only; solid edges show annotation-time data flow. SAM supplies foreground filtering within offline sweep association. Instance-prior-constrained search encompasses both VLM size/orientation estimation and subsequent MHT: the VLM does not directly output the final 3D position.

Key Designs

1. Guideline-adapted detection: convert expert definitions into prompts and few-shot supervision

A vision-language model (VLM) first reads a category's textual definition and example images to generate five descriptive terms or synonyms, together with a prior for its average real-world length, width, and height. The terms and their combinations are then passed to GroundingDINO, and the prompt giving the best category-specific 2D detection precision on validation data is selected. For police-officer, the selected prompt is law enforcement officer; pushable-pullable can be covered by descriptions of garbage containers and hand trucks. Predictions remain in the original guideline taxonomy: the model-facing expression is refined, not the labels themselves.

The VLM does not simply declare which term is best, nor does semantic similarity alone determine selection: candidates must be checked through actual detections. This maps specialized guideline terminology to visual concepts familiar to the foundational detector, but subsequent few-shot training is still needed to calibrate box extent and fine-grained categories. In Table 2, refined prompts increase zero-shot 2D mAP from 16.9 to 18.2, while the corresponding 3D mAP falls from 16.1 to 15.7; prompt optimization alone therefore does not guarantee downstream improvement.

Fine-tuning combines the selected terms with 4โ€“8 visual examples per class to form a multimodal few-shot training set. Its crucial constraint is federated annotation: a police example guarantees annotations for police officers, not for every car or truck in the background. Loss is computed only for the image's target category, without penalizing detections of other categories as false positives. Here, federated means category-local annotation, not distributed federated learning. The 2D detector is genuinely trained; the method is neither entirely zero-shot nor training-free.

2. Offline sweep association: obtain better geometric evidence through calibration and category-aware aggregation

After 2D detection, SAM produces an instance foreground mask. Known sensor parameters project LiDAR points into the image, and the frustum and mask establish associations. A mask identifies foreground pixels, but does not guarantee that every 3D point projecting there belongs to the object: windows can reveal distant backgrounds, fences can occlude a vehicle, and bicycle wheel openings can reveal other objects. Associated points therefore remain contaminated evidence, so their enclosing box should not directly determine object dimensions.

A single sweep can also be too sparse. The authors compare the current sweep alone, past sweeps, future sweeps, and symmetric aggregation. Future frames are accessible because this is offline data annotation, not a directly transferable causal online-driving setting. Aggregation is category-aware rather than using the same sweep count for every object. For bicycle in Table 3, the current sweep yields 30.1, adding two future sweeps yields 32.4, and adding two past sweeps yields only 28.6; traffic-cone gives 52.1, 54.1, and 50.3, respectively.

The authors also track individual instances and aggregate only those deemed static, but this achieves 22.1 3D mAP, below category-aware aggregation's 22.8. Static-object decisions can be wrong, and moving objects still need denser point clouds. Category-aware aggregation introduces motion smears, but subsequent search prevents arbitrary cuboid expansion, limiting the inflation of geometric extent by those smears. This does not amount to explicitly eliminating motion distortion.

3. Instance-prior-constrained search: determine cuboid size before finding its placement

v-MHT, or VLM-Guided Multi-Hypothesis Testing, highlights the target detection with a green box and gives the VLM its category's average dimensions. The category prior anchors scale, while the instance image lets the model distinguish a sedan from a larger van within the car class and adjust length, width, and height. Orientation is not obtained merely by asking the VLM to guess an absolute angle: it identifies the object's image location and visible front, rear, or side, which are combined with camera extrinsics to estimate initial orientation. The appendix prompt also requests checks using lanes, road geometry, and traffic flow.

An instance-specific cuboid is then initialized, and center locations and yaw angles are enumerated within its frustum. Length, width, and height remain fixed throughout search; location and rotation can change. Rotation is constrained around the initial orientation instead of blindly searching a full circle for every object. Fixed dimensions prevent backgrounds, occluders, or aggregation smears from inflating candidate boxes, while semantic orientation helps distinguish geometrically similar vehicle fronts and rears, reducing 180-degree flips. Appendix Table 10 reports 19.2 mAP3D and 23.8 NDS for category-level MHT, versus 21.9 and 25.2 for instance-level v-MHT, indicating that the gain is not simply from more search iterations.

Candidates must cover associated points while remaining consistent with the original 2D detection. Appendix E defines the following objective, where \(P\) is the associated point set, \(B_{3D}(\mathbf{\Theta})\) is the candidate cuboid, \(\pi\) is camera projection, \(B_{2D}\) is the 2D detection box, and the search state \(\mathbf{\Theta}=[x,y,z,\psi]\) contains location and yaw:

\[ R(\mathbf{\Theta})=\frac{1}{|P|}\sum_{p_i\in P}\mathbb{I}\big(p_i\in B_{3D}(\mathbf{\Theta})\big),\qquad \mathbf{\Theta}^{*}=\arg\max_{\mathbf{\Theta}}\left[R(\mathbf{\Theta})+\operatorname{IoU}\big(\pi(B_{3D}(\mathbf{\Theta})),B_{2D}\big)\right]. \]

Point coverage supplies 3D evidence, while projection IoU discourages candidates from drifting away from the visual object. This is more appropriate than merely enclosing as many points as possible, but still depends on association and 2D-box quality and cannot guarantee correctness under every occlusion. Default translation and rotation steps are 0.5 m and \(\pi/10\), respectively; Numba and GPU parallelization accelerate candidate evaluation.

Appendix F further describes confidence-aware routing: detections above 0.3 use VLM instance priors and local search, while those below 0.3 bypass the VLM and fall back to category-average dimensions and full rotation search. The paper does not specify the branch at exactly the threshold. This implementation boundary matters: not every detection in the final system necessarily receives an instance-specific VLM estimate.

4. Geometric-temporal scoring: combine visual confidence, point-cloud support, and track consistency

Once a cuboid is generated, 2D detection confidence need not reflect its 3D quality. The authors project the cuboid onto the bird's-eye-view plane, divide its rectangle into a 7ร—7 grid, and count cells containing at least one LiDAR point. Occupancy measures whether the cuboid has reasonably distributed point-cloud support. It is the proportion of occupied cells, not point count, 3D IoU, or the MHT point-coverage ratio above. With \(N\) denoting occupied cells, the fused score is:

\[ S_{\text{3D}}=\frac{N}{7^{2}},\qquad S=\alpha S_{\text{2D}}+(1-\alpha)S_{\text{3D}}. \]

The coefficient \(\alpha\) is selected using validation-set 3D mAP; a fixed numerical value should not be invented. Spatially close cuboids with the same class are then heuristically linked across consecutive sweeps, and individual scores are replaced by the track's mean score. This changes confidence, not the cuboid positions, and does not train a motion model.

The unified 3D coordinate space naturally accommodates multiple cameras and is more suitable here than maintaining identities separately in 2D with SAM2, which must also handle perspective-dependent scale, occlusion, and cross-camera matching. Category-aware aggregation densifies evidence; geometric scoring and track averaging stabilize ranking. These should not be conflated: adding points, generating cuboids, and revising confidence are distinct stages.

A Worked Example

Consider a bicycle detection that includes a rider. Guideline adaptation teaches the detector to include the rider rather than outputting only the bicycle frame. SAM supplies the corresponding foreground, and calibration associates relevant LiDAR points. Table 3 compares the current sweep with two future sweeps to demonstrate category-level aggregation effects, but those aggregate results do not guarantee improvement for this particular instance.

Associated points may still include backgrounds visible through wheel openings, so their spatial extent cannot directly define bicycle length. The VLM uses the image and category scale to estimate instance dimensions and visible direction; MHT fixes those dimensions and changes location and yaw to jointly satisfy point coverage and 2D projection. BEV occupancy then enters the fused confidence, followed by 3D track-score averaging. If the scene contains a bicycle group that the guidelines require annotating collectively, independent detections can still violate the rules, as illustrated by the paper's failure cases.

Loss & Training

The 2D stage fine-tunes foundational model weights with few-shot supervision and augmentations including random rotation and cropping; category-local annotation determines which detection supervision is valid. Validation supports prompt selection, model selection, and hyperparameter tuning. For nuScenes, 570 frames from the official training split form validation, and the official validation split's 6,019 frames become the benchmark test set. The main text specifies 8 validation and 192 test frames for PandaSet.

Appendix O reports AdamW with learning rate and weight decay both set to \(10^{-4}\), using four NVIDIA A100 GPUs. The v-MHT geometric objective evaluates annotation-time candidates; it is not an end-to-end back-propagated training loss. The main method uses no target-task 3D training boxes. Appendix I additionally trains PointNet to refine 3D dimension residuals using a few extra 3D annotations; this relaxes the supervision setting and must not be folded into the main results obtained without 3D training annotations.

Key Experimental Results

Main Results

mAP3D averages AP over 18 categories and ground-plane center-distance thresholds \(\{0.5,1.0,2.0,4.0\}\) m; it is not 3D IoU AP. mAP2D averages category AP at IoU=0.5. NDS additionally incorporates translation, scale, orientation, velocity, and attribute errors. The following excerpts Table 1; mAP and NDS are shown on the paper's percentage scale.

AutoExpert-nuScenes method mAP3D NDS Comparison condition
CM3D 12.1 16.6 Original 2D detector
CM3D + ft-GD 18.2 23.1 Guideline-adapted 2D detector replacement
CPD + frustum 17.9 22.3 Category detection and targeted frustum
AnnofreeOD + ft-GD 17.3 22.1 Likewise enhanced 2D detection
auto3D 25.4 27.2 All components

auto3D improves over original CM3D by 13.3 mAP3D points and over CM3D already using ft-GD by 7.2 points. The latter comparison better isolates 3D generation and post-processing benefits; the entire gain cannot be attributed to v-MHT. Some self-supervised baselines require additional 2D category matching, so these are comparisons of adapted complete annotation pipelines.

Ablation Study

The 2D adaptation ablation excerpts Table 2 with the corresponding 3D generation pipeline; its rows are not complete auto3D variants.

2D detector configuration mAP2D mAP3D NDS
GD + original names 16.9 16.1 21.3
ft-GD + original names 20.0 16.6 21.2
GD + refined names 18.2 15.7 22.1
ft-GD + refined names 20.8 18.2 23.1

The cumulative 3D experiment excerpts Table 4, starting from CM3D + ft-GD and progressively adding components.

Cumulative configuration mAP3D NDS mAP3D gain over previous row
Baseline 18.2 23.1 Not applicable
+ v-MHT 21.9 25.2 3.7
+ Category-aware sweep aggregation 22.8 25.9 0.9
+ Geometric scoring 23.6 26.4 0.8
+ 3D track scoring 25.4 27.2 1.8

Key Findings

  • v-MHT contributes the largest gain in this accumulation order, but the study does not test every ordering and cannot establish order-independent effects. Refined prompts combined with few-shot fine-tuning outperform either alone.
  • Table 12 reports total time per sweep on four A100 GPUs: 0.74 s for CM3D, 0.95 s for ordinary MHT, and 0.64 s for v-MHT. v-MHT includes 0.40 s of VLM processing and 0.15 s of 3D generation; these are batched results on specified hardware, not online per-object latency.
  • PandaSet Table 22 reports 18.4 mAP3D and 27.6 NDS for auto3D, versus 12.3 and 16.0 for CM3D. Its NDS omits velocity and attribute errors and is renormalized, so it is not directly comparable to nuScenes NDS.

Highlights & Insights

  • The task separates label recognition from compliance with annotation rules. Few-shot boxes communicate domain-specific extent and conventions, not just category appearance.
  • Fixed instance dimensions protect against noise rather than merely initializing geometry. Search cannot absorb background points by expanding a cuboid, assigning distinct responsibilities to semantic priors and point-cloud evidence.
  • A VLM can reduce total computation by narrowing downstream geometric search. This depends on reliable priors, confidence-aware routing, and parallel implementation; adding reasoning alone does not imply acceleration.

Limitations & Future Work

  • Independent bicycle detections where collective annotation is required, and incorrect cuboids caused by fences, show that refined categories and isolated instance search do not fully implement guideline logic. Group rules and contextual constraints are promising extensions.
  • Future sweeps, expert validation, and known calibration are important conditions. This is an offline assisted-annotation pipeline, not a fully unattended, training-free, real-time driving system; validation supervision also has an operational cost.
  • The source contains numerical or wording inconsistencies: Table 17's caption claims some categories prefer two past sweeps, whereas its numbers and Table 3 support two future sweeps; current-sweep construction-worker values are 25.7 and 25.6, respectively. The examples here explicitly use main-text Table 3 rather than silently reconciling them.
  • Appendix Table 16 gives 1.133 orientation error for its unrefined baseline, compared with 0.992 in Table 1. The main text specifies an 8/192 validation/test split for PandaSet, while Table 22's caption states 200 frames. Results are reported according to their respective tables; these discrepancies require author or code clarification.
  • Rare classes remain difficult: auto3D's nuScenes AP values for child, police-officer, and debris are 5.4, 3.6, and 0.1. The paper provides no error bars; more repeated experiments, guideline-transfer tests, and active expert correction would be useful.
  • vs GroundingDINO / ordinary open-vocabulary detection: these recognize categories using general visual-language knowledge; this work additionally selects guideline prompts and fine-tunes with category-local annotations. Open-vocabulary capacity is a starting point, not expert-rule compliance.
  • vs CM3D / OpenBox: CM3D uses maps and lanes among its geometric cues, while OpenBox uses category-size priors and point clustering. This work constrains fixed-size search using image-derived instance priors and adds offline aggregation and temporal scoring.
  • vs Oyster / LISO / CPD: these learn 3D proposals from unlabeled points but generally do not output guideline categories. This work adds categories and frustums to strengthen the baselines, demonstrating that 2D rule adaptation remains important.

Rating

  • Novelty: 4/5 โ€” Authentic guideline-driven cross-modal annotation is closer to production requirements than ordinary open-vocabulary detection.
  • Experimental Thoroughness: 4/5 โ€” Two benchmarks, strengthened baselines, and component analyses provide broad coverage, but error bars and appendix consistency are lacking.
  • Writing Quality: 3/5 โ€” The pipeline and mechanisms are understandable, while some captions, error values, and evaluation-frame counts require clarification.
  • Value: 4/5 โ€” A reusable route to offline assisted LiDAR annotation, not yet a replacement for expert quality control.