Skip to content

Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/reml-lab/DETRProbe3D/
Area: 3D Vision
Keywords: DETR, Probing, Monocular Depth Estimation, 3D Object Localization, Representation Learning

TL;DR

This paper introduces a probing framework to investigate the implicit geometric representations of 2D detection transformers, revealing that DETR query embeddings spontaneously encode accurate object-level depth and camera-frame 3D coordinates without any 3D supervision during pre-training.

Background & Motivation

Detection transformer models, pioneered by DETR and its variants, have fundamentally revolutionized the 2D object detection paradigm by pairing an encoder-decoder transformer with learned object queries. Trained end-to-end on standard 2D detection benchmarks like COCO with only discrete category labels and 2D bounding boxes, these architectures eliminate handcrafted components such as anchor generation and non-maximum suppression. In parallel, representation analysis on modern vision foundation models (such as DINOv2 and Stable Diffusion) has shown that self-supervised vision encoders inherently capture rich 3D geometric and physical priors in their dense feature maps, sparking intense curiosity regarding the geometric awareness of models trained without 3D supervision.

However, prior probing literature has focused almost exclusively on generic backbone features, dense multi-scale feature maps, or pooled global scene representations. Even within object detection interpretability studies, investigations have remained limited to convolutional architectures. For detection transformers, whether and to what degree their core decision carriers—the object-level query embeddings output by the transformer decoder—spontaneously encode 3D physical and geometric properties has remained completely unexplored.

Understanding the latent 3D information inside 2D DETR query embeddings is critical for clarifying how transformers represent objects in space and evaluating the transferability of 2D representations to downstream 3D spatial perception tasks. The core idea of this paper is to design an object-level probing framework for monocular depth estimation and camera-frame 3D object center localization, using linear and non-linear probes on frozen 2D detection transformer queries to uncover their emergent 3D spatial understanding.

Method

Overall Architecture

The paper formulates a structured probing pipeline to systematically evaluate the latent 3D object-level properties encoded within pre-trained 2D DETRs. The complete workflow spans four distinct phases: pre-trained detector forward inference, foreground object confidence filtering, cross-model geometric alignment, and probe training with frozen detector backbones and decoders. Given a single monocular RGB image, the pre-trained 2D detector produces a collection of object query embeddings along with corresponding predicted 2D bounding boxes and class probabilities; the filtered query embeddings are then passed to independent lightweight probe networks to regress object-centric metric depth or 3D bounding box center coordinates.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Monocular RGB Image Input"] --> B["Pre-Trained 2D DETR Inference<br/>Extract final decoder layer query embeddings q"]
    B --> C["Objectness Confidence Filtering<br/>Discard background queries below probability threshold"]
    C --> D["Cross-Model Alignment & GT Matching<br/>IoU greedy matching to establish probe pairs"]
    D --> E["Probe Representation Prediction<br/>Linear mapping and two-layer MLP probe regression"]
    E --> F["Predicted 3D Geometric Properties<br/>Metric depth d and 3D object center (x, y, z)"]

Key Designs

1. Objectness Confidence Filtering and Property Pairing: Constructing Unbiased Probe Datasets Because DETR models inherently output a fixed set of query embeddings (e.g., 100 or 300 queries per image), the vast majority of query slots correspond to empty background or unconfident candidates that cannot be mapped to actual physical objects. To build valid probe training pairs \((q_n, p_n)\), the pipeline first imposes a foreground probability threshold of \(\tau = 0.5\) on the predicted classification distribution, retaining only high-confidence object queries. For the monocular depth estimation task, the ground-truth scalar depth \(d_n\) is sampled from the dense depth map at the center coordinates of the predicted 2D bounding box. For the 3D location prediction task, the detector's predicted 2D bounding boxes are greedily matched with ground-truth 2D bounding boxes based on Intersection-over-Union (IoU); matched instances then inherit the bottom-center coordinates \((x_n, y_n, z_n)\) of the ground-truth 3D bounding box.

2. Cross-Model Geometric Alignment Filtering: Eliminating Confounding Detection Discrepancies Different detection transformer variants diverge in backbones, multi-scale feature hierarchies, and overall 2D detection recall, yielding differing sets of surviving detections across images. Evaluating probes across unaligned detection sets would conflate a detector's 2D recall performance with the true expressiveness of its latent geometric representations. To enforce strict comparability, the authors implement an anchor-based cross-model filtering procedure with standard DETR as the anchor. For every test image, predicted boxes from alternative models are greedily matched against the anchor's predicted boxes at an IoU threshold of 0.5. The pipeline retains only the strict intersection of object instances that are simultaneously detected by all evaluated models, establishing an identical, balanced benchmark for cross-architecture geometric decoding.

3. Two-Layer Non-Linear Spatial Probing: Decoding Non-Linear Manifold Geometries To dissect the structural complexity of latent representations, the probing protocol distinguishes between linearly readable information and non-linearly accessible geometric cues. The paper evaluates both a single affine linear probe \(f_{\theta}^{\text{lin}}(q) = W q + b\) and a two-layer multi-layer perceptron (MLP) probe \(f_{\theta}^{\text{mlp}}(q) = W_2 \text{ReLU}(W_1 q + b_1) + b_2\) with a hidden dimension of 256. Throughout probe training, all detector parameters remain completely frozen, and the probe is trained purely via Mean Squared Error (MSE) minimization. Comparing linear against non-linear probe performance directly gauges the curvature and non-linear structure of the geometric manifolds formed after multiple layers of cross-attention.

Loss & Training

During the probing stage, detector parameters are entirely frozen. The probe network parameters are optimized using the Mean Squared Error (MSE) objective:

\[\mathcal{L}_{\text{probe}}(\theta) = \frac{1}{M_{\text{tr}}} \sum_{n=1}^{M_{\text{tr}}} \| f_{\theta}(q_n) - p_n \|_2^2\]

Optimization uses the Adam optimizer with a batch size of 512 and an initial learning rate of \(10^{-3}\). For the monocular depth MLP probe experiments, training runs for 1000 epochs with a 20-epoch linear warm-up followed by cosine learning rate decay. Evaluation metrics include Mean Absolute Error (MAE), threshold accuracy (\(\delta_1 < 1.25\)), and Absolute Relative Error (AbsRel).

Key Experimental Results

Main Results

The authors evaluate five representative 2D DETR variants (DETR, Conditional DETR, Deformable DETR, LW-DETR, RT-DETR v2) across the synthetic outdoor driving dataset Virtual KITTI 2, the real-world indoor RGB-D dataset NYUv2, and the physical benchmark KITTI. Baselines include the foundation depth model Depth Anything 2 and the dedicated monocular 3D detector MonoDETR.

Monocular depth estimation results are summarized below:

Model Probe Virtual KITTI 2 (Outdoor) MAE (m) ↓ Virtual KITTI 2 AbsRel (%) ↓ Virtual KITTI 2 \(\delta_1\) (%) ↑ NYUv2 (Indoor) MAE (m) ↓ NYUv2 AbsRel (%) ↓ NYUv2 \(\delta_1\) (%) ↑
DETR Linear 1.01 8.29 93.10 0.58 24.30 57.22
Conditional-DETR Linear 1.13 9.07 92.57 0.53 22.52 60.96
Deformable-DETR Linear 1.06 8.34 91.78 0.56 23.45 59.91
LW-DETR Linear 1.26 9.84 91.38 0.59 25.30 58.70
RT-DETR v2 Linear 1.56 12.49 86.21 0.64 26.80 55.22
DETR MLP 0.54 3.62 98.81 0.51 20.20 65.48
Conditional-DETR MLP 0.69 4.39 98.67 0.50 19.99 66.78
Deformable-DETR MLP 0.65 4.05 98.94 0.54 21.49 63.65
LW-DETR MLP 0.88 5.63 98.01 0.55 22.65 62.70
RT-DETR v2 MLP 1.05 6.73 97.21 0.56 22.70 60.43
Depth Anything 2 (Zero-shot) - 1.49 9.13 97.88 0.61 24.21 59.57
MonoDETR (Zero-shot) - 3.83 39.82 64.72 10.59 497.72 0.00
MonoDETR (MLP head) - 0.47 3.14 99.42 0.75 21.49 62.50
Depth Anything 2 (MLP head) - 0.50 3.30 99.47 0.36 13.36 84.00

For 3D object center localization on the KITTI benchmark, coordinate-wise and overall center errors are reported below:

Model x MAE (m) ↓ y MAE (m) ↓ z MAE (m) ↓ Center MAE (m) ↓ Center AbsRel (%) ↓
DETR 0.33 0.12 1.03 0.50 2.40
Conditional-DETR 0.68 0.16 1.09 0.64 3.01
Deformable-DETR 0.66 0.16 1.04 0.62 2.90
RT-DETR v2 2.27 0.18 1.39 1.28 6.49
LW-DETR 2.28 0.17 1.20 1.22 6.01
MonoDETR (Full 3D Supervision) 0.19 0.06 0.64 0.30 1.41

Ablation Study

To confirm that probe accuracy originates from genuine high-dimensional query representations rather than trivial geometric bounding box heuristics (such as vertical image position and scale priors), the authors compare query embedding probes against baseline probes fed solely with predicted 2D bounding boxes \((x, y, w, h)\):

Input Representation Probe Architecture Virtual KITTI 2 MAE (m) ↓ Virtual KITTI 2 AbsRel (%) ↓ Virtual KITTI 2 \(\delta_1\) (%) ↑ NYUv2 MAE (m) ↓ NYUv2 AbsRel (%) ↓ NYUv2 \(\delta_1\) (%) ↑
DETR Bounding Box (Bbox) Linear 4.27 32.67 44.69 0.87 37.61 41.74
DETR Query Embedding Linear 1.01 8.29 93.10 0.58 24.30 57.22
DETR Bounding Box (Bbox) MLP 1.02 7.09 95.76 0.77 32.98 51.83
DETR Query Embedding MLP 0.54 3.62 98.81 0.51 20.20 65.48

Key Findings

  • Non-linear probes substantially outperform linear probes: On Virtual KITTI 2, non-linear MLP probes cut the MAE of DETR embeddings from 1.01m to 0.54m and reduce AbsRel from 8.29% to 3.62%, indicating that 3D geometric information resides in non-linear query manifold structures.
  • Outperforming zero-shot 3D foundation models: Standard 2D DETR models equipped with MLP probes attain lower depth estimation MAE than zero-shot Depth Anything 2 across both Virtual KITTI 2 and NYUv2 (e.g., 0.54m vs. 1.49m on Virtual KITTI 2).
  • Narrow gap to fully supervised 3D detectors: With a median object distance of 25.3m on KITTI, frozen 2D DETR achieves a 3D center MAE of 0.50m (2.40% relative error), trailing MonoDETR (0.30m MAE, trained with full 3D bounding box supervision) by merely 0.20 meters.
  • Real-time and lightweight variants degrade 3D representations: RT-DETR v2 and LW-DETR exhibit notably higher depth and localization errors (KITTI center MAE of 1.28m and 1.22m), suggesting that reduced decoder depth and restricted cross-scale attention impair the capture of latent geometric cues.
  • Geometric information concentrates in low-dimensional subspaces: PCA compression shows that probing accuracy plateaus once retaining roughly 40-60 principal components, confirming that depth-relevant signals are densely concentrated in a low-dimensional subspace of the query embeddings.

Highlights & Insights

  • Shifting the interpretability paradigm: Moving beyond dense backbone feature maps, this work probes the fundamental discrete decision unit of modern vision transformers—the object query embedding—verifying that high-level object semantics implicitly encode metric physical geometry.
  • Proving representation richness beyond 2D boxes: The ablation study decisively rules out the hypothesis that probes merely exploit 2D box perspective geometry, demonstrating that query embeddings contain rich semantic context, category size priors, and occlusion relationships.
  • Architectural implications for efficient detectors: Real-time simplifications that sacrifice cross-scale and full-decoder interactions incur significant representational penalties in latent spatial geometry, providing critical design considerations for embodied AI and unified multi-task detectors.

Limitations & Future Work

  • Author-admitted limitations: The investigated 3D attributes are currently restricted to scalar depth and 3D bounding box center coordinates, leaving full 3D bounding box orientations (yaw/pitch/roll) and physical dimensions (length/width/height) unexplored. Furthermore, NYUv2 indoor evaluations are constrained by sensor noise in monocular depth maps.
  • Potential limitations: While cross-model geometric alignment guarantees fair benchmarking, filtering for mutual detection intersections inevitably selects for more visible objects, potentially overestimating geometric performance under severe occlusion or extreme lighting. Additionally, the two-layer MLP probe may partially function as a coordinate frame calibrator between latent space and metric units.
  • Future directions: Incorporating frozen 2D DETR query embeddings as plug-and-play geometric priors into monocular 3D detection and autonomous driving pipelines, bypassing costly from-scratch 3D pre-training.
  • vs Depth Anything / DINOv2 Probing: Earlier studies (e.g., El Banani et al.) demonstrated dense 3D awareness in self-supervised backbones requiring dense decoders; this paper establishes that discrete object queries in task-specific 2D detectors also spontaneously retain precise physical 3D localization.
  • vs MonoDETR: MonoDETR demands a dual-pathway depth-guided encoder and expensive 3D bounding box annotations; this work proves standard 2D DETRs already construct high-quality 3D geometric representations, requiring only a lightweight readout head to approach supervised performance.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ First systematic probing study exploring the emergent 3D object-level understanding inside 2D detection transformer query embeddings.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous coverage across five DETR architectures, indoor/outdoor benchmarks, linear/non-linear probes, PCA compression, and 2D box ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Lucid experimental protocols, meticulous variable isolation, and persuasive comparative analysis.
  • Value: ⭐⭐⭐⭐⭐ Delivers vital theoretical and empirical foundations for transfer learning, monocular 3D detection, and embodied spatial perception.