Skip to content

The 3D Mirage: Probing and Taming 3D Hallucinations

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/hdnndh/The-3D-Mirage-Probing-and-Taming-3D-Hallucinations
Area: 3D Vision / Hallucination Detection
Keywords: monocular depth estimation, 3D hallucination, contextual stability, self-distillation, parameter-efficient fine-tuning

TL;DR

Addressing the systemic vulnerability where monocular depth foundation models hallucinate prominent non-existent 3D obstacles on planar optical illusions when surrounding context is restricted, this paper introduces the 3D-Mirage benchmark, reference-free second-order metrics (DCS and CCS), and a parameter-efficient Grounded Self-Distillation framework that cuts geometric hallucination by over 94% without catastrophic forgetting.

Background & Motivation

Monocular depth estimation (MDE) foundation modelsโ€”such as Depth Anything V2, ZoeDepth, and Depth Proโ€”have achieved remarkable zero-shot generalization across diverse visual scenes. To infer depth from a single image, an inherently ill-posed problem where projective geometry discards metric scale, these models rely heavily on large-scale statistical and semantic priors learned from diverse datasets. However, this impressive capability conceals an unexamined safety-critical flaw: the network frequently trades strict local geometric fidelity for semantic consistency, substituting brittle visual priors for physical evidence.

This reliance manifests as a dramatic failure mode when encountering real-world visual illusions (e.g., 3D street chalk art, forced-perspective murals, and deceptive billboards). On low-curvature or flat carrier surfaces, models perceive phantom 3D bumps, hollows, or obstacles. More critically, this vulnerability is exacerbated under context variations: while a global view might provide sufficient surrounding context to infer a flat roadway, restricting the field-of-view (FOV) via cropping or localized occlusion removes broader contextual anchors, triggering severe hallucinated geometry. Existing benchmarks (KITTI, ScanNet, NYUv2) lack such adversarial optical configurations, and standard pixel-averaged metrics like MAE, RMSE, and AbsRel cannot isolate localized structural distortions or quantify cross-view instability without ground-truth depth.

To bridge this diagnostic and methodological gap, a holistic probing, scoring, and taming paradigm is required. Core idea: Introduce the 3D-Mirage paired-view benchmark along with reference-free second-order metrics (DCS and CCS), and develop Grounded Self-Distillation (GSD)โ€”a LoRA-based dual-branch framework that suppresses spurious ROI curvature via multi-surface plane-mixture fitting while preserving frozen teacher geometry on background regions to eliminate 3D hallucinations without catastrophic forgetting.

Method

Overall Architecture

The framework encompasses three coordinated components: the 3D-Mirage benchmark designed to induce contextual instability, reference-free second-order evaluation metrics to quantify structural hallucination, and a parameter-efficient Grounded Self-Distillation (GSD) mitigation pipeline.

During GSD training, each instance consists of a full image \(x_{\text{full}}\) and a context-restricted crop \(x_{\text{crop}}\) centered on the illusion region. A frozen Depth-Anything-V2-Large acts as the teacher network \(T\), providing stable baseline depth representations. The student network \(S\) introduces low-rank adaptation (LoRA) modules into the patch embedding and MLP layers of the Vision Transformer (ViT) encoder (comprising only ~4M parameters, ~1.2% of the backbone), while keeping the decoder frozen. The student processes both views through shared weights. The Hallucination Knowledge Re-editing (HKR) loss penalizes second-order curvature within annotated illusion ROIs and aligns predictions with local surface hypotheses fitted from teacher ring statistics. Concurrently, the Non-hallucination Knowledge Preservation (NKP) loss distills the teacher's stable geometry on background pixels and edge boundaries, preserving general depth estimation capabilities.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Paired Views<br/>(Full Image x_full & Restricted Crop x_crop)"] --> B["Dual-Branch Feature Propagation<br/>(Frozen Teacher T & LoRA-adapted Student S)"]
    B --> C["Dual-View Projection & Second-Order Probing<br/>(DCS Structural Deviation + CCS Contextual Confusion)"]
    C --> D["Local Surface Fitting & Gating Mixture<br/>(Ring-Fitted Hypotheses ฯ€_k + Gating Weights w_j)"]
    D --> E["Hallucination Knowledge Re-editing HKR<br/>(Curvature Penalty + Multi-Hypothesis Alignment)"]
    D --> F["Non-hallucination Knowledge Preservation NKP<br/>(Background Distillation + Seam/Edge Boundary Regularizer)"]
    E --> G["Combined Objective Optimization<br/>(Parameter-Efficient Update + Regularizer Interleaving)"]
    F --> G

Key Designs

1. 3D-Mirage Benchmark and Second-Order Reference-Free Probing Metrics

To expose contextual instability on deceptive surfaces without relying on unobtainable dense 3D ground truth, the 3D-Mirage benchmark curates 468 real-world scenes containing 3D paintings, anamorphic street art, and murals. Precise polygonal region-of-interest (ROI) masks delineate the low-curvature support planes, incorporating nested exclusion masks to isolate authentic objects situated atop the illusion. For each scene, four random crops retaining at least 40% of the ROI diagonal simulate restricted fields-of-view, producing 1,872 full-crop evaluation pairs.

For a depth model \(f_\theta\), predictions \(D_{\text{full}}\) and \(D_{\text{crop}}\) are computed. The full-view prediction is aligned to the crop coordinate frame (\(\tilde{D}_{\text{full}\rightarrow\text{crop}}\)). Applying quantile normalization \(\mathrm{QN}(\cdot; M_i)\) and a separable second-difference magnitude operator \(\mathcal{L}(\cdot)\) produces normalized second-order response maps \(R_i^{\text{full}}\) and \(R_i^{\text{crop}}\). Two summary scalars are extracted per branch: the top-10% response sum \(t_{v,i}\) and the 10%-trimmed mean \(\mu_{v,i}\). Two complementary metrics are formulated: 1. Deviation Composite Score (DCS): Measures the absolute hallucination intensity by capturing spurious second-order surface curvature: $\(\mathrm{DCS} = \|\bar{\mathbf{t}}\|_2 + \frac{1}{N}\sum_{i=1}^N \|\mathbf{t}_i\|_2\)$ where \(\mathbf{t}_i = [t_{\text{full},i}, t_{\text{crop},i}]^\top\). Lower DCS indicates that the low-curvature carrier is correctly reconstructed without artificial geometric protrusions. 2. Confusion Composite Score (CCS): Quantifies contextual instability across varying FOVs by projecting mean responses along the off-diagonal unit vector \(\mathbf{u} = \frac{1}{\sqrt{2}}[1, -1]^\top\): $\(\mathrm{CCS} = |\mathbf{u}^\top \bar{\boldsymbol{\mu}}| + \frac{1}{N}\sum_{i=1}^N |\mathbf{u}^\top \boldsymbol{\mu}_i|\)$ A low CCS guarantees that the model's structural prediction remains robust and invariant to the presence or absence of peripheral scene context.

2. Hallucination Knowledge Re-editing (HKR): Curvature Suppression and Gated Hypothesis Mixture

Enforcing a naive global flat-plane constraint fails when an illusion spans multi-faceted support surfaces, such as perpendicular walls or road-curb transitions. HKR resolves this by combining explicit curvature penalties with an adaptive mixture of local planar hypotheses.

For each illusion polygon \(j \in \mathcal{R}\), up to \(K\) simple geometric plane hypotheses \(\{\pi_{j,k}\}_{k=1}^K\) are fitted to the frozen teacher's depth within an adjacent dilated ring band, with corresponding residual scales \(\sigma_{j,k}\). A compact gating MLP maps ROI and ring visual statistics to mixture logits \(w_j = \text{softmax}(G(\cdot))\), supervised via soft cross-entropy against residual targets \(q_j\) derived from \(\sigma_{j,k}\). The deviation between student normalized depth \(z\) and each fitted plane \(\pi_{j,k}\) is computed as \(\ell_{j,k} = \overline{|z - \pi_{j,k}|}_{m_j}\). The total HKR loss penalizes the union ROI curvature while pulling instances toward plausible surface mixtures: $\(\mathcal{L}_{\text{HKR}} = \alpha \overline{|\mathcal{L}(s)|}_m + \sum_{j \in \mathcal{R}} \sum_{k=1}^K w_{j,k} \ell_{j,k}\)$ Computed across both full and cropped views, this formulation flattens spurious 3D curvature while accommodating naturally angled or piecewise-planar environments.

3. Non-hallucination Knowledge Preservation (NKP): Multi-Band Boundary Alignment and Background Distillation

To prevent the flattening objective from bleeding into non-illusion areas and degrading general representation capabilities, the NKP loss establishes surgical boundary protection. A clean background mask \(m_{\text{bg}} = (1 - m)(1 - r)(1 - r_g)\) excludes the illusion ROI \(m\), an adjacent ring band \(r\), and a protective guard band \(r_g\).

On \(m_{\text{bg}}\), the student's depth and second-order Laplacian magnitudes are tightly bound to the teacher. At the tricky transition boundary, the ring \(r\) is bifurcated into a high-gradient edge subset \(r_e\) (preserving authentic object silhouettes) and a low-gradient seam subset \(r_f\). Student depth is anchored to locally smoothed teacher predictions \(\tilde{z}_T\) on seam \(r_f\), while second-order responses are matched on \(r_e\) and guard ring \(r_g\): $\(\mathcal{L}_{\text{NKP}} = \alpha_3 \overline{|z - z_T|}_{m_{\text{bg}}} + \alpha_4 \overline{|\mathcal{L}(z_T) - \mathcal{L}(z)|}_{m_{\text{bg}}} + \alpha_5 \overline{|z - \tilde{z}_T|}_{r_f} + \alpha_6 \overline{|\mathcal{L}(z) - \mathcal{L}(z_T)|}_{r_e} + \alpha_7 \overline{|\mathcal{L}(z) - \mathcal{L}(z_T)|}_{r_g}\)$ This multi-band regularizer eliminates edge bleeding and boundary halo artifacts, ensuring that corrections remain strictly confined to the illusion carrier.

Loss & Training

The overall training objective combines both views: \(\mathcal{L} = \lambda_C \mathcal{L}_{\text{crop}} + \lambda_F \mathcal{L}_{\text{full}}\), where the crop view carries dominant weight to counteract FOV instability. Each branch integrates \(\mathcal{L}_{\text{HKR}}\), \(\mathcal{L}_{\text{NKP}}\), and gating regularization terms (\(\text{CE}(w, q)\) and an anchor loss \(\min_k \ell_{j,k}\)).

To prevent degenerative over-smoothing on standard scenes, positive illusion batches are interleaved in a 4:1 ratio with non-illusion regularizer data from Penn-Fudan (pedestrian scenes) and CamVid (urban driving frames). For non-illusion samples, \(R = \emptyset\) deactivate HKR, reducing supervision strictly to NKP preservation. The network is optimized using AdamW (learning rate \(1 \times 10^{-4}\), weight decay 0.01) for a single epoch with batch size 8 on an NVIDIA A100 GPU, achieving rapid convergence without runtime inference overhead.

Key Experimental Results

Main Results

Evaluation on the 3D-Mirage test split (47 unseen scenes, 188 paired full-crop instances) demonstrates that all leading monocular depth foundation models suffer from severe 3D mirages (Table 1). Lower values are better.

Model Architecture \(d_{\text{cluster}} \downarrow\) \(d_{\text{avg}} \downarrow\) DCS \(\downarrow\) \(D_{\text{cluster}} \downarrow\) \(D_{\text{avg}} \downarrow\) CCS \(\downarrow\)
DepthPro (ICLR'25) ViT Multi-scale 701.1 726.2 1427.3 \(2.294 \times 10^{-3}\) \(2.402 \times 10^{-3}\) \(4.696 \times 10^{-3}\)
Marigold (CVPR'24) Diffusion (LDM) 317.8 331.4 649.1 \(6.680 \times 10^{-4}\) \(9.290 \times 10^{-4}\) \(1.597 \times 10^{-3}\)
DepthFM (AAAI'25) Flow Matching 1020.0 1063.0 2083.0 \(4.914 \times 10^{-3}\) \(5.215 \times 10^{-3}\) \(1.013 \times 10^{-2}\)
ZoeDepth (arXiv'23) Hybrid Relative/Metric 291.5 297.8 589.3 \(7.486 \times 10^{-4}\) \(7.560 \times 10^{-4}\) \(1.505 \times 10^{-3}\)
MiDaS (TPAMI'22) CNN / ViT Hybrid 330.2 340.0 670.2 \(4.120 \times 10^{-4}\) \(5.090 \times 10^{-4}\) \(9.220 \times 10^{-4}\)
DAv2-S (NeurIPS'24) ViT-Small 495.9 511.2 1007.1 \(7.210 \times 10^{-4}\) \(7.900 \times 10^{-4}\) \(1.512 \times 10^{-3}\)
DAv2-B (NeurIPS'24) ViT-Base 431.7 449.3 881.0 \(6.320 \times 10^{-4}\) \(7.270 \times 10^{-4}\) \(1.359 \times 10^{-3}\)
DAv2-L (Teacher Baseline) ViT-Large 488.8 505.8 994.6 \(6.840 \times 10^{-4}\) \(7.820 \times 10^{-4}\) \(1.466 \times 10^{-3}\)
Ours (GSD) ViT-Large + LoRA 28.55 30.09 58.64 \(9.174 \times 10^{-5}\) \(9.894 \times 10^{-5}\) \(1.907 \times 10^{-4}\)
Relative Improvement \(\Delta\) - -94.16% -94.05% -94.10% -86.59% -87.35% -86.99%

On the external 3D-Visual-Illusion (3DVI, NeurIPS'25) test set, GSD demonstrates outstanding zero-shot transfer in the monocular track. While the DAv2-Metric baseline incurs AbsRel of 0.52 and EPE of 16.24, our adapted model achieves Disparity EPE of 1.75, bad2 of 26.67%, and Depth AbsRel of 0.03, with \(\delta_1\) reaching 99.50%, matching or exceeding many multi-view stereo baselines.

Ablation Study

Ablation experiments evaluate the impact of loss components and adaptation paradigms across 3D-Mirage, NYUv2, KITTI 15, and the pairwise depth benchmark DA-2K (combining Table 4 and Table 5).

Adaptation Setting / Variant 3D-Mirage DCS \(\downarrow\) 3D-Mirage CCS \(\downarrow\) NYUv2 AbsRel \(\downarrow\) KITTI 15 AbsRel \(\downarrow\) DA-2K Pairwise Acc \(\uparrow\) Note
Ours (Full GSD Pipeline) 58.64 \(1.907 \times 10^{-4}\) 0.1597 0.3424 96.08% Best balance of anti-hallucination and retention
w/o Re-editing (w/o \(\mathcal{L}_{\text{HKR}}\)) 988.60 \(1.470 \times 10^{-3}\) 0.1623 0.3409 97.05% Preserves depth but leaves mirage unresolved
w/o Preservation (w/o \(\mathcal{L}_{\text{NKP}}\)) 42.83 \(1.410 \times 10^{-4}\) 0.1533 0.3510 94.00% Over-flattening bleeds into background geometry
Simple Plane L1 Fitting 59.68 \(2.118 \times 10^{-4}\) 0.5299 (RMSE) 7.10 (RMSE) 93.04% Single plane constraint fails on multi-surface folds
Naive Full Encoder Fine-tuning 27.85 \(9.389 \times 10^{-5}\) 0.9014 (RMSE) 8.54 (RMSE) 58.51% Catastrophic forgetting; DA-2K accuracy collapses
Decoder-Only Fine-tuning 70.26 \(1.986 \times 10^{-4}\) 0.6487 (RMSE) 6.93 (RMSE) 87.48% Fails to modulate deep semantic ViT representations

Key Findings

  • Cooperative Necessity of HKR and NKP: Removing re-editing (\(\mathcal{L}_{\text{HKR}}\)) leaves the DCS score near baseline levels (988.60), showing zero hallucination reduction. Conversely, removing preservation (\(\mathcal{L}_{\text{NKP}}\)) achieves a lower DCS (42.83) at the expense of corrupting non-illusion structures and degrading DA-2K accuracy.
  • Vulnerability of Full Fine-tuning: Unconstrained encoder fine-tuning causes devastating catastrophic forgetting: DA-2K pairwise accuracy plummets from 96.08% to 58.51%, and DIW WHDR degrades to 32.15%. Parameter-efficient LoRA constraints are critical to surgically confining modifications.
  • Cross-Backbone Generalization: GSD readily transfers to alternative architectures. Adapting ZoeDepth reduces DCS from 589.27 to 335.12 and cuts EPE from 9.555 to 7.286 on 3DVI; adapting the Marigold diffusion backbone reduces DCS by 44% and CCS by 40% while maintaining crisp foreground object boundaries.

Highlights & Insights

  • Diagnosing the Context-Dependent 3D Mirage: Systematically uncovers the fundamental failure of modern depth models trading geometric fidelity for contextual priors, demonstrating that FOV restriction acts as a primary catalyst for hallucinated 3D obstacles.
  • Reference-Free Curvature and Invariance Metrics: DCS and CCS offer a ground-truth-free diagnostic framework based on second-order derivatives, enabling real-world deployments to audit structural depth integrity without expensive LiDAR scans.
  • Surgical Self-Distillation via Multi-Hypothesis Gating: Instead of imposing rigid single-plane constraints, the framework estimates local surface hypotheses from teacher neighborhood rings and gates them dynamically, resolving multi-surface illusions cleanly without background degradation.

Limitations & Future Work

  • Limited Scope of Optical Traps: The 3D-Mirage benchmark focuses predominantly on textured planar and low-curvature surfaces, leaving transparent surfaces, specular reflections, extreme shadows, and severe adverse weather conditions for future exploration.
  • Absence of Negative Obstacle Evaluations: The study primarily tackles phantom protrusions (positive obstacles); verifying that the smoothing objective does not erase critical negative hazards (such as potholes, road dips, or low curbs) remains a vital direction for autonomous driving safety.
  • vs. Yao et al. (3DVI, NeurIPS 2025): 3DVI requires heavy vision-language models (VLMs) and stereo matching to mitigate visual illusions, introducing prohibitive latency; GSD operates purely on monocular models via 1.2% LoRA adapters, adding zero latency or memory overhead at inference.
  • vs. Conventional Fine-Tuning & Adversarial Defenses: Standard fine-tuning leads to severe catastrophic forgetting of pre-trained depth geometry; GSD circumvents this via frozen-teacher self-distillation and ring-boundary regularization, maintaining high ordinal accuracy across NYUv2 and KITTI.

Rating

  • Novelty: โญโญโญโญโญ [Pioneering formulation and systematic mitigation of the 3D Mirage failure mode in monocular depth foundation models]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluations spanning the new benchmark, zero-shot 3DVI transfer, general depth retention, and multiple architectures]
  • Writing Quality: โญโญโญโญโญ [Exceptional clarity, rigorous mathematical formulations, and compelling motivation-to-solution narrative]
  • Value: โญโญโญโญโญ [Directly resolves a hazardous blind spot in vision foundation models for autonomous driving and robotics]