The 3D Mirage: Probing and Taming 3D Hallucinations¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/hdnndh/The-3D-Mirage-Probing-and-Taming-3D-Hallucinations
Area: 3D Vision / Hallucination Detection
Keywords: monocular depth estimation, 3D hallucination, contextual stability, self-distillation, parameter-efficient fine-tuning
TL;DR¶
Addressing the systemic vulnerability where monocular depth foundation models hallucinate prominent non-existent 3D obstacles on planar optical illusions when surrounding context is restricted, this paper introduces the 3D-Mirage benchmark, reference-free second-order metrics (DCS and CCS), and a parameter-efficient Grounded Self-Distillation framework that cuts geometric hallucination by over 94% without catastrophic forgetting.
Background & Motivation¶
Monocular depth estimation (MDE) foundation modelsโsuch as Depth Anything V2, ZoeDepth, and Depth Proโhave achieved remarkable zero-shot generalization across diverse visual scenes. To infer depth from a single image, an inherently ill-posed problem where projective geometry discards metric scale, these models rely heavily on large-scale statistical and semantic priors learned from diverse datasets. However, this impressive capability conceals an unexamined safety-critical flaw: the network frequently trades strict local geometric fidelity for semantic consistency, substituting brittle visual priors for physical evidence.
This reliance manifests as a dramatic failure mode when encountering real-world visual illusions (e.g., 3D street chalk art, forced-perspective murals, and deceptive billboards). On low-curvature or flat carrier surfaces, models perceive phantom 3D bumps, hollows, or obstacles. More critically, this vulnerability is exacerbated under context variations: while a global view might provide sufficient surrounding context to infer a flat roadway, restricting the field-of-view (FOV) via cropping or localized occlusion removes broader contextual anchors, triggering severe hallucinated geometry. Existing benchmarks (KITTI, ScanNet, NYUv2) lack such adversarial optical configurations, and standard pixel-averaged metrics like MAE, RMSE, and AbsRel cannot isolate localized structural distortions or quantify cross-view instability without ground-truth depth.
To bridge this diagnostic and methodological gap, a holistic probing, scoring, and taming paradigm is required. Core idea: Introduce the 3D-Mirage paired-view benchmark along with reference-free second-order metrics (DCS and CCS), and develop Grounded Self-Distillation (GSD)โa LoRA-based dual-branch framework that suppresses spurious ROI curvature via multi-surface plane-mixture fitting while preserving frozen teacher geometry on background regions to eliminate 3D hallucinations without catastrophic forgetting.
Method¶
Overall Architecture¶
The framework encompasses three coordinated components: the 3D-Mirage benchmark designed to induce contextual instability, reference-free second-order evaluation metrics to quantify structural hallucination, and a parameter-efficient Grounded Self-Distillation (GSD) mitigation pipeline.
During GSD training, each instance consists of a full image \(x_{\text{full}}\) and a context-restricted crop \(x_{\text{crop}}\) centered on the illusion region. A frozen Depth-Anything-V2-Large acts as the teacher network \(T\), providing stable baseline depth representations. The student network \(S\) introduces low-rank adaptation (LoRA) modules into the patch embedding and MLP layers of the Vision Transformer (ViT) encoder (comprising only ~4M parameters, ~1.2% of the backbone), while keeping the decoder frozen. The student processes both views through shared weights. The Hallucination Knowledge Re-editing (HKR) loss penalizes second-order curvature within annotated illusion ROIs and aligns predictions with local surface hypotheses fitted from teacher ring statistics. Concurrently, the Non-hallucination Knowledge Preservation (NKP) loss distills the teacher's stable geometry on background pixels and edge boundaries, preserving general depth estimation capabilities.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Paired Views<br/>(Full Image x_full & Restricted Crop x_crop)"] --> B["Dual-Branch Feature Propagation<br/>(Frozen Teacher T & LoRA-adapted Student S)"]
B --> C["Dual-View Projection & Second-Order Probing<br/>(DCS Structural Deviation + CCS Contextual Confusion)"]
C --> D["Local Surface Fitting & Gating Mixture<br/>(Ring-Fitted Hypotheses ฯ_k + Gating Weights w_j)"]
D --> E["Hallucination Knowledge Re-editing HKR<br/>(Curvature Penalty + Multi-Hypothesis Alignment)"]
D --> F["Non-hallucination Knowledge Preservation NKP<br/>(Background Distillation + Seam/Edge Boundary Regularizer)"]
E --> G["Combined Objective Optimization<br/>(Parameter-Efficient Update + Regularizer Interleaving)"]
F --> G
Key Designs¶
1. 3D-Mirage Benchmark and Second-Order Reference-Free Probing Metrics
To expose contextual instability on deceptive surfaces without relying on unobtainable dense 3D ground truth, the 3D-Mirage benchmark curates 468 real-world scenes containing 3D paintings, anamorphic street art, and murals. Precise polygonal region-of-interest (ROI) masks delineate the low-curvature support planes, incorporating nested exclusion masks to isolate authentic objects situated atop the illusion. For each scene, four random crops retaining at least 40% of the ROI diagonal simulate restricted fields-of-view, producing 1,872 full-crop evaluation pairs.
For a depth model \(f_\theta\), predictions \(D_{\text{full}}\) and \(D_{\text{crop}}\) are computed. The full-view prediction is aligned to the crop coordinate frame (\(\tilde{D}_{\text{full}\rightarrow\text{crop}}\)). Applying quantile normalization \(\mathrm{QN}(\cdot; M_i)\) and a separable second-difference magnitude operator \(\mathcal{L}(\cdot)\) produces normalized second-order response maps \(R_i^{\text{full}}\) and \(R_i^{\text{crop}}\). Two summary scalars are extracted per branch: the top-10% response sum \(t_{v,i}\) and the 10%-trimmed mean \(\mu_{v,i}\). Two complementary metrics are formulated: 1. Deviation Composite Score (DCS): Measures the absolute hallucination intensity by capturing spurious second-order surface curvature: $\(\mathrm{DCS} = \|\bar{\mathbf{t}}\|_2 + \frac{1}{N}\sum_{i=1}^N \|\mathbf{t}_i\|_2\)$ where \(\mathbf{t}_i = [t_{\text{full},i}, t_{\text{crop},i}]^\top\). Lower DCS indicates that the low-curvature carrier is correctly reconstructed without artificial geometric protrusions. 2. Confusion Composite Score (CCS): Quantifies contextual instability across varying FOVs by projecting mean responses along the off-diagonal unit vector \(\mathbf{u} = \frac{1}{\sqrt{2}}[1, -1]^\top\): $\(\mathrm{CCS} = |\mathbf{u}^\top \bar{\boldsymbol{\mu}}| + \frac{1}{N}\sum_{i=1}^N |\mathbf{u}^\top \boldsymbol{\mu}_i|\)$ A low CCS guarantees that the model's structural prediction remains robust and invariant to the presence or absence of peripheral scene context.
2. Hallucination Knowledge Re-editing (HKR): Curvature Suppression and Gated Hypothesis Mixture
Enforcing a naive global flat-plane constraint fails when an illusion spans multi-faceted support surfaces, such as perpendicular walls or road-curb transitions. HKR resolves this by combining explicit curvature penalties with an adaptive mixture of local planar hypotheses.
For each illusion polygon \(j \in \mathcal{R}\), up to \(K\) simple geometric plane hypotheses \(\{\pi_{j,k}\}_{k=1}^K\) are fitted to the frozen teacher's depth within an adjacent dilated ring band, with corresponding residual scales \(\sigma_{j,k}\). A compact gating MLP maps ROI and ring visual statistics to mixture logits \(w_j = \text{softmax}(G(\cdot))\), supervised via soft cross-entropy against residual targets \(q_j\) derived from \(\sigma_{j,k}\). The deviation between student normalized depth \(z\) and each fitted plane \(\pi_{j,k}\) is computed as \(\ell_{j,k} = \overline{|z - \pi_{j,k}|}_{m_j}\). The total HKR loss penalizes the union ROI curvature while pulling instances toward plausible surface mixtures: $\(\mathcal{L}_{\text{HKR}} = \alpha \overline{|\mathcal{L}(s)|}_m + \sum_{j \in \mathcal{R}} \sum_{k=1}^K w_{j,k} \ell_{j,k}\)$ Computed across both full and cropped views, this formulation flattens spurious 3D curvature while accommodating naturally angled or piecewise-planar environments.
3. Non-hallucination Knowledge Preservation (NKP): Multi-Band Boundary Alignment and Background Distillation
To prevent the flattening objective from bleeding into non-illusion areas and degrading general representation capabilities, the NKP loss establishes surgical boundary protection. A clean background mask \(m_{\text{bg}} = (1 - m)(1 - r)(1 - r_g)\) excludes the illusion ROI \(m\), an adjacent ring band \(r\), and a protective guard band \(r_g\).
On \(m_{\text{bg}}\), the student's depth and second-order Laplacian magnitudes are tightly bound to the teacher. At the tricky transition boundary, the ring \(r\) is bifurcated into a high-gradient edge subset \(r_e\) (preserving authentic object silhouettes) and a low-gradient seam subset \(r_f\). Student depth is anchored to locally smoothed teacher predictions \(\tilde{z}_T\) on seam \(r_f\), while second-order responses are matched on \(r_e\) and guard ring \(r_g\): $\(\mathcal{L}_{\text{NKP}} = \alpha_3 \overline{|z - z_T|}_{m_{\text{bg}}} + \alpha_4 \overline{|\mathcal{L}(z_T) - \mathcal{L}(z)|}_{m_{\text{bg}}} + \alpha_5 \overline{|z - \tilde{z}_T|}_{r_f} + \alpha_6 \overline{|\mathcal{L}(z) - \mathcal{L}(z_T)|}_{r_e} + \alpha_7 \overline{|\mathcal{L}(z) - \mathcal{L}(z_T)|}_{r_g}\)$ This multi-band regularizer eliminates edge bleeding and boundary halo artifacts, ensuring that corrections remain strictly confined to the illusion carrier.
Loss & Training¶
The overall training objective combines both views: \(\mathcal{L} = \lambda_C \mathcal{L}_{\text{crop}} + \lambda_F \mathcal{L}_{\text{full}}\), where the crop view carries dominant weight to counteract FOV instability. Each branch integrates \(\mathcal{L}_{\text{HKR}}\), \(\mathcal{L}_{\text{NKP}}\), and gating regularization terms (\(\text{CE}(w, q)\) and an anchor loss \(\min_k \ell_{j,k}\)).
To prevent degenerative over-smoothing on standard scenes, positive illusion batches are interleaved in a 4:1 ratio with non-illusion regularizer data from Penn-Fudan (pedestrian scenes) and CamVid (urban driving frames). For non-illusion samples, \(R = \emptyset\) deactivate HKR, reducing supervision strictly to NKP preservation. The network is optimized using AdamW (learning rate \(1 \times 10^{-4}\), weight decay 0.01) for a single epoch with batch size 8 on an NVIDIA A100 GPU, achieving rapid convergence without runtime inference overhead.
Key Experimental Results¶
Main Results¶
Evaluation on the 3D-Mirage test split (47 unseen scenes, 188 paired full-crop instances) demonstrates that all leading monocular depth foundation models suffer from severe 3D mirages (Table 1). Lower values are better.
| Model | Architecture | \(d_{\text{cluster}} \downarrow\) | \(d_{\text{avg}} \downarrow\) | DCS \(\downarrow\) | \(D_{\text{cluster}} \downarrow\) | \(D_{\text{avg}} \downarrow\) | CCS \(\downarrow\) |
|---|---|---|---|---|---|---|---|
| DepthPro (ICLR'25) | ViT Multi-scale | 701.1 | 726.2 | 1427.3 | \(2.294 \times 10^{-3}\) | \(2.402 \times 10^{-3}\) | \(4.696 \times 10^{-3}\) |
| Marigold (CVPR'24) | Diffusion (LDM) | 317.8 | 331.4 | 649.1 | \(6.680 \times 10^{-4}\) | \(9.290 \times 10^{-4}\) | \(1.597 \times 10^{-3}\) |
| DepthFM (AAAI'25) | Flow Matching | 1020.0 | 1063.0 | 2083.0 | \(4.914 \times 10^{-3}\) | \(5.215 \times 10^{-3}\) | \(1.013 \times 10^{-2}\) |
| ZoeDepth (arXiv'23) | Hybrid Relative/Metric | 291.5 | 297.8 | 589.3 | \(7.486 \times 10^{-4}\) | \(7.560 \times 10^{-4}\) | \(1.505 \times 10^{-3}\) |
| MiDaS (TPAMI'22) | CNN / ViT Hybrid | 330.2 | 340.0 | 670.2 | \(4.120 \times 10^{-4}\) | \(5.090 \times 10^{-4}\) | \(9.220 \times 10^{-4}\) |
| DAv2-S (NeurIPS'24) | ViT-Small | 495.9 | 511.2 | 1007.1 | \(7.210 \times 10^{-4}\) | \(7.900 \times 10^{-4}\) | \(1.512 \times 10^{-3}\) |
| DAv2-B (NeurIPS'24) | ViT-Base | 431.7 | 449.3 | 881.0 | \(6.320 \times 10^{-4}\) | \(7.270 \times 10^{-4}\) | \(1.359 \times 10^{-3}\) |
| DAv2-L (Teacher Baseline) | ViT-Large | 488.8 | 505.8 | 994.6 | \(6.840 \times 10^{-4}\) | \(7.820 \times 10^{-4}\) | \(1.466 \times 10^{-3}\) |
| Ours (GSD) | ViT-Large + LoRA | 28.55 | 30.09 | 58.64 | \(9.174 \times 10^{-5}\) | \(9.894 \times 10^{-5}\) | \(1.907 \times 10^{-4}\) |
| Relative Improvement \(\Delta\) | - | -94.16% | -94.05% | -94.10% | -86.59% | -87.35% | -86.99% |
On the external 3D-Visual-Illusion (3DVI, NeurIPS'25) test set, GSD demonstrates outstanding zero-shot transfer in the monocular track. While the DAv2-Metric baseline incurs AbsRel of 0.52 and EPE of 16.24, our adapted model achieves Disparity EPE of 1.75, bad2 of 26.67%, and Depth AbsRel of 0.03, with \(\delta_1\) reaching 99.50%, matching or exceeding many multi-view stereo baselines.
Ablation Study¶
Ablation experiments evaluate the impact of loss components and adaptation paradigms across 3D-Mirage, NYUv2, KITTI 15, and the pairwise depth benchmark DA-2K (combining Table 4 and Table 5).
| Adaptation Setting / Variant | 3D-Mirage DCS \(\downarrow\) | 3D-Mirage CCS \(\downarrow\) | NYUv2 AbsRel \(\downarrow\) | KITTI 15 AbsRel \(\downarrow\) | DA-2K Pairwise Acc \(\uparrow\) | Note |
|---|---|---|---|---|---|---|
| Ours (Full GSD Pipeline) | 58.64 | \(1.907 \times 10^{-4}\) | 0.1597 | 0.3424 | 96.08% | Best balance of anti-hallucination and retention |
| w/o Re-editing (w/o \(\mathcal{L}_{\text{HKR}}\)) | 988.60 | \(1.470 \times 10^{-3}\) | 0.1623 | 0.3409 | 97.05% | Preserves depth but leaves mirage unresolved |
| w/o Preservation (w/o \(\mathcal{L}_{\text{NKP}}\)) | 42.83 | \(1.410 \times 10^{-4}\) | 0.1533 | 0.3510 | 94.00% | Over-flattening bleeds into background geometry |
| Simple Plane L1 Fitting | 59.68 | \(2.118 \times 10^{-4}\) | 0.5299 (RMSE) | 7.10 (RMSE) | 93.04% | Single plane constraint fails on multi-surface folds |
| Naive Full Encoder Fine-tuning | 27.85 | \(9.389 \times 10^{-5}\) | 0.9014 (RMSE) | 8.54 (RMSE) | 58.51% | Catastrophic forgetting; DA-2K accuracy collapses |
| Decoder-Only Fine-tuning | 70.26 | \(1.986 \times 10^{-4}\) | 0.6487 (RMSE) | 6.93 (RMSE) | 87.48% | Fails to modulate deep semantic ViT representations |
Key Findings¶
- Cooperative Necessity of HKR and NKP: Removing re-editing (\(\mathcal{L}_{\text{HKR}}\)) leaves the DCS score near baseline levels (988.60), showing zero hallucination reduction. Conversely, removing preservation (\(\mathcal{L}_{\text{NKP}}\)) achieves a lower DCS (42.83) at the expense of corrupting non-illusion structures and degrading DA-2K accuracy.
- Vulnerability of Full Fine-tuning: Unconstrained encoder fine-tuning causes devastating catastrophic forgetting: DA-2K pairwise accuracy plummets from 96.08% to 58.51%, and DIW WHDR degrades to 32.15%. Parameter-efficient LoRA constraints are critical to surgically confining modifications.
- Cross-Backbone Generalization: GSD readily transfers to alternative architectures. Adapting ZoeDepth reduces DCS from 589.27 to 335.12 and cuts EPE from 9.555 to 7.286 on 3DVI; adapting the Marigold diffusion backbone reduces DCS by 44% and CCS by 40% while maintaining crisp foreground object boundaries.
Highlights & Insights¶
- Diagnosing the Context-Dependent 3D Mirage: Systematically uncovers the fundamental failure of modern depth models trading geometric fidelity for contextual priors, demonstrating that FOV restriction acts as a primary catalyst for hallucinated 3D obstacles.
- Reference-Free Curvature and Invariance Metrics: DCS and CCS offer a ground-truth-free diagnostic framework based on second-order derivatives, enabling real-world deployments to audit structural depth integrity without expensive LiDAR scans.
- Surgical Self-Distillation via Multi-Hypothesis Gating: Instead of imposing rigid single-plane constraints, the framework estimates local surface hypotheses from teacher neighborhood rings and gates them dynamically, resolving multi-surface illusions cleanly without background degradation.
Limitations & Future Work¶
- Limited Scope of Optical Traps: The 3D-Mirage benchmark focuses predominantly on textured planar and low-curvature surfaces, leaving transparent surfaces, specular reflections, extreme shadows, and severe adverse weather conditions for future exploration.
- Absence of Negative Obstacle Evaluations: The study primarily tackles phantom protrusions (positive obstacles); verifying that the smoothing objective does not erase critical negative hazards (such as potholes, road dips, or low curbs) remains a vital direction for autonomous driving safety.
Related Work & Insights¶
- vs. Yao et al. (3DVI, NeurIPS 2025): 3DVI requires heavy vision-language models (VLMs) and stereo matching to mitigate visual illusions, introducing prohibitive latency; GSD operates purely on monocular models via 1.2% LoRA adapters, adding zero latency or memory overhead at inference.
- vs. Conventional Fine-Tuning & Adversarial Defenses: Standard fine-tuning leads to severe catastrophic forgetting of pre-trained depth geometry; GSD circumvents this via frozen-teacher self-distillation and ring-boundary regularization, maintaining high ordinal accuracy across NYUv2 and KITTI.
Rating¶
- Novelty: โญโญโญโญโญ [Pioneering formulation and systematic mitigation of the 3D Mirage failure mode in monocular depth foundation models]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluations spanning the new benchmark, zero-shot 3DVI transfer, general depth retention, and multiple architectures]
- Writing Quality: โญโญโญโญโญ [Exceptional clarity, rigorous mathematical formulations, and compelling motivation-to-solution narrative]
- Value: โญโญโญโญโญ [Directly resolves a hazardous blind spot in vision foundation models for autonomous driving and robotics]