Skip to content

Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/startnew/flexdepth
Area: Autonomous Driving
Keywords: depth estimation, dynamic scene decoupling, scale-driven decoder, edge deployment, self-supervised learning

TL;DR

To tackle photometric reprojection corruption caused by moving objects and boundary blurring in lightweight models within autonomous driving, FlexDepth establishes a flexible model family scaling from Nano to X-Large. It introduces a pre-convolution-free Scale-Driven Decoder alongside a two-stage static-dynamic decoupled training strategy derived from cross-epoch prediction variance, securing SOTA accuracy and ultra-low edge latency without requiring semantic or optical flow priors.

Background & Motivation

Self-supervised monocular depth estimation recovers dense 3D scene geometry from consecutive video sequences via photometric consistency, bypassing the prohibitive data collection expenses of LiDAR ground truth and becoming a cornerstone of autonomous driving and robotic perception. Nevertheless, prevailing self-supervised paradigms face severe real-world deployment bottlenecks. Heavy architectures grounded in Vision Transformers or diffusion priors incur massive computational overhead and latency, making them impractical for resource-constrained automotive edge chips. Conversely, existing lightweight networks rely on static, single-scale designs and simplistic upsampling operators to suppress FLOPs, which incurs structural fidelity degradation, especially around thin structures and object silhouettes. Crucially, complex driving scenes are populated by dynamic traffic participants whose non-rigid movements inherently violate the multi-view static scene assumption, corrupting the photometric reprojection loss with erroneous gradient backpropagation.

Existing approaches designed to tackle dynamic scenarios suffer from notable drawbacks. Several methods introduce auxiliary supervisory tasks such as semantic segmentation, optical flow fields, or cross-dataset pseudo-depth distillation. These auxiliary components substantially bloat inference pipelines, destroy the clean elegance of end-to-end self-supervision, and prove fragile when confronted with casting shadows, motion blur, and out-of-distribution environments. Other strategies rely on fixed global thresholds to suppress dynamic residuals, failing to generalize across driving scenarios where the proportion and velocity of dynamic agents fluctuate drastically. Concurrently, conventional lightweight encoder-decoder architectures routinely apply pre-convolutions before skip concatenation to compress channel dimensions, which prematurely discards vital global context and high-dimensional depth cues before spatial expansion.

To overcome these challenges, this work unifies architectural scalability and intrinsic dynamic-scene robustness within a self-supervised regime without relying on any external supervision or auxiliary priors. The core idea is to eliminate information-constricting pre-convolutions through an adaptive Scale-Driven Decoder (SDD) and exploit cross-epoch predictive inconsistency between early and late training stages to construct an adaptive static-dynamic decoupled mask for two-stage robust depth retraining.

Method

Overall Architecture

The FlexDepth family encompasses five distinct scales—Nano, Small, Medium, Large, and X-Large—tailored for deployment spanning low-power mobile platforms to high-performance edge compute platforms. The depth estimation network adopts an efficient YOLOv11 backbone as the multi-scale feature encoder, progressively extracting feature hierarchies across five spatial scales (\(1/2\) down to \(1/32\) of the original resolution). The training framework is organized into two stages: Stage 1 jointly optimizes the PoseNet (a ResNet-18 encoder coupled with a four-layer decoder predicting relative 6-DoF camera poses \(P_{t \to s}\)) and the depth prediction network using standardized photometric reprojection error; Stage 2 freezes the converged PoseNet and constructs a pixel-wise adaptive Static-Dynamic Decoupled Mask \(M\) based on disparity fluctuation between initial and final Stage-1 weights, subsequently retraining the depth network under a masked uncertainty distribution loss and geometric depth smoothing to eliminate moving-object corruption.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Monocular Driving Video Frames<br/>I_t, I_s"] --> B["Stage 1: Joint Geometric Pre-training<br/>Co-optimizing PoseNet and YOLOv11 Depth Encoder"]
    B --> C["Cross-Epoch Prediction Inconsistency Sensing<br/>Discrepancy between disp_i and disp_l"]
    C --> D["Adaptive Static-Dynamic Decoupled Mask M<br/>Dynamic thresholding isolates moving foreground"]
    D --> E["Stage 2: Robust Depth Retraining<br/>Frozen PoseNet + Masked Normal Distribution Loss L_m"]
    E --> F["Scale-Driven Decoder SDD<br/>Adaptive assembly of HEB/HPB and dynamic resampling"]
    F --> G["High-Fidelity Metric Depth Output<br/>Sharp structural contours and robust dynamic depth"]

Key Designs

1. Scale-Driven Decoder (SDD): Reconstructing Lightweight Upsampling and Feature Fusion

Traditional lightweight depth decoders follow a rigid sequence of "pre-convolution \(\to\) bilinear upsampling \(\to\) skip concatenation \(\to\) post-convolution". Applying convolution filters before scale recovery forces the compression of sparse semantic information, irreversibly discarding critical cues and causing contour blur and small-object misses. SDD eliminates the pre-convolution module entirely, feeding high-dimensional features directly into upsampling and concatenation: $\(D_i = \mathcal{P}\{\mathcal{C}(\text{Concat}[\mathcal{U}(F_{i+1}), F_{\text{enc},i}])\}\)$ Within this formulation, components adaptively match model capacities: - Post-Convolution Unit \(\mathcal{C}\): Flex-Nano/Small employ High-Efficiency Bottleneck (HEB) units. A uniform split operator \(S\) divides the feature into \(f_a\) and \(f_b\), routing features through dense residual bottleneck chains whose outputs concatenate directly with initial slices. This shortens gradient propagation paths and stabilizes training for compact models. Flex-Medium/Large/X-Large switch to High-Performance Bottlenecks (HPB), integrating Cross Stage Partial (CSP) blocks with nested convolutions to capture subtle local structural boundaries without incurring notable latency. - Dynamic Resampling Operator \(\mathcal{U}\): Rather than standard bilinear interpolation—which acts as an isotropic low-pass filter that washes out high-frequency contours and causes halo artifacts—SDD predicts dynamic 2D coordinate offsets \(\Delta P\) via convolution. After spatial reorganization through Pixel Shuffle (\(\text{PS}\)), offsets combine with a uniform grid \(C\) for non-uniform grid sampling: $\(\mathcal{U}(X) = \text{GS}\left(X, \text{PS}\left(\frac{2}{\text{Norm}} \cdot (C + \Delta P) - 1, s\right)\right)\)$ This enables content-dependent sampling aligned with object silhouettes, recovering crisp object boundaries. - Inverted Prediction Head \(\mathcal{P}\): For larger model variants, the prediction sequence is reversed: upsampling precedes depth prediction. Because high-channel feature maps preserve semantic and structural consistency during upsampling, subsequent prediction heads produce sharper depth maps. Smaller variants retain conventional ordering to guarantee real-time mobile throughput.

2. Cross-Epoch Inconsistency Sensing and Adaptive Static-Dynamic Decoupled Mask

Under self-supervised multi-view geometric supervision, moving traffic participants violate photometric consistency, causing their backpropagated error gradients to oscillate continuously throughout training. Consequently, dynamic objects exhibit severe prediction discrepancies between early and late training epochs, whereas static backgrounds converge monotonically. Rather than adopting fragile global thresholds, FlexDepth extracts disparity predictions from the initial epoch (\(\text{disp}_i\)) and final epoch (\(\text{disp}_l\)) of Stage 1 to form a pixel-adaptive decoupling mask: $\(M = \left[ |\text{disp}_i^{\beta} - \text{disp}_l^{\beta}| < g(\phi(f_{si}, f_{sl})) \right]\)$ Here \(\beta = 0.7\) scales foreground-background dynamic sensitivity, and \([\cdot]\) denotes the Iverson bracket. The function \(\phi\) quantifies structural representation shifts across feature spaces, which function \(g\) maps into an adaptive pixel-level threshold field. In regions exhibiting rapid motion or complex textures, the threshold dynamically contracts to strictly isolate moving objects, while static roadway and architectural structures remain unmasked, achieving zero-cost unsupervised dynamic entity separation without semantic or flow priors.

3. Static-Dynamic Decoupled Retraining and Uncertainty-Guided Smoothing

Once the binary static mask \(M\) is computed, Stage 2 freezes PoseNet to prevent contaminated motion gradients from degrading ego-motion estimations. To handle residual photometric errors near object boundaries and shadow transitions, an auxiliary variance network predicts pixel-wise uncertainty \(\sigma\), formulating the masked normal distribution objective \(L_m\): $\(L_m = M \cdot \left[ \log(\sigma + 1) + \frac{\lambda L_r^2}{2\sigma^2} \right]\)$ The photometric loss \(L_r\) incorporates structural similarity (SSIM, \(\alpha=0.85\)) and \(L_1\) color difference under per-pixel minimum reprojection with \(\lambda=0.5\). To regularize dynamic regions (\(1-M\)) where photometric loss is excluded, an anisotropic Geometric Depth Smoothing (\(L_{\text{GDS}}\)) term enforces structural continuity: $\(L_{\text{GDS}} = |\partial_x \hat{d}_t| e^{-|\partial_x I_t|} + [\eta(1 - M) + M] |\partial_y \hat{d}_t| e^{-|\partial_y I_t|}\)$ By scaling smoothing penalties on dynamic regions with \(\eta = 100\), the network propagates smooth geometric manifolds from adjacent static ground planes and backgrounds, eliminating floating artifacts and holes on moving vehicles.

Key Experimental Results

Main Results

Quantitative evaluations are conducted on the standard KITTI benchmark (Eigen split, \(640 \times 192\)) and Cityscapes dynamic urban driving benchmark (\(416 \times 128\)). All models are trained purely on monocular video sequences (M) without ground-truth depth or auxiliary annotations.

Dataset Method Category / Scale FLOPs (G) ↓ Abs Rel ↓ Sq Rel ↓ RMSE ↓ \(\delta < 1.25\) ↑
KITTI Monodepth2 Baseline 8.0 0.115 0.903 4.863 0.877
KITTI Lite-Mono-Tiny Lightweight 2.8 0.110 0.837 4.710 0.880
KITTI PuriLight Lightweight 7.1 0.106 0.747 4.536 0.890
KITTI Flex-Nano (Ours) Lightweight 0.7 0.110 0.794 4.678 0.878
KITTI Flex-Small (Ours) Lightweight 2.8 0.104 0.713 4.458 0.890
KITTI TinyDepth Medium 13.3 0.096 0.665 4.249 0.904
KITTI Flex-Medium (Ours) Medium 10.0 0.096 0.639 4.253 0.903
KITTI Flex-Large (Ours) Large 11.5 0.095 0.642 4.199 0.906
KITTI MonoViT Heavyweight 59.7 0.099 0.708 4.372 0.900
KITTI DSI-MonoViT Dynamic SOTA 59.7 0.096 0.711 4.321 0.905
KITTI Flex-X-Large (Ours) Large 24.6 0.093 0.605 4.114 0.910
Cityscapes ManyDepth Multi-frame 12.1 0.114 1.193 6.223 0.875
Cityscapes DynamicDepth M + Semantic 18.5+ 0.103 1.000 5.867 0.895
Cityscapes PuriLight Single-frame 5.7 0.099 1.014 5.892 0.901
Cityscapes Flex-Nano (Ours) Single-frame 0.6 0.107 1.261 6.133 0.893
Cityscapes Flex-Small (Ours) Single-frame 2.2 0.100 1.078 5.813 0.904
Cityscapes ProDepth Multi-frame 35.5 0.095 0.876 5.531 0.908
Cityscapes DSI-MonoViT Single-frame 47.7 0.088 0.872 5.410 0.918
Cityscapes Flex-X-Large (Ours) Single-frame 19.6 0.086 0.877 5.268 0.926

Ablation Study

Component-wise ablation experiments on KITTI using the Flex-X-Large architecture systematically evaluate post-convolution variations (HPB vs. HEB), upsampling strategies (Dynamic Resampling vs. Bilinear), prediction head topologies (Inverted vs. Conventional), and second-stage dynamic handling options:

ID Post-Conv \(\mathcal{C}\) Upsampling \(\mathcal{U}\) Head \(\mathcal{P}\) Dynamic Mask FLOPs (G) Params (M) Abs Rel ↓ Sq Rel ↓ RMSE ↓ \(\delta < 1.25\) ↑ Note
1 HEB (◦) Dynamic (✓) Inverted (✓) Mask \(M\) (✓) 24.0 32.0 0.093 0.614 4.188 0.907 HEB slightly drops compute but slightly lowers accuracy
2 HPB (✓) Bilinear (◦) Inverted (✓) Mask \(M\) (✓) 24.3 32.3 0.094 0.627 4.204 0.906 Bilinear filtering suppresses high-frequency edge details
3 HPB (✓) Dynamic (✓) Conventional (◦) Mask \(M\) (✓) 24.3 32.3 0.094 0.631 4.205 0.906 Conventional order harms high-channel geometry semantics
4 HPB (✓) Dynamic (✓) Inverted (✓) W/o Stage 2 (×) 24.6 32.3 0.100 0.729 4.431 0.900 Severe performance collapse without dynamic decoupling
5 HPB (✓) Dynamic (✓) Inverted (✓) Fixed Thresh (◦) 24.6 32.3 0.093 0.629 4.195 0.909 Fixed threshold lacks cross-scene environmental adaptability
6 HPB (✓) Dynamic (✓) Inverted (✓) Mask \(M\) (✓) 24.6 32.3 0.093 0.605 4.114 0.910 Full model setting achieving optimal performance

Key Findings

  • Crucial Impact of Two-Stage Decoupled Training: Removing Stage 2 dynamic retraining (ID 4 vs. ID 6) triggers a sharp 7.5% degradation in Abs Rel (0.093 to 0.100) and inflates RMSE from 4.114 to 4.431, confirming that moving-object photometric violations severely corrupt gradient updates. Furthermore, the adaptive thresholding mask \(M\) surpasses fixed thresholds (Sq Rel 0.605 vs. 0.629), confirming the necessity of scene-adaptive thresholds.
  • Superiority in Dynamic Region Isolation: In decoupled evaluations on Cityscapes, Flex-X-Large achieves an Abs Rel of 0.091 and \(\delta < 1.25\) of 0.935 in isolated dynamic regions, outperforming Mining-BrNet (Abs Rel 0.119, \(\delta\) 0.872) and ProDepth (Abs Rel 0.115, \(\delta\) 0.884)—which depend on pixel motion or multi-frame inputs—while using only \(55\%\) to \(72\%\) of their computational budget.
  • Unprecedented Edge Inference Throughput: On the Snapdragon 8 Elite mobile processor, Flex-Nano (0.718 GFLOPs) achieves 37.6 FPS real-time execution, whereas Lite-Mono-Tiny (2.842 GFLOPs) only reaches 15.9 FPS. On an NVIDIA RTX 4090 GPU, Flex-Nano achieves an extraordinary 1583 FPS, verifying its exceptional viability for resource-constrained automotive chips.
  • Outperforming Heavy Foundation Models In-Domain: Under an aligned least-squares alignment protocol on KITTI against the 335M-parameter foundation model Depth Anything V2 (ViT-L), the in-domain self-supervised Flex-X-Large (32M parameters, 25 GFLOPs) secures superior metric accuracy (Abs Rel 0.063 vs. 0.070), highlighting the compelling efficiency and precision of task-specialized self-supervision over bloated foundation architectures.

Highlights & Insights

  • Cross-Epoch Convergence Inconsistency for Unsupervised Motion Discovery: Leveraging the observation that moving objects oscillate and fail to converge under photometric losses provides a principled, zero-cost mechanism to discover dynamic regions without relying on off-the-shelf optical flow or instance segmentation models.
  • Scale-Driven Architectural Customization: Deviating from traditional uniform scaling, configuring distinct post-convolution blocks (HEB vs. HPB), prediction head sequences, and dynamic resampling across model capacities yields Pareto-optimal accuracy-throughput trade-offs across all deployment tiers.
  • Elimination of Pre-Convolution Bottlenecks: Revealing that premature \(1 \times 1\) pre-convolutions choke global context before spatial upsampling provides an effective architectural blueprint for lightweight dense prediction decoders.

Limitations & Future Work

  • Performance Drop in Unstructured Indoor Environments: FlexDepth relies on depth priors and road geometry common to driving perspectives; when tested in complex indoor scenes characterized by chaotic textures, non-planar ground, and extreme near-field occlusions, its geometric smoothness regularizer can induce over-smoothing.
  • Photometric Breakdown Under Extreme Illumination: Being purely self-supervised, the pipeline relies on photometric consistency; sudden lighting shifts, nocturnal headlight glare, and torrential rain can degrade Stage-1 convergence, impacting the quality of Stage-2 temporal masks.
  • Future Directions: Integrating lightweight transient neural representation fields or 3D Gaussian Splatting dynamics into the two-stage framework represents a promising path to refine ego-motion and dynamic scene geometry jointly.
  • vs. Monodepth2 & Lite-Mono: Monodepth2’s static auto-masking only filters vehicles moving at identical velocity to the camera, completely failing on orthogonal traffic. Lite-Mono improves encoder efficiency via cross-channel attention but retains rigid, lossy upsampling decoders. FlexDepth resolves both issues via content-aware dynamic resampling and cross-epoch motion decoupling.
  • vs. DSI-MonoViT: While DSI pioneered depth self-inhibition, it relies on static empirical thresholds and a heavy 60 GFLOPs Vision Transformer. FlexDepth upgrades this to an adaptive feature-conditioned thresholding function within a lightweight scale-driven architecture, reducing compute by more than half while surpassing dynamic-region accuracy on Cityscapes.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Exploiting training convergence variance across epochs for dynamic decoupling is elegant and self-contained; the Scale-Driven Decoder design is well-rationalized.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Extensive multi-dataset verification (KITTI, Cityscapes dynamic split, Make3D zero-shot), rigorous ablation, and real-world mobile/GPU runtime benchmarking.
  • Writing Quality: ⭐⭐⭐⭐⭐ Well-structured narrative with rigorous mathematical derivations and clear visual comparisons.
  • Value: ⭐⭐⭐⭐⭐ Delivering 0.7 GFLOPs / 37.6 FPS edge perception alongside 0.093 Abs Rel SOTA performance offers direct value for automotive production systems.