Unified Video Dense Prediction from Disjoint Data¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://unid-video.github.io
Area: Segmentation
Keywords: unified dense prediction, video dense prediction, latent distillation, diffusion priors, temporal consistency
TL;DR¶
To tackle the fragmentation of dense prediction datasets across incompatible domains and the acute scarcity of multi-task video annotations, UniD leverages pretrained diffusion visual priors to train per-task specialists on disjoint datasets, distilling them into a unified video backbone via lightweight latent projectors in a single forward pass without co-annotation or pseudo-labeling.
Background & Motivation¶
Holistic scene understanding requires vision systems to simultaneously infer 3D geometry (depth, surface normals), physical appearance (albedo, shading, material segmentation), and high-level scene semantics (semantic segmentation, human part parsing, instance boundaries). For embodied agents navigating dynamic real-world environments, these predictions must be delivered continuously and coherently over streaming video. However, existing dense annotations are deeply fragmented across disparate domains: geometric labels originate largely from indoor RGB-D sensors or synthetic scans with limited semantic diversity, semantic segmentation relies on manual pixel annotations with rich categories but no paired geometry, and intrinsic decompositions are virtually impossible to capture in the wild, remaining confined to synthetic datasets.
This data divergence creates a severe dilemma for unified perception: existing unified models either restrict their scope to narrow co-annotated subsets or incur enormous computational overhead to generate pseudo-labels over large-scale unlabeled corpora. When scaling to video domains, obtaining dense annotations across all eight tasks becomes prohibitively expensive. Furthermore, whenever a new task is introduced, pseudo-labeling pipelines must re-annotate the entire dataset and retrain from scratch, precluding modular extensibility.
The key insight of this paper is that diffusion models pretrained on internet-scale images possess rich generative priors that naturally bridge cross-domain discrepancies while exhibiting intrinsic correspondence across frames. Core idea: train per-task specialists on their own disjoint, domain-specific datasets using pretrained diffusion priors, and then distill their latent representations into a shared temporal backbone via lightweight task projectors in latent space, eliminating the need for co-annotated data or pseudo-labeling while transferring temporal consistency across tasks.
Method¶
Overall Architecture¶
UniD formulates unified video dense prediction as follows: given an input video stream \(\mathbf{v} = \{\mathbf{x}_1, \dots, \mathbf{x}_T\}\) with \(\mathbf{x}_i \in \mathbb{R}^{H \times W \times 3}\), a shared backbone network executes a single forward pass to simultaneously produce dense prediction maps \(\{\mathbf{y}_i^k\}_{k=1}^K\) across \(K=8\) tasks, where \(\mathbf{y}_i^k \in \mathbb{R}^{H \times W \times C_k}\). The architecture comprises three core stages: first, per-task specialists are trained on their respective disjoint datasets using a single-step latent diffusion model (LDM); second, a unified backbone \(\mathbf{G}_\phi\) with Extended Self-Attention (ESA) temporal memory is trained to reconstruct all task latents via lightweight latent projectors \(\Psi^k\) under latent \(\ell_1\) reconstruction and temporal gradient matching; third, during streaming inference, the unified representation is decoded through task projectors, the frozen VAE decoder prefix, and pixel projectors to output all predictions simultaneously.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400, 'subGraphTitleMargin': {'top': 8, 'bottom': 16}}}}%%
flowchart TD
A["Input video frame sequence<br/>x_1, ..., x_T"] --> B["VAE encoder E mapping<br/>compact latent code z_i"]
B --> C["Single-step diffusion specialists<br/>trained on disjoint datasets"]
B --> D["Unified backbone G_phi<br/>with ESA temporal memory"]
C -.->|provide target latents l_i^k| E["Latent joint distillation<br/>lightweight DPT projectors Psi^k"]
D --> E
E --> F["Temporal gradient matching<br/>suppressing flicker across frames"]
F --> G["Streaming multi-task decoding<br/>Psi^k + VAE D' + Pk single pass"]
G --> H["8 video dense prediction outputs<br/>geometry / semantics / intrinsics"]
Key Designs¶
1. Single-step diffusion specialists: bridging domain gaps with generative visual priors
Training on disjoint datasets with heterogeneous domain characteristics (synthetic indoor scans for geometry versus web images for semantics) causes standard discriminative backbones (such as DINOv3) to overfit to spurious appearance-geometry correlations. UniD instead adopts Stable Diffusion v2 to build single-step diffusion specialists \((F_\theta^k, P^k)\). Each input frame \(\mathbf{x}_i\) is encoded into a compact latent representation \(\mathbf{z}_i = \mathcal{E}(\mathbf{x}_i) \in \mathbb{R}^{h \times w \times d}\) by a frozen VAE encoder. Rather than iterative multi-step denoising, the U-Net backbone \(F_\theta^k\) performs a single-step prediction of the task latent code \(\mathbf{l}_i^k \in \mathbb{R}^{h \times w \times d}\). This latent code is decoded by the frozen VAE decoder truncated before its final convolutional layer \(\mathcal{D}'(\cdot)\), and mapped to target dimension \(C_k\) via a single \(3 \times 3\) convolution pixel projector \(P^k\):
$$ (F_\theta^k, P^k) = \arg\min_{F^k, P^k} \mathbb{E}_{(\mathbf{x}, \mathbf{y}^k) \sim \mathcal{D}_k} \left[ \mathcal{L}_k\left( P^k(\mathcal{D}'(F^k(\mathcal{E}(\mathbf{x})))), \mathbf{y}^k \right) \right] $$
This design allows each specialist to inherit internet-scale visual priors and appearance-invariant geometry, unlocking robust out-of-distribution generalization.
2. Latent joint distillation: bypassing pseudo-labeling storage and memory bottlenecks
Supervising a unified model with per-pixel predictions across \(K=8\) tasks on shared data requires either pre-computing terabytes of multi-task label maps or retaining \(K\) simultaneous decoder computational graphs in GPU memory, which explodes peak memory to an estimated 650GB per device. UniD completely sidesteps this bottleneck by distilling task representations directly in latent space. The unified backbone \(\mathbf{G}_\phi\) mirrors the specialist U-Net architecture and is coupled with \(K\) lightweight Dense Prediction Transformer (DPT) latent projectors \(\{\Psi^k\}_{k=1}^K\) that fuse three intermediate features with the output representation at a channel dimension of 256. For any input video sequence, frozen specialists compute latent targets \(\mathbf{l}_i^k = F_\theta^k(\mathbf{z}_i)\) on the fly, while the unified model predicts \(\hat{\mathbf{l}}_i^k = \Psi^k(\mathbf{G}_\phi(\mathbf{z}_i))\). By operating purely in latent space without backpropagating through the high-resolution VAE decoder \(\mathcal{D}'\), peak training memory is slashed from 650GB to 78GB (an 88% reduction), while adding a \((K+1)\)-th task requires only training a new projector without retraining the backbone.
3. Temporal gradient matching: propagating temporal consistency across unannotated tasks
Because video annotations are heavily asymmetric across domainsโavailable for depth or panoptic segmentation but entirely absent for intrinsics and materialsโa naive per-frame reconstruction loss fails to impart temporal coherence to static tasks. UniD first equips the backbone with Extended Self-Attention (ESA), concatenating keys and values from \(M\) historical frames stored in a memory bank \(\mathcal{M}_t\) with the current frame without introducing extra parameters. To regularize dynamic feature continuity during distillation, UniD generalizes Temporal Gradient Matching (TGM) to the latent space:
$$ \mathcal{L}{\text{TGM}}(\hat{\mathbf{l}}, \mathbf{l}) = \sum}^{T-1} \mathbb{I{[\cdot]} \cdot \left| \frac{\hat{\mathbf{l}}} - \hat{\mathbf{l}i}{|\hat{\mathbf{l}}} - \hat{\mathbf{l}i|_2} - \frac{\mathbf{l}} - \mathbf{li}{|\mathbf{l} \right|_1 $$
where } - \mathbf{l}_i|_2\(\mathbb{I}_{[\cdot]}\) suppresses learning in fast-changing or occluded regions. The overall latent distillation objective is \(\mathcal{L}_{\text{rec}} = \|\hat{\mathbf{l}}_i^k - \mathbf{l}_i^k\|_1 + \lambda \mathcal{L}_{\text{TGM}}(\hat{\mathbf{l}}^k, \mathbf{l}^k)\). Through this joint constraint, tasks lacking video supervision (albedo, shading, material) effectively inherit temporal stability from video-supervised tasks.
Loss & Training¶
The unified backbone \(\mathbf{G}_\phi\) and latent projectors \(\{\Psi^k\}_{k=1}^K\) are trained on the union of all task datasets, sampled with equal 50:50 weight between images and videos. Optimization uses AdamW with a learning rate of \(3 \times 10^{-5}\), weight decay of \(10^{-2}\), and batch size 32 for 50k iterations. For discrete classification tasks (semantic segmentation), the authors note that \(\ell_1\) latent reconstruction tends to attenuate high-frequency boundaries (cross-entropy specialist latents show a spatial Laplacian norm of 2.52 versus 1.40 for regression tasks). To compensate, a lightweight fine-tuning (FT) stage is applied exclusively to the task projectors \((\Psi^k, P^k)\) while keeping the unified backbone \(\mathbf{G}_\phi\) completely frozen.
Key Experimental Results¶
Main Results¶
Evaluation covers all 8 dense prediction tasks with a primary focus on unseen out-of-distribution (OOD) benchmarks. In monocular depth and surface normal estimation, UniD demonstrates decisive advantages over multi-task unified baselines and the strong discriminative DINOv3-H model.
| Dataset (Depth) | Metric | UniD (Gฯ) [Ours] | DINOv3-H (Multi-task) | 4M-21-XL (Unified) | DICEPTION (Unified) |
|---|---|---|---|---|---|
| NYUv2 (Indoor OOD) | AbsRel โ / \(\delta_1\) โ | 0.059 / 96.3 | 0.092 / 92.9 | 0.068 / 95.1 | 0.060 / 96.3 |
| KITTI (Outdoor OOD) | AbsRel โ / \(\delta_1\) โ | 0.110 / 89.5 | 0.121 / 86.5 | 0.105 / 89.6 | 0.085 / 91.8 |
| ETH3D (Real OOD) | AbsRel โ / \(\delta_1\) โ | 0.073 / 95.5 | 0.087 / 93.7 | 0.070 / 95.3 | 0.070 / 96.6 |
| ScanNet (Indoor OOD) | AbsRel โ / \(\delta_1\) โ | 0.063 / 96.2 | 0.089 / 93.6 | 0.065 / 95.5 | 0.071 / 94.8 |
| DIODE (Mixed OOD) | AbsRel โ / \(\delta_1\) โ | 0.301 / 77.3 | 0.314 / 77.5 | 0.331 / 73.4 | 0.304 / 72.2 |
| Dataset (Surface Normal) | Metric | UniD (Gฯ) [Ours] | DINOv3-H (Multi-task) | 4M-21-XL (Unified) | DICEPTION (Unified) |
|---|---|---|---|---|---|
| NYUv2 (Indoor OOD) | Mean โ / \(11.25^\circ\) โ | 16.2 / 59.9 | 16.8 / 57.0 | 18.4 / 49.8 | 19.7 / 51.4 |
| ScanNet (Indoor OOD) | Mean โ / \(11.25^\circ\) โ | 14.6 / 65.1 | 16.0 / 58.7 | 17.4 / 48.9 | 19.3 / 54.5 |
| iBims-1 (Real OOD) | Mean โ / \(11.25^\circ\) โ | 16.6 / 67.4 | 16.7 / 66.8 | 18.6 / 61.4 | 17.8 / 66.2 |
| Sintel (Synthetic OOD) | Mean โ / \(11.25^\circ\) โ | 33.0 / 21.4 | 32.3 / 18.9 | 41.6 / 12.4 | 37.2 / 19.6 |
On intrinsic decomposition and boundary tasks, UniD matches in-domain baselines while substantially outperforming them on OOD benchmarks (IIW albedo WHDR reduced to 0.207 vs. 0.289 for DINOv3-H; SAW shading [email protected] reaching 94.6 vs. 87.8; Mapillary boundary odsF reaching 57.2 vs. 55.0).
Ablation Study¶
The ablation investigates the interaction between video-level training (with TGM loss) and inference-time temporal memory (ESA) across video consistency benchmarks (mVC-4/16) and depth accuracy.
| Train w/ Video | Inf. Memory | Depth ScanNet (AbsRel โ) | Normal InteriorNet (mVC-4/16 โ) | Albedo InteriorNet (mVC-4/16 โ) | Shading InteriorNet (mVC-4/16 โ) | Boundary YT-VIS (mVC-4 โ) | Semantic VIPSeg (mVC-4 โ) |
|---|---|---|---|---|---|---|---|
| โ (static image distillation) | \(\emptyset\) (single frame) | 0.121 | 79.0 / 65.6 | 84.4 / 65.0 | 92.7 / 81.5 | 92.9 | 85.9 |
| โ (static image distillation) | [16, 1] (16-frame queue) | 0.115 | 82.9 / 72.7 | 87.3 / 70.7 | 93.6 / 82.9 | 93.0 | 89.1 |
| โ (with temporal TGM) | \(\emptyset\) (single frame) | 0.109 | 80.9 / 68.5 | 86.0 / 67.1 | 93.7 / 83.5 | 93.6 | 86.5 |
| โ (full UniD model) | [16, 1] (16-frame queue) | 0.103 | 84.7 / 75.8 | 88.8 / 73.7 | 94.5 / 85.2 | 93.7 | 90.3 |
In terms of inference efficiency on an A100 GPU at 432ร768 resolution: the specialist ensemble (\(K=8\) sequential U-Net runs) achieves only 0.44 FPS at \(M=16\). In contrast, UniD executes the U-Net backbone only once (235.0 ms), reducing per-frame latency from 2.27s to 0.64s to achieve 1.56 FPS, delivering an over 3.5ร speedup.
Key Findings¶
- Cross-task temporal spillover: While only depth and semantic segmentation contain video annotations during training, incorporating video training and TGM improves temporal consistency for tasks without any video supervisionโalbedo mVC rises from 84.4 to 88.8 and shading mVC rises from 92.7 to 94.5, confirming cross-task temporal feature sharing.
- Diffusion priors guard against domain shifts: Under synthetic color jittering on Hypersim validation data, the diffusion specialist exhibits 3ร greater stability than DINOv3-H (depth AbsRel deviation 0.011 vs. 0.034; normal mean error deviation \(2.2^\circ\) vs. \(5.8^\circ\)), proving that generative pretraining captures invariant geometry rather than memorizing surface textures.
- Cross-task mutual regularization: On ScanNet geometric consistency between predicted depth and normal, the unified model UniD (\(22.6^\circ\)) outperforms both DINOv3-H (\(36.0^\circ\)) and the independent specialist ensemble (\(23.9^\circ\)), demonstrating that unified multi-task distillation enforces mutually coherent geometric constraints.
Highlights & Insights¶
- Disjoint multi-task latent distillation paradigm: UniD demonstrates that unified dense prediction across disparate modalities does not require unified datasets or heavy offline pseudo-labeling; single-task specialists trained on isolated domains can supervise a unified backbone purely through on-the-fly latent feature reconstruction.
- Single-step diffusion with shared backbone efficiency: Replacing multi-step iterative diffusion sampling with single-step direct latent regression, combined with a shared backbone and lightweight DPT heads, enables unified 8-task streaming inference at 1.56 FPS under 16-frame temporal context.
- Modular horizontal extensibility: Introducing an additional dense prediction task (e.g., optical flow or 3D bounding boxes) only requires training an isolated specialist and a single lightweight latent projector \(\Psi^{K+1}\), without modifying the shared backbone or re-labeling legacy data.
Limitations & Future Work¶
- Frequency degradation in classification latents under \(\ell_1\) distillation: Because cross-entropy-trained segmentation latents exhibit significantly higher spatial Laplacian norms (2.52 vs. 1.40 for regression), pure \(\ell_1\) latent reconstruction attenuates high-frequency semantic boundaries (Cityscapes mIoU drops from 64.1 to 37.8, restored to 58.4 via projector fine-tuning). Future work could explore frequency-aware or contrastive latent reconstruction objectives.
- Memory scaling under long temporal windows: The Extended Self-Attention mechanism caches past key-value representations in GPU memory, which scales with sequence length; future iterations could incorporate state-space models (Mamba) or recurrent latent memories to achieve 30+ FPS real-time streaming throughput.
Related Work & Insights¶
- vs DICEPTION / 4M-21: DICEPTION routes all predictions through the RGB image space conditioned on task prompts, demanding \(K\) separate diffusion passes to predict \(K\) tasks; 4M-21 discretizes outputs into tokens, constraining continuous regression precision. UniD operates in continuous latent space, predicting all 8 tasks in a single pass with arbitrary output dimensionalities.
- vs DINOv3-H Multi-task Baseline: While discriminative self-supervised backbones achieve strong in-domain performance, they degrade sharply under out-of-distribution shifts due to overfitting to appearance-geometry correlations; UniD leverages internet-scale generative priors to maintain robust generalization across unseen environments.
Rating¶
- Novelty: โญโญโญโญโ (Creative use of pretrained diffusion generative priors and latent distillation to unite completely disjoint datasets)
- Experimental Thoroughness: โญโญโญโญโญ (Thorough evaluation across 8 dense tasks, OOD generalization, temporal metrics, cross-task consistency, and memory profiling)
- Writing Quality: โญโญโญโญโญ (Clear problem exposition with principled analysis of Laplacian frequency characteristics across task spaces)
- Value: โญโญโญโญโญ (Provides an efficient, extensible blueprint for unified video perception across fragmented real-world datasets)