Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/Sta8is/Re2Pix
Area: Autonomous Driving
Keywords: video prediction, vision foundation models, hierarchical video generation, latent diffusion models, nested dropout
TL;DR¶
Re2Pix decouples future video prediction into two stagesโvision foundation model (VFM) semantic representation forecasting followed by semantics-guided latent diffusion synthesisโand introduces nested dropout alongside mixed supervision to resolve autoregressive train-test mismatch, delivering 7ร to 14ร training convergence speedups and state-of-the-art temporal semantic consistency and visual fidelity on driving benchmarks.
Background & Motivation¶
In autonomous driving and complex physical environments, anticipating how a visual scene evolves over time is essential for long-horizon reasoning, risk assessment, and safe trajectory planning. However, learning predictive world models directly from raw, uncurated video sequences requires mastering high-level semantic dynamics (such as the trajectories and interactions of vehicles and pedestrians) and fine-grained visual details (such as lighting shifts, reflections, occlusions, and road textures) simultaneously. Most contemporary approaches adopt an end-to-end paradigm, predicting future frames within the latent space of a Variational Autoencoder (VAE) using latent diffusion transformers. The fundamental shortcoming of this paradigm is that high-level semantic topology and low-level appearance details are deeply entangled within the exact same latent space, forcing the generator to simultaneously infer scene physics and render photorealistic pixels. This entanglement frequently results in temporal semantic collapse, object identity drift, flickering artifacts, and prohibitively slow training convergence.
Recent approaches have attempted to alleviate this burden by aligning intermediate diffusion features with pretrained representations via auxiliary distillation objectives (e.g., REPA and VideoREPA). Yet, because these diffusion models still perform both forecasting and rendering within a single latent space, dynamics and appearance remain implicitly coupled. Conversely, pure semantic forecasting frameworks (such as DINO-Foresight and DINO-WM) predict future representations in feature space without synthesizing RGB pixels, leaving the downstream generative loop unaddressed.
The core tension is whether semantic dynamics forecasting and generative appearance synthesis can be explicitly separated while maintaining temporally coherent, photorealistic generation. Re2Pix resolves this by adopting a coarse-to-fine hierarchical formulation: it first models future scene structure within the representation space of a frozen vision foundation model (VFM), and subsequently conditions a latent diffusion model on these predicted representations to synthesize photorealistic RGB frames. Core idea: decompose future video prediction into frozen VFM semantic representation forecasting and semantics-guided latent diffusion generation, employing nested dropout and mixed supervision to bridge the distribution shift caused by autoregressively accumulated prediction errors.
Method¶
Overall Architecture¶
Re2Pix addresses the video prediction task by predicting future frames \(x_{M+1:K}\) given historical context observations \(x_{1:M}\). The pipeline strictly follows a two-stage hierarchical paradigm: 1. Stage 1: High-Level Semantic Feature Prediction. A frozen vision foundation model encoder (DINOv2-Reg ViT-B/14) independently extracts high-level semantic feature maps from context frames. A lightweight Masked Feature Transformer autoregressively forecasts the future semantic feature maps one frame at a time. 2. Stage 2: Semantics-Guided Video Generation. A causal 3D VAE compresses historical context frames into compact latent codes. A diffusion transformer (DiT backbone) takes clean context latents, noisy future latents, and the sequence of semantic features, fusing them at the input level via early semantic alignment. Guided by the predicted semantic skeleton, the model denoises future latents, which are ultimately decoded into photorealistic RGB frames by the 3D VAE decoder.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Context video frames x(1:M)"] --> B["Stage 1: VFM Feature Prediction<br/>Frozen DINOv2 extraction + Masked Transformer autoregression"]
A --> C["Causal 3D VAE Encoding<br/>Compress to compact latents z(1:M)"]
B --> D["Stage 2: Early Semantic Alignment<br/>Bilinear resize + independent projection + channel-wise sum"]
C --> D
D --> E["Nested Dropout & Mixed Supervision<br/>Stochastic channel truncation + 90% GT / 10% Pred mixture"]
E --> F["Diffusion Transformer Denoising<br/>DiT backbone predicts clean future latents z(M+1:K)"]
F --> G["3D VAE Decoder Reconstruction<br/>Synthesize photorealistic, semantically consistent RGB frames"]
Key Designs¶
1. VFM Feature Prediction: Decoupling Dynamics from Appearance in Semantic Space Predicting directly in pixel or generic VAE latent spaces forces the network to juggle high-level physical trajectories alongside high-frequency textural noise. To isolate structural scene dynamics, Re2Pix employs a frozen DINOv2 encoder \(E_h\), extracting intermediate features from transformer blocks 3, 6, 9, and 12, concatenated and projected to \(C_h = 1152\) channels via PCA. This feature representation inherently organizes information hierarchically, with top components capturing broad geometric layouts and lower components encoding finer boundaries. A masked transformer \(G_h\) is trained to predict future representations: during training, it receives \(M\) unmasked context frames and regresses the fully masked \((M+1)\)-th frame feature via Smooth L1 loss. At inference time, \(G_h\) rolls out predictions autoregressively in a frame-wise manner, establishing a temporally coherent semantic trajectory prior to visual rendering.
2. Early Semantic Alignment: Zero-Token-Overhead Input-Level Conditioning Injecting high-dimensional temporal semantic guidance into diffusion backbones traditionally relies on cross-attention mechanisms, which introduce massive parameter overhead (adding approximately 265M parameters in baseline tests) and inflate computational complexity. Global modulation approaches such as AdaLN, on the other hand, struggle to enforce fine-grained spatial correspondences. Re2Pix instead designs a direct early fusion strategy: VAE latents \(z\) are patchified with a spatial patch size of \(2 \times 2\), and the VFM semantic feature maps \(h\) are bilinearly interpolated to match this exact spatial patch grid. Both modalities are independently projected into the transformer embedding dimension and fused via element-wise, channel-wise addition at the input layer. This provides dense, localized structural conditioning throughout all self-attention layers without adding any tokens to the sequence length.
3. Nested Dropout and Mixed Supervision: Bridging the Autoregressive Train-Test Gap A naive two-stage pipeline suffers from severe train-test mismatch: during training, the diffusion generator has access to pristine, ground-truth future semantic features from \(E_h(x_{M+1:K})\) via teacher forcing; at test time, it must rely exclusively on features autoregressively predicted by \(G_h\), which inevitably accumulate drift and noise. When trained solely on clean semantics, the diffusion generator overfits to perfect high-frequency features, resulting in blurry, degraded outputs at test time. Re2Pix overcomes this distribution shift using two complementary mechanisms: - Nested Feature Dropout: Leveraging the variance hierarchy of PCA features, training stochastically retains only the first \(c \in \{8, 16, 32, 64, 128, 256, 512, 1152\}\) channels for all semantic feature maps with equal probability, zeroing out the remaining \(C_h - c\) channels: $\(\tilde{h}_{1:K} = \left[ h_{1:K}^{1:c},\, \mathbf{0}^{C_h - c} \right]\)$ This forces the diffusion model to rely on coarse, robust semantic subspaces rather than memorizing brittle, fine-grained channel correlations. - Mixed Supervision: During training, batches are stochastically sampled such that 10% of instances use predicted features from \(G_h\) while 90% use ground-truth features from \(E_h\). This 90/10 mixture regularizes the model against over-reliance on idealized conditioning while preserving high visual sharpness and semantic alignment.
Loss & Training¶
The two stages are trained independently, ensuring modularity and stable convergence: 1. Feature Prediction Objective: Masked transformer \(G_h\) is trained using a Smooth L1 loss against ground-truth future features: $\(\mathcal{L}_{\text{feat}} = \text{SmoothL1}\left( G_h(h_1, \ldots, h_M),\, h_{M+1} \right)\)$ 2. Diffusion Generation Objective: The latent video diffusion model \(G_z\) (DiT backbone with 14 layers, 16 heads, 2048 hidden dimensions, 800M parameters, AdaLN-LoRA, and RMSNorm) follows the EDM formulation. Denoising loss is computed exclusively on future frame latents (\(t = M+1, \ldots, K\)): $\(\mathcal{L}_{\text{diffusion}} = \mathbb{E}_{\epsilon, n} \left[ \lambda_n \left\| G_z(z_{M+1:K}^{(n)};\, z_{1:M},\, \tilde{h}_{1:K},\, n) - z_{M+1:K} \right\|^2 \right]\)$ Training on 8 NVIDIA H200 GPUs takes approximately 7 hours (40k iterations) for single-dataset benchmarks, achieving a 7ร to 14ร speedup in convergence over end-to-end baselines.
Key Experimental Results¶
Main Results¶
On the Cityscapes benchmark (conditioning on 13 frames to predict 12 future frames; evaluating semantic segmentation mIoU and depth estimation on frame 19, alongside video-level FID and FVD across all predicted frames), Re2Pix outperforms standard diffusion baselines, REPA/VideoREPA feature alignment methods, and a larger 1.5B parameter baseline:
| Method | Parameters | mIoU(A)โ | IoU(M)โ | ฮด1โ | AbsRโ | FIDโ | FVDโ |
|---|---|---|---|---|---|---|---|
| Re2Pix (Stage 1 Feature Bound) | - | 69.76 | 69.66 | 88.03 | 0.1262 | - | - |
| Baseline | 782M | 60.55 | 57.64 | 85.15 | 0.1460 | 12.86 | 60.70 |
| Baseline w/ REPA | 792M | 61.45 (+0.90) | 59.63 (+1.99) | 85.35 (+0.20) | 0.1465 (+0.0005) | 12.34 (-0.52) | 55.15 (-5.55) |
| Baseline w/ VideoREPA | 792M | 60.98 (+0.43) | 58.61 (+0.97) | 85.22 (+0.07) | 0.1466 (+0.0006) | 12.95 (+0.09) | 59.03 (-1.67) |
| Baseline-Large | 1.5B | 61.69 (+1.14) | 59.65 (+2.01) | 85.43 (+0.28) | 0.1453 (-0.0070) | 11.99 (-0.87) | 56.81 (-3.89) |
| Re2Pix (Ours) | 1.1B | 63.53 (+2.98) | 62.29 (+4.65) | 85.72 (+0.57) | 0.1413 (-0.0047) | 9.90 (-2.96) | 52.66 (-8.04) |
Ablation Study¶
Comprehensive ablation experiments validate the critical role of conditioning strategies and architectural fusion mechanisms on Cityscapes:
| Ablation Category | Configuration | mIoU(A)โ | IoU(M)โ | ฮด1โ | AbsRโ | FIDโ | FVDโ | Note |
|---|---|---|---|---|---|---|---|---|
| Nested Dropout | Fixed 1152 Channels | 62.38 | 60.67 | 85.45 | 0.1467 | 12.63 | 58.91 | Overfits to fine-grained feature channels |
| Nested Dropout | 63.53 | 62.29 | 85.72 | 0.1413 | 9.90 | 52.66 | Substantial gain in FID/FVD and semantic metrics | |
| Supervision Mixture | Ground Truth Only (100% GT) | 64.42 | 63.15 | 86.06 | 0.1407 | 10.43 | 80.85 | Severe train-test gap leads to blurry rollout and high FVD |
| Predicted Only (100% Pred) | 62.77 | 61.38 | 85.58 | 0.1420 | 10.21 | 55.81 | Degraded semantic fidelity | |
| Mixed (85/15) | 62.49 | 60.00 | 85.52 | 0.1440 | 10.54 | 54.89 | Suboptimal trade-off; reduced semantic guidance | |
| Mixed (90/10) | 63.53 | 62.29 | 85.72 | 0.1413 | 9.90 | 52.66 | Optimal balance of fidelity and semantic stability | |
| Mixed (95/05) | 63.49 | 62.33 | 85.75 | 0.1319 | 10.27 | 53.69 | Slight degradation in perceptual quality | |
| Fusion Strategy | Cross-Attention (+265M params) | 61.34 | 59.03 | 85.30 | 0.1462 | 11.83 | 56.37 | Heavier compute, inferior performance |
| AdaLN Conditioning | 63.10 | 61.75 | 85.71 | 0.1425 | 10.75 | 55.75 | Lacks dense spatial alignment | |
| Early-Fusion (Input Sum) | 63.53 | 62.29 | 85.72 | 0.1413 | 9.90 | 52.66 | Zero added tokens or parameters; best overall |
Key Findings¶
- Decoupling dynamics from synthesis accelerates convergence: Re2Pix reaches an FID of 15 within 20k iterations, whereas the end-to-end baseline requires 140k iterationsโrepresenting a 7ร training speedup for generation quality. Semantic segmentation mIoU achieves an even faster 14ร speedup.
- Nested dropout prevents over-reliance on fine-grained features: Introducing nested dropout slashes FID from 12.63 to 9.90 and FVD from 58.91 to 52.66. At inference, dropping channels from 1152 down to 128 retains strong performance (FID 9.80, FVD 52.82), confirming that the top principal components govern scene structure.
- Mixed supervision prevents distribution collapse: Training solely on ground-truth semantics causes test-time FVD to surge to 80.85 due to blurry reconstructions from noisy inputs. The 90/10 mixture resolves this distribution shift completely.
- Structural advantage surpasses raw parameter scaling: Baseline-Large (1.5B parameters) remains distinctly inferior to Re2Pix (1.1B parameters) across all semantic and generative metrics, showing that structural decoupling outweighs brute-force scaling.
Highlights & Insights¶
- From auxiliary distillation to generative intermediate: Unlike REPA-based approaches that treat foundation models as static feature distillation targets, Re2Pix converts VFM representations into an explicit generative intermediate, establishing a physical division of labor between dynamics modeling and pixel rendering.
- Lightweight early fusion design: Channel-wise input addition after spatial alignment avoids memory-heavy cross-attention layers while providing dense structural conditioning across all DiT layers at zero extra token cost.
- PCA-aligned nested dropout: Stochastically truncating PCA feature channels directly leverages the natural variance distribution of orthogonal components, creating an effective multi-granularity regularizer without auxiliary routing networks.
- Broad cross-task transferability: The semantics-first hierarchical framework is directly applicable to physical world models, robot policy rollouts, and video prediction tasks where dynamics reasoning must precede appearance rendering.
Limitations & Future Work¶
- Unidirectional error propagation: If the autoregressive feature predictor \(G_h\) fails to predict an unexpected entity (such as an occluded pedestrian stepping into view), the downstream diffusion generator cannot independently synthesize the missing semantic object.
- Domain bias of frozen VFMs: The framework relies heavily on DINOv2 representations; severe domain shifts (e.g., adverse weather, dense fog, night glare) that degrade VFM feature extraction will propagate errors through the generative pipeline.
- Future directions: Integrating bidirectional feedback or uncertainty-aware gating between the diffusion latent space and the feature predictor, as well as end-to-end joint fine-tuning strategies.
Related Work & Insights¶
- vs Cosmos-Predict / Vista (End-to-End Diffusion World Models): Cosmos and Vista predict future frames directly in latent/pixel space, demanding massive web-scale video datasets and compute while remaining susceptible to temporal drift. Re2Pix explicitly isolates semantic dynamics into feature forecasting, achieving superior semantic consistency with dramatically lower training overhead.
- vs REPA / VideoREPA (Representation Alignment): REPA enforces feature similarity in hidden layers but leaves the diffusion model responsible for temporal forecasting from scratch. Re2Pix autoregressively forecasts VFM features first, supplying the diffusion model with an explicit structural blueprint.
- vs DINO-Foresight / DINO-WM (Feature-Only World Models): DINO-Foresight models scene evolution solely within representation space, lacking pixel generation capabilities. Re2Pix completes the loop by synthesizing high-fidelity, photorealistic RGB video from forecasted features.
Rating¶
- Novelty: โญโญโญโญโญ First systematic framework to guide hierarchical video diffusion via VFM representation forecasting, supported by elegant nested dropout and mixed supervision designs.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluations across Cityscapes, nuScenes, CoVLA, and zero-shot KITTI, accompanied by exhaustive ablations and convergence analyses.
- Writing Quality: โญโญโญโญโญ Clear exposition, thorough motivation of the representation-pixel tension, and rigorous mathematical and architectural formulations.
- Value: โญโญโญโญโญ Provides a highly practical, computationally efficient, and semantically grounded hierarchical paradigm for autonomous driving world models and video prediction.