Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: EventHosts PDF
Area: Image Generation
Keywords: Video Generation, Adaptive Tokenization, Variational Autoencoder, Flow Matching, Diffusion Model
TL;DR¶
KATok introduces an adaptive video VAE that jointly learns dynamic token selection via Gumbel-Softmax gating and soft attention masking, coupled with a cascaded mask-prior diffusion strategy to eliminate spatio-temporal misalignment, achieving higher generation quality and several-fold acceleration.
Background & Motivation¶
Latent diffusion models (LDMs) have become the cornerstone of high-fidelity image and video synthesis, relying on variational autoencoders (VAEs) to compress high-dimensional pixels into compact continuous latent spaces to lower generative computational costs while preserving visual quality. However, conventional convolutional and transformer-based VAEs deployed in video generation adhere to fixed compression ratios across spatial and temporal dimensions. As a result, the number of latent tokens scales linearly with video resolution and sequence length, severely constraining scalable high-resolution and long-duration video generation.
Natural video data exhibits extreme spatio-temporal redundancy: static backgrounds and textureless regions demand substantially less representational capacity than dynamic actions and fine motion textures. Uniform fixed-grid tokenization inevitably squanders capacity on uninformative visual areas. While recent flexible-length tokenizers allow manual adjustment of token budgets, they lack true content adaptivity, requiring heuristic schedules or costly inference-time search to determine sample-specific token allocations. Crucially, directly applying sparse tokens resulting from adaptive dropping to downstream diffusion models causes severe content-position misalignment: without a regular grid, the generative model struggles to correctly bind visual content to physical coordinates, resulting in structural distortion and temporal flicker.
To overcome these challenges, KATok investigates an end-to-end differentiable token-selection mechanism that autonomously decides which tokens to retain or prune based on content complexity, while resolving the spatial alignment issue during generative diffusion. Core idea: develop an end-to-end differentiable adaptive transformer VAE (KATok) utilizing Gumbel-Softmax gating and soft attention masking for data-dependent compression, paired with a cascaded mask-prior diffusion model to eliminate spatio-temporal misalignment while providing several-fold generative speedups.
Method¶
Overall Architecture¶
The KATok pipeline encompasses adaptive tokenization and downstream sparse flow-matching generation. During tokenization, input video frames are partitioned into spatio-temporal 3D patches and fed into a transformer encoder equipped with 3D RoPE and global register tokens. A lightweight classification head predicts per-token keep/drop logits; during training, continuous soft masks are sampled via Gumbel-Softmax relaxation and shared with the decoder cross-attention layers via soft attention biasing. The decoder reconstructs the video through asymmetric fine-grained query tokens. During downstream generation, to prevent geometric misalignment on sparse tokens, the model decouples structure prediction from content synthesis: a lightweight mask-prior DiT first generates the binary active grid coordinates, and the primary flow-matching DiT synthesizes sparse latent content conditioned on these explicit spatial positions before KATok's decoder reconstructs the final video.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Video (T×H×W×C)"] --> B["Adaptive Keep-or-Drop Tokenization<br/>Gumbel-Softmax Gating & Soft Masks"]
B --> C["Soft Attention Mask Decoding<br/>Asymmetric Multi-Grid Query Reconstruction"]
C --> D["Cascaded Mask-Prior Generation<br/>Lightweight Prior Predicts Active Coordinates"]
D --> E["Sparse Flow-Matching Generation<br/>Coordinate-Conditioned Compact Latents"]
E --> F["Output High-Fidelity Video"]
Key Designs¶
1. Adaptive Keep-or-Drop Tokenization with Soft Attention Masking: Learning Data-Driven Spatio-Temporal Pruning To resolve the inefficiency of fixed-rate compression over redundant spatio-temporal backgrounds, KATok places a lightweight head after the encoder to predict keep/drop logits \(\alpha_i \in \mathbb{R}^2\). During training, Gumbel-Softmax relaxation samples differentiable soft masks \(\tilde{m}_i \in (0, 1)\), modulating the continuous reparameterized latents into \(z_i = \tilde{m}_i \cdot \hat{z}_i\). To guarantee that dropped tokens do not inadvertently leak information into the reconstruction process, the decoder cross-attention matrix is modified with a soft additive bias: $\(A_{ij} \leftarrow A_{ij} + \log(\tilde{m}_j + \varepsilon)\)$ Coupled with an \(\ell_1\) sparsity penalty \(\mathcal{L}_{\text{sparse}} = \mathbb{E}_X [\sum_{i=1}^N \tilde{m}_i(X)]\) whose weight \(\lambda_{\text{sparse}}\) is gradually annealed, the network autonomously learns to retain dense tokens across complex dynamic regions while aggressively discarding static background tokens, eliminating the need for manual token budgets.
2. Asymmetric Coarse-to-Fine Decoding with Global Register Tokens: Compact Latents and High-Fidelity Reconstruction To preserve fine geometric contours and high-frequency textures under heavy token pruning, KATok introduces an asymmetric resolution scheme. The encoder operates on coarse 3D patches of size \(16^2 \times 8\), generating a highly compact sequence of latent tokens for efficient sparsification and downstream diffusion. Conversely, the decoder expands reconstruction granularity using an \(8^2 \times 4\) fine grid of learnable query tokens that attend to the surviving sparse latents via cross-attention. Concurrently, global register tokens in the encoder serve as persistent semantic anchors; ablations confirm that removing them degrades token efficiency and harms spatial coherence, proving their role in stabilizing global structures during aggressive pruning.
3. Cascaded Mask-Prior Conditioning: Resolving Spatio-Temporal Misalignment in Sparse Diffusion Training flow-matching models directly on unordered sparse latents causes severe spatial misalignment, as the network struggles to simultaneously discover unstructured coordinate positions and synthesize detailed content. While joint content-position prediction partially alleviates this issue, it remains overly sensitive to decoupled noise schedules. KATok instead introduces a cascaded architecture: an 8.3M lightweight mask-prior model first predicts the binary selection mask across the full spatio-temporal grid, from which active coordinates are extracted and injected as deterministic positional embeddings into the 686M primary flow generator. This complete decoupling of "where to generate" from "what visual details to synthesize" ensures consistent spatio-temporal geometry even under extreme token sparsity.
Loss & Training¶
The tokenizer training objective combines multiple objectives: $\(\mathcal{L} = \mathcal{L}_{\text{recon}} + \lambda_{\text{KL}}\mathcal{L}_{\text{KL}} + \lambda_{\text{sparse}}\mathcal{L}_{\text{sparse}} + \lambda_{\text{adv}}\mathcal{L}_{\text{adv}} + \lambda_{\text{align}}\mathcal{L}_{\text{align}}\)$ \(\mathcal{L}_{\text{recon}}\) incorporates an S3D-based Video-LPIPS to enforce temporal motion consistency. \(\mathcal{L}_{\text{align}}\) aligns representations using features from a frozen vJEPA-2 encoder, while latent Gaussian noise sampled from \([0, 0.2]\) is injected during training to enhance downstream generative robustness. Training follows a three-stage curriculum: single-resolution pretraining (\(256^2 \times 16\)), multi-resolution training, and perceptual GAN fine-tuning.
Key Experimental Results¶
Main Results¶
Evaluated on approximately 5,000 validation videos from Panda-70M for reconstruction fidelity, and on UCF-101, SkyTimelapse, and Kinetics-600 using a SiT-XL backbone for generative performance.
Table 1: Video Reconstruction and Compression Comparison on Panda-70M Validation Set
| Method | Resolution | #Tokens | Channels | Comp. \(\uparrow\) | PSNR \(\uparrow\) | SSIM \(\uparrow\) | LPIPS \(\downarrow\) | rFVD \(\downarrow\) |
|---|---|---|---|---|---|---|---|---|
| Omni-VAE | \(256^2 \times 17^*\) | 5120.00 | 8 | 96.00 | 28.10 | 0.88 | 0.05 | 7.84 |
| Omni-VAE | \(512^2 \times 33^*\) | 36864.00 | 8 | 96.00 | 24.07 | 0.80 | 0.06 | 16.85 |
| ElasticTok-KL | \(256^2 \times 16\) | 3845.56 | 8 | 102.25 | 30.52 | 0.91 | 0.06 | 12.37 |
| KATok (Ours) | \(256^2 \times 16\) | 366.24 | 64 | 134.21 | 31.24 | 0.94 | 0.04 | 5.12 |
| KATok (Ours) | \(512^2 \times 32\) | 1554.24 | 64 | 253.00 | 33.23 | 0.95 | 0.05 | 6.40 |
*Note: \(*\) indicates evaluation on the first 16/32 frames; ElasticTok does not support \(512^2\) resolution.
Table 2: Video Generation Quality and Throughput Comparison (gFVD \(\downarrow\))
| Model | SkyTimelapse | UCF-101 | Kinetics-600 | Training Speedup | Throughput (videos/s) |
|---|---|---|---|---|---|
| Omni-VAE + SiT-XL | 23.28 | 100.00 | 206.58 | 1.0× (Baseline) | 4.91 |
| Elastic-KL + SiT-XL | 95.53 | 712.56 | — | — | 4.24 |
| Ours-Joint | 21.36 | 73.16 | 193.78 | — | 5.01 |
| Ours-Cascaded | 23.19 | 61.53 | 160.84 | 6.9× | 15.71 |
Ablation Study¶
Component-wise ablations conducted under controlled 100K training steps on the VAE and diffusion strategies.
Table 3: Ablation on KATok VAE Components (100K Training Steps, Max Tokens = 514)
| Configuration | PSNR \(\uparrow\) | LPIPS \(\downarrow\) | SSIM \(\uparrow\) | Active Tokens \(\downarrow\) | Note |
|---|---|---|---|---|---|
| Full model | 30.85 | 0.07 | 0.93 | 365.57 | Baseline configuration |
| − latent reg. | 31.26 | 0.07 | 0.94 | 361.70 | Slight reconstruction gain, hurts generation |
| − asymmetric decoding (\(16^2 \times 8\)) | 29.61 | 0.11 | 0.92 | 377.80 | Coarse reconstruction degrades fine details |
| − Video-LPIPS (image LPIPS) | 29.28 | 0.11 | 0.91 | 404.00 | Lacks temporal cues, retains redundant tokens |
| − register tokens | 29.17 | 0.12 | 0.91 | 410.20 | Loss of global anchors reduces token efficiency |
| − soft attention mask | 19.00 | 0.51 | 0.54 | 2.00 | Training collapse; only register tokens active |
| − Gumbel-Softmax | 18.91 | 0.52 | 0.53 | 2.00 | Gradient breakage causes complete collapse |
Table 4: Generation Strategy Ablation on UCF-101 (gFVD \(\downarrow\))
| Group | Variant | gFVD \(\downarrow\) | Core Analysis |
|---|---|---|---|
| Latent Regularization | With latent reg. | 101.23 | Smooths latent distribution, boosting generation |
| w/o latent reg. | 161.23 | Overfits reconstruction, degrading generative fidelity | |
| Diffusion Architecture | Naïve Generation | 95.69 | Suffers heavily from content-position misalignment |
| Joint Generation | 73.16 | Decoupled denoising aligns positions, hyperparameter-sensitive | |
| Cascaded Generation | 61.53 | Mask prior decouples position from content, achieving optimal quality |
Key Findings¶
- Soft Attention Masking and Differentiable Relaxation Are Vital: Disabling soft attention masking or Gumbel-Softmax causes immediate training collapse down to 2 register tokens, demonstrating that explicit attention modulation and end-to-end gradient flow are mandatory for learning adaptive selection.
- Token Allocation Correlates Strongly with Temporal Entropy: Empirical correlation reveals that the number of active tokens correlates strongly with temporal entropy (\(r = 0.865\)) and spatio-temporal entropy (\(r = 0.872\)), but only moderately with spatial entropy (\(r = 0.618\)). Static or blank clips require as few as 28 tokens for near-perfect reconstruction.
- Dual Leap in Speed and Fidelity via Cascaded Prior: Reducing token counts from 5,120 to ~366 accelerates training convergence by 6.9× and yields a 3.2× inference throughput gain (15.71 videos/s), while improving UCF-101 gFVD by 38.5 points compared to dense OmniTokenizer.
Highlights & Insights¶
- Differentiable Soft Gating in Cross-Attention: Incorporating \(\log(\tilde{m}_j + \varepsilon)\) directly into the attention logit acts as a smooth, continuous surrogate for hard pruning, allowing backpropagation from reconstruction loss to directly inform token selection.
- Emergent Motion Controllability: At inference time, modulating the sequence length of the sampled noise tokens (e.g., from 200 to 400) provides continuous, intuitive control over visual detail and motion intensity without requiring task-specific retraining.
- Decoupled Structural Prior for Sparse Diffusion: Decomposing sparse generation into a lightweight 8.3M mask-prior stage and a deep content generator offers a robust, generalizable blueprint for handling unstructured, sparse representations in continuous flow matching.
Limitations & Future Work¶
- Token Count Inflation in Highly Dynamic Scenes: In videos dominated by global camera motion or chaotic particle dynamics (e.g., heavy storms), the active token count increases significantly, reducing computational savings relative to fixed dense models.
- Cascaded Error Propagation: If the lightweight mask prior generates sub-optimal or sparse spatial positions, the primary generator cannot correct positional topology, suggesting the need for joint bidirectional feedback in future work.
- Broader Modality Extension: While focused on video synthesis, this adaptive keep-or-drop paradigm is readily extensible to 3D point cloud generation and 4D dynamic neural field modeling.
Related Work & Insights¶
- vs OmniTokenizer: OmniTokenizer enforces a fixed spatio-temporal grid across all video regions; KATok prunes over 90% of tokens while surpassing OmniTokenizer in both reconstruction and generative FVD.
- vs ElasticTok: ElasticTok adopts a truncated sequence approach requiring per-sample binary search at inference; KATok is fully content-adaptive and search-free, resulting in a 46.7× faster training setup.
- vs Token Pruning (DynamicViT / ToMe): Traditional token pruning methods are tailored for inference acceleration in discriminative tasks; KATok establishes a unified training-and-inference framework for generative continuous latent spaces.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First principled exploration of end-to-end adaptive token pruning and cascaded positional conditioning for continuous video diffusion.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation spanning Panda-70M, UCF-101, SkyTimelapse, and Kinetics-600 with entropy analysis and exhaustive ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear motivation, rigorous mathematical formulation, and well-structured empirical validation.
- Value: ⭐⭐⭐⭐⭐ Establishes a highly efficient, scalable foundation for next-generation high-resolution and long-duration video generation models.