Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/showlab/PadCaptioner
Area: Video Understanding
Keywords: Dense video captioning, Parallel decoding, Latent planning, Omni-modal Video-LLM, Causal dependency restructuring
TL;DR¶
Addressing the severe inference latency and cross-event interference in standard token-by-token dense video captioning, PadCaptioner restructures the causal dependency graph via latent global planning and achieves lossless parallel autoregressive decoding with a 3.8ร actual wall-clock speedup using event-factorized attention.
Background & Motivation¶
Dense Video Captioning (DVC) requires jointly localizing densely occurring temporal events in untrimmed long videos and generating detailed descriptions for each event. With the remarkable generative capabilities and cross-modal reasoning abilities of Multimodal Large Language Models (MLLMs), autoregressive Video-LLMs have emerged as the dominant paradigm. However, standard autoregressive frameworks adhere to a strictly sequential, token-by-token decoding pattern. Because DVC involves multi-sentence narratives across multiple densely occurring events within long-form video, the total generated tokens frequently reach hundreds or thousands, leading to prohibitive inference latency. While diffusion-based LLMs (dLLMs) explore position-wise parallel generation, their iterative refinement steps conflict with KV cache acceleration, yielding limited practical speedup and inferior generation quality on video understanding tasks.
The root deficiency of standard sequential autoregressive modeling lies in forcing all generated tokens into a single monolithic causal chain. In untrimmed videos, content is inherently structured into temporally distinct events: fine-grained semantic dependencies are tightly concentrated within event boundaries (local intra-event coherence), whereas cross-event relations are primarily governed by high-level narrative and temporal structure, meaning local token-level dependencies across events are exceptionally weak. Enforcing full token-level causal dependency across events introduces massive serialization overhead and injects irrelevant cross-event context into attention calculations, diluting the model's focus on the active event.
This paper's angle of attack is to break the monolithic causal chain by incorporating video temporal priors into causal dependency graph restructuring. Core idea: exploit the weak local dependencies across temporally distinct events to restructure the causal dependency graph, first using autoregressive latent global planning to generate event tokens and adaptively aggregate audio-visual semantics, and then performing event-factorized parallel decoding across subchains to achieve lossless grounded captioning with balanced global awareness and local focus.
Method¶
Overall Architecture¶
PadCaptioner maps an omni-modal prefix (temporally interleaved video-audio frames and textual instruction tokens) into structured, parallel dense captions. Instead of decoding all tokens in a single sequential chain, generation is split into two coordinated stages: (1) Latent Global Planning, where the model autoregressively generates a compact sequence of global event tokens with an adaptively determined event count; and (2) Dependency-Restructured Parallel Decoding, where event-specific caption subchains are anchored by their corresponding global event tokens and decoded synchronously in parallel across time steps.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Audio-visual streams & prompt input<br/>Interleaved multimodal prefix P"] --> B["Autoregressive latent planning<br/>Predict global event token sequence"]
B --> C["Adaptive semantic aggregation<br/>Aggregate salient audio-visual features"]
C --> D["Event-factorized parallel decoding<br/>Synchronous multi-branch subchains"]
D --> E["Output omni-modal event timestamps & captions"]
Key Designs¶
1. Adaptive Latent Global Planning: Autoregressively generating explicitly grounded global event anchors
To establish a global temporal structure prior to parallel decoding without fixing the number of events beforehand, PadCaptioner introduces a special token <G> into the vocabulary. Conditioned on the multimodal prefix \(P\), the model autoregressively predicts a sequence of global event tokens \(\{G_1, G_2, \dots, G_K\}\) followed by a termination switch token \(S\). The event count \(K\) is adaptively inferred from the video content. To ensure that each latent token \(G_i\) reliably anchors a valid temporal event, the model optimizes an explicit grounding constraint during training. Using ground-truth event intervals, normalized dot-product similarity \(\text{sim}_t^i\) is computed between \(G_i\) and the multimodal prefix tokens across all time segments \(t \in [1, T]\), optimized under a binary cross-entropy loss:
$\(\mathcal{L}_{\text{gnd}}^{G_i} = - \frac{1}{T} \sum_{t=1}^T \Big( y_t^i \log \text{sim}_t^i + (1 - y_t^i) \log (1 - \text{sim}_t^i) \Big)\)$
This constraint embeds intrinsic grounding capabilities into each \(G_i\). At inference time, event localization is directly derived from feature similarity matching against the prefix tokens without requiring separate regression heads.
2. Adaptive Semantic Aggregation: Injecting enriched multimodal event semantics into latent tokens
A bare planning token lacks the fine-grained details needed to guide descriptive captioning. To provide rich local cues at the start of parallel branch decoding, PadCaptioner dynamically aggregates audio-visual prefix features from the grounded temporal interval into \(G_i\). Facing spatial, temporal, and cross-modal heterogeneity, the authors evaluated uniform mean pooling, a lightweight MLP scoring head, and attention-guided aggregation. Attention-guided aggregation proved superior: it directly repurposes the cross-attention weights between the query's terminal prompt token and the multimodal prefix as importance scores, requiring zero additional parameters while weighting informative visual and acoustic tokens and adding the aggregated representation back into the \(G_i\) embedding.
3. Event-Factorized Parallel Decoding: Decoupling cross-event local interference while preserving global awareness
During text generation, all \(K\) event branches decode their local tokens synchronously at each step \(j\). The joint conditional distribution is factorized as: $\(\{L_j^i\}_{i=1}^K \sim \mathbb{P}\Big( \{L_j^i\}_{i=1}^K \;\Big|\; P, G^{1:K}, \{G^i, L_{<j}^i\}_{i=1}^K \Big)\)$ In attention modeling, an event-factorized causal mask is applied: each local token \(L_j^i\) attends to the shared multimodal prefix \(P\) and all global event tokens \(G^{1:K}\) (maintaining global narrative awareness), as well as its own event anchor \(G^i\) and intra-branch past tokens \(L_{<j}^i\) (ensuring intra-event autoregressive coherence). Crucially, local tokens from different event branches are completely masked from one another. This eliminates spurious cross-event token interactions while retaining global structural guidance.
4. Structure-Agnostic Positional Encoding & Independent Branch Termination: Efficient batched parallel inference
Because event caption lengths vary, synchronized batch decoding requires proper alignment. PadCaptioner assigns identical relative positional indices to tokens at decoding step \(j\) across all subchains, eliminating artificial sequential ordering between parallel branches. During training, subchains shorter than the maximum length \(N\) are padded with <EOS> tokens whose loss is included to incentivize prompt termination. During inference, when any subchain produces an <EOS> token, its KV-cache updates terminate immediately without waiting for longer branches, preventing redundant computation.
Key Experimental Results¶
Main Results¶
On the omni-modal benchmarks LongVALE and ChronusAV, PadCaptioner (3B) surpasses previous 7B SOTA models across all metrics, setting a new benchmark in event grounding (F1) and caption quality (Sim, SODA_c, CIDEr, METEOR).
| Benchmark | Method | Model Scale | F1 (Grounding) โ | Sim (Similarity) โ | SODA_c โ | CIDEr โ | METEOR โ |
|---|---|---|---|---|---|---|---|
| LongVALE | LongVALE-LLM | 7B | 31.2 | 37.2 | 2.8 | 7.9 | 4.7 |
| LongVALE | ChronusOmni | 7B | 49.7 | 52.4 | 3.7 | 5.6 | 5.2 |
| LongVALE | PadCaptioner (Ours) | 3B | 56.4 | 58.5 | 6.4 | 13.7 | 8.6 |
| ChronusAV | LongVALE-LLM | 7B | 18.4 | 25.4 | 3.1 | 4.6 | 8.8 |
| ChronusAV | ChronusOmni | 7B | 60.1 | 36.8 | 8.6 | 2.9 | 7.6 |
| ChronusAV | PadCaptioner (Ours) | 3B | 63.2 | 40.0 | 12.4 | 9.6 | 12.4 |
In wall-clock decoding latency evaluations (measured on a single NVIDIA A6000 GPU), PadCaptioner demonstrates dramatic throughput acceleration:
| Model | LongVALE Latency/Video (ms) โ | LongVALE Latency/Token (ms) โ | ChronusAV Latency/Video (ms) โ | ChronusAV Latency/Token (ms) โ |
|---|---|---|---|---|
| ChronusOmni (7B) | 16162.1 | 41.3 | 28093.0 | 41.0 |
| video-SALMONN-2+ (3B) | 2706.8 | 22.5 | 3083.4 | 22.8 |
| PadCaptioner (3B, Ours) | 4283.8 (3.8ร speedup) | 13.4 (3.1ร speedup) | 7526.7 (3.7ร speedup) | 13.7 (3.0ร speedup) |
Note: video-SALMONN-2+ exhibits lower raw video latency primarily because it outputs terse, low-quality captions (F1 35.5/21.5, Sim 26.6/17.6). When normalized per generated token (T/token), PadCaptioner achieves a decisive speedup.
Ablation Study¶
On ChronusAV (trained on a 12K video subset), ablation experiments confirm the exact role of each architectural component:
| Setting / Variant | F1 (Grounding) โ | Sim (Semantics) โ | T/video (ms) โ | T/token (ms) โ | Note |
|---|---|---|---|---|---|
| Baseline (Standard Sequential) | 32.7 | 17.6 | 7798.7 | 22.9 | Monolithic sequential chain, prone to cross-event noise |
| + Textual planning | 37.4 | 19.4 | - | - | Sequential text planning of event boundaries |
| + Latent planning | 61.5 | 38.4 | 12405.1 | 22.9 | Massive accuracy boost, but sequential latency increases |
| ++ Parallel decoding (PadCaptioner full) | 61.8 | 38.8 | 7598.2 | 13.8 | Quality improves further, decoding speedup 1.66ร |
| Factorized attention replaced by causal attention | 38.7 | 19.9 | - | - | Unmasked cross-event tokens introduce severe interference |
| Access to self \(G^i\) only (no global \(G^{1:K}\)) | 59.8 | 37.6 | - | - | Drops inter-event global context, degrading coherence |
| w/o semantic aggregation | 47.7 | 24.3 | - | - | Latent tokens lack event-specific audio-visual grounding |
| Mean pooling aggregation | 58.5 | 35.7 | - | - | Uniform pooling fails to distinguish token importance |
| Scoring head aggregation | 61.4 | 38.2 | - | - | Lightweight MLP predicts token importance |
| Attention-guided aggregation | 61.8 | 38.8 | - | - | Zero parameters, optimal task-relevant feature pooling |
Key Findings¶
- Parallel decoding provides lossless speedup with slight performance gains: Transitioning from sequential latent planning to factorized parallel decoding improves F1 from 61.5 to 61.8 and Sim from 38.4 to 38.8. Eliminating unnecessary cross-event token attention prevents long-context semantic dilution and sharpens intra-event focus.
- Factorized attention masking is essential: Substituting the factorized mask with standard causal attention causes F1 to collapse to 38.7 and Sim to 19.9, demonstrating that uncoordinated cross-talk across concurrently generated event tokens corrupts generation.
- Attention-guided aggregation is lightweight and expressive: Leveraging cross-attention maps for feature pooling achieves higher accuracy (61.8 / 38.8) than mean pooling (58.5 / 35.7) and outpaces a parameterized scoring head (61.4 / 38.2) without extra weights.
Highlights & Insights¶
- Mapping physical temporal structure to causal graph sparsity: By recognizing that distinct video events enjoy conditional independence given global context, the paper restructures attention into block-diagonal subchains, escaping sequential latency without sacrificing autoregressive fluency.
- Dual-role latent planning tokens: The global event tokens simultaneously act as geometric temporal anchors (via similarity-based localization) and semantic anchors that broadcast global context to parallel subchains.
- Generalizable decoupled generation paradigm: The "latent global planning + factorized parallel decoding" structure offers a compelling blueprint for multi-view video captioning, long-form document section synthesis, and embodied multi-step action planning.
Limitations & Future Work¶
- Event-level granularity constraint: The dependency graph restructuring operates strictly at the event level. For long-duration single events, intra-event decoding remains sequential; exploring sub-event or spatiotemporal object-level parallelism is a promising direction.
- Semantic coarsening in long events: Adaptive aggregation over prolonged intervals can smooth out fine-grained temporal dynamics, occasionally yielding generic descriptions for extended segments.
- Hardware-level kernel optimization: The custom block-diagonal attention mask currently relies on general PyTorch attention implementations; engineering specialized FlashAttention kernels could further amplify physical throughput.
Related Work & Insights¶
- vs PDVC (ICCV 2021): PDVC adopts a DETR-like query architecture for parallel event generation but cannot inherit modern MLLM pretrained weights and language priors; PadCaptioner restructures MLLM autoregressive attention natively, preserving foundation model capabilities.
- vs Diffusion LLMs (dLLMs): While dLLMs support sequence-level parallelism, multi-step iterative denoising breaks standard KV caching, impairing real-world speed and generation quality in video understanding; PadCaptioner retains causal autoregression within subchains, retaining full KV-cache compatibility.
Rating¶
- Novelty: โญโญโญโญโญ Elegant, theoretically grounded causal graph restructuring tailored to the temporal structure of video.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluations across LongVALE, ChronusAV, and YouCook2, reporting wall-clock latency, parameter efficiency, and detailed ablations.
- Writing Quality: โญโญโญโญโญ Exceptionally clear framing of motivation, dependency formulation, and empirical validations.
- Value: โญโญโญโญโญ Provides a highly practical, performant roadmap for accelerating inference in long-context multimodal video models.