VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/ShareLab-SII/VLZip
Area: Multimodal VLM
Keywords: Multimodal Token Compression, Long-Context VLM, Interleaved Image-Text Reasoning, Soft Prefix Injection, Q-Former
TL;DR¶
VLZip addresses the quadratic attention complexity bottleneck in interleaved image-text long-context modeling by hierarchically distilling visual and textual chunks into layer-specific soft prefixes and injecting them into decoder hidden states, enabling 6x longer training sequences up to 120K tokens and superior performance on the narrative-driven LongVLBench benchmark.
Background & Motivation¶
Vision-language models (VLMs) have demonstrated impressive mastery over single-image recognition and short conversational tasks, yet real-world applications—such as deciphering illustrated operation manuals, analyzing longitudinal visual medical histories, parsing long-form video narratives, and tracking long-horizon multimodal GUI agent trajectories—routinely require reasoning over ultra-long, interleaved sequences of images and text. Unlocking genuine long-context multimodal intelligence is hindered by two intertwined hurdles: computationally, the quadratic scaling of self-attention causes prohibitive GPU memory explosion and latency bottlenecks; evaluatively, prevailing multimodal benchmarks often repurpose disjointed short-context datasets, rely on synthetic "needle-in-a-haystack" (NIAH) retrieval tasks that test trivial fact recall rather than holistic comprehension, or pad inputs with irrelevant filler text that distorts authentic narrative modeling.
Existing efforts to mitigate the architectural bottleneck predominantly accept compromises between throughput and reasoning fidelity. One prominent line of research focuses on aggressive token pruning (e.g., GlimpsePrune, VisionZip), which permanently drops tokens almost exclusively from the visual modality, risking irreversible loss of fine-grained spatial evidence while completely neglecting the burgeoning textual volume in interleaved inputs. A second line explores architectural overhauls, replacing standard attention with State Space Models (SSMs) or linear recurrent variants; however, these designs often sacrifice the sharp, non-local associative reasoning at which Transformers excel. A third line relies on data-centric curation without altering the underlying quadratic computational footprint. Furthermore, while soft prompt compression methods in NLP demonstrate the possibility of distilling long text into compact prefixes, they remain restricted to text-only inputs and typically apply shallow, single-layer injection that degrades over deep transformer layers.
The entry point of this work is that one does not need to abandon the pure Transformer architecture nor permanently discard crucial context tokens; rather, extensive interleaved multimodal contexts can be distilled externally into compact, multi-layer "soft prefixes" and delivered directly into every decoder layer. Core idea: propose VLZip, a unified visual and textual compression framework that hierarchically distills interleaved image and text chunks into layer-specific soft prefixes and injects them via residual addition into placeholder hidden states across every decoder layer, trading a modest external compression cost for a 25x reduction in attention sequence length and high-fidelity long-context reasoning.
Method¶
Overall Architecture¶
VLZip handles an input sequence \(X = \{S, C, Q\}\) comprising a system prompt \(S\), an inquiry question \(Q\), and an extensive interleaved context \(C = \{t_0, v_0, \dots, t_n, v_n\}\). To minimize computational cost, the prompt and question remain uncompressed, while only the lengthy context \(C\) undergoes compression. The overall workflow consists of three consecutive stages: first, modality-specific hierarchical compressors independently distill image chunks and long text segments into compact feature representations specialized for all \(L\) decoder layers; second, the original context chunks are substituted with a minimal number of lightweight placeholder tokens (text placeholders T-P and image placeholders I-P), compressing the self-attention sequence length by a factor of 25; finally, during the LLM decoder's forward pass, the corresponding layer-specific soft prefix features are injected via element-wise residual addition into the placeholder hidden states prior to self-attention at each layer, supplying uncompromised global context across the entire network depth.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Interleaved Long-Context Input<br/>System Prompt S + Interleaved Context C + Question Q"] --> B["Dual-Modality Chunking & Soft Prefix Distillation<br/>Visual chunking & light text encoding + Q-Former L-layer queries"]
B --> C["Cross-Modal Placeholder Compact Sequence Construction<br/>Replace raw image/text chunks with compact T-P / I-P tokens"]
C --> D["Layer-Wise Hidden State Residual Injection<br/>Inject layer-specific soft prefixes into placeholder hidden states"]
D --> E["Long-Context Self-Attention & Generation<br/>Short-sequence self-attention + fine-grained global reasoning"]
Key Designs¶
1. Dual-Modality Chunking & Soft Prefix Distillation: Unified Extraction of Layer-Specific Representations To overcome the limitations of prior techniques that prune only visual tokens while ignoring interleaved text growth, and single-step compression that loses nuanced details, VLZip establishes symmetric hierarchical compression modules for both modalities. For the visual modality, each image is processed by the vision encoder into \(N_{v,j}\) visual tokens and partitioned into uniform chunks of size \(C_v\) (default 100). Each visual chunk is passed through a shared visual Q-Former equipped with \(L \times M_v\) learnable queries \(\mathbf{q}_v \in \mathbb{R}^{(L \cdot M_v) \times d}\), extracting in a single forward pass layer-specific soft prefix features \(\mathbf{Z}_{v,j,p} \in \mathbb{R}^{L \times M_v \times d}\) tailored for all \(L\) decoder layers. Symmetrically, for the textual modality, long text segments are partitioned into uniform chunks of size \(C_t\) (default 100) tokens, encoded by a lightweight text encoder (Qwen2.5-0.5B), projected through a linear matrix \(\mathbf{W}_{\text{proj}}\), and processed by a text Q-Former with \(L \times M_t\) learnable queries \(\mathbf{q}_t \in \mathbb{R}^{(L \cdot M_t) \times d}\) to generate layer-specific text prefixes \(\mathbf{Z}_{t,i,k} \in \mathbb{R}^{L \times M_t \times d}\). This single-pass query design eliminates the latency overhead of re-invoking compressors per layer while decoupling layer-wise contextual semantics.
2. Cross-Modal Placeholder Compact Sequence Construction: Drastic Reduction in Attention Sequence Length To alleviate the quadratic memory and latency explosion inside the backbone LLM, VLZip reconstructs the input sequence using a placeholder substitution mechanism. For each raw image chunk containing \(C_v\) tokens, the sequence retains only \(M_v\) (default 4) visual placeholder tokens (I-P); identically, each raw text chunk of \(C_t\) tokens is represented by only \(M_t\) (default 4) textual placeholder tokens (T-P). This mechanism shrinks the effective context length entering the LLM embedding layer by a factor of \(C / M = 25\). For instance, an expansive 100K-token multimodal context is condensed to a mere 4K placeholder sequence, leaving system prompts and questions intact at full fidelity. This formulation preserves the exact chronological interleaving topology while keeping self-attention computational overhead minimal.
3. Layer-Wise Hidden State Residual Injection: Sustained Fine-Grained Global Context Across Depth Addressing the drawback of conventional soft prompting—where prepended prompt vectors become progressively diluted across deep Transformer layers—VLZip deploys a layer-wise residual injection mechanism across all decoder layers. At decoder layer \(l\), letting \(\mathbf{H}_{v,j,p}^{(l)} \in \mathbb{R}^{M_v \times d}\) and \(\mathbf{H}_{t,i,k}^{(l)} \in \mathbb{R}^{M_t \times d}\) denote the placeholder hidden states for the respective visual and textual chunks, the corresponding layer-specific prefix slices are added directly to the hidden states before self-attention:
The updated placeholder representations are subsequently processed alongside prompt and question tokens through self-attention and feed-forward networks. By repeatedly injecting freshly distilled layer-specific contextual representations, the model enables compact 4-token placeholders to continuously carry the descriptive capacity of 100 original tokens across all network depths.
Loss & Training¶
To train the disparate compression and reasoning modules stably, VLZip adopts a four-stage progressive curriculum training pipeline: 1. Stage 1: Visual Compressor Pre-training. Trained on single-image instruction-following data (LLaVA-OneVision single-image subset) with the vision backbone and LLM decoder frozen; only the projection adapter and vision Q-Former are updated to learn distilling visual chunks into layer-specific prefixes. 2. Stage 2: Textual Compressor Pre-training. Pre-trained via a self-supervised text reconstruction objective on long-document corpora (ChatQA2); the lightweight text encoder, projection matrix \(\mathbf{W}_{\text{proj}}\), and text Q-Former are optimized to compress and reconstruct text chunks accurately. 3. Stage 3: Joint Interleaved Fine-tuning. Conducted on interleaved multi-image and multi-text corpora, unfreezing the LLM decoder while keeping both compressors trainable, training the unified backbone to fuse and reason over compressed dual-modality soft prefixes. 4. Stage 4: Long-Context Consolidation. Continual training on ultra-long textual and interleaved sequences with all Stage 3 parameters active, fully adapting the model's global dependency modeling to extreme context lengths.
Key Experimental Results¶
Main Results¶
The authors conduct comprehensive evaluations on the newly proposed LongVLBench (spanning 7 context length bins from 0-4k to >128k, comprising 140 video narrative QA samples) and the standard MMLongBench (evaluating VRAG, NIAH, and ICL tasks), using Qwen2.5-VL-Instruct-3B as the primary backbone.
The table below reports comparative results across context length bins on LongVLBench, including inference success rates in parentheses (indicating avoidance of OOM crashes):
| Model | Size | 0-4k | 4k-8k | 8k-16k | 16k-32k | 32k-64k | 64k-128k | >128k | Avg |
|---|---|---|---|---|---|---|---|---|---|
| Long-VITA | 14B | 5.25 (100%) | 6.10 (100%) | 6.94 (80%) | 3.80 (75%) | 5.78 (45%) | 5.33 (15%) | 0.00 (0%) | 33.1 |
| LongLLaVA | 9B | 4.70 (100%) | 5.15 (100%) | 4.95 (100%) | 4.20 (100%) | 6.10 (100%) | 5.20 (100%) | 1.20 (100%) | 45.0 |
| Mantis | 8B | 0.70 (100%) | 2.00 (100%) | 1.88 (85%) | 0.00 (10%) | 2.00 (5%) | 0.00 (0%) | 0.00 (0%) | 6.3 |
| Qwen2.5-VL | 7B | 4.90 (100%) | 5.55 (100%) | 5.80 (100%) | 5.25 (100%) | 5.70 (100%) | 4.25 (100%) | 1.58 (95%) | 47.1 |
| InternVL2.5 | 4B | 4.95 (100%) | 5.05 (100%) | 5.85 (100%) | 4.90 (100%) | 3.15 (100%) | 2.90 (100%) | 0.95 (95%) | 39.1 |
| Ovis2 | 4B | 4.85 (100%) | 6.25 (100%) | 6.25 (100%) | 4.35 (100%) | 5.95 (100%) | 0.40 (100%) | 0.50 (100%) | 40.8 |
| GlimpsePrune | 3B | 5.35 (100%) | 4.70 (100%) | 6.20 (100%) | 4.45 (100%) | 5.10 (100%) | 2.21 (95%) | 0.00 (65%) | 39.9 |
| VisionZip | 3B | 4.95 (100%) | 4.70 (100%) | 6.89 (90%) | 3.75 (20%) | 2.00 (5%) | 4.00 (15%) | 0.00 (0%) | 24.7 |
| Qwen2.5-VL (Baseline) | 3B | 5.35 (100%) | 4.90 (100%) | 6.20 (100%) | 4.30 (100%) | 5.60 (100%) | 2.58 (95%) | 0.16 (95%) | 41.4 |
| VLZip (Ours) | 3B | 4.60 (100%) | 5.65 (100%) | 5.75 (100%) | 5.60 (100%) | 6.70 (100%) | 7.45 (100%) | 7.60 (100%) | 61.9 |
The table below summarizes performance at extreme 128k context lengths on MMLongBench:
| Model | Size | VRAG-128k | NIAH-128k | ICL-128k |
|---|---|---|---|---|
| LongLLaVA | 9B | 18.9 | 36.5 | OOM / - |
| Qwen2.5-VL | 7B | 31.1 | 27.0 | 44.0 |
| InternVL2.5 | 4B | 21.3 | 25.4 | 0.8 |
| GlimpsePrune | 3B | 8.2 | 13.5 | 6.8 |
| Qwen2.5-VL (Baseline) | 3B | 10.8 | 13.4 | 7.5 |
| VLZip (Ours) | 3B | 23.3 | 25.3 | 58.8 |
Ablation Study¶
The ablation experiments investigate the necessity of unified compression across modalities, injection depth strategies, and chunk configurations.
The table below details the maximum trainable sequence length achieved on 8×A100 GPUs with DeepSpeed ZeRO-2 under different compression combinations:
| Visual Compression | Textual Compression | Max Trainable Sequence Length (Tokens) |
|---|---|---|
| ✗ | ✗ | 20K |
| ✓ | ✗ | 50K |
| ✗ | ✓ | 40K |
| ✓ | ✓ | 120K |
The table below analyzes the impact of different layer injection strategies with fixed \(C=100, M=4\):
| Injection Strategy | VRAG avg | VRAG 128k | NIAH avg | NIAH 128k | ICL avg | ICL 128k | LongVL avg | LongVL >128k |
|---|---|---|---|---|---|---|---|---|
| First-layer | 31.2 | 31.3 | 31.3 | 27.8 | 75.8 | 55.5 | 59.6 | 8.00 |
| First-half | 30.0 | 28.0 | 28.9 | 27.1 | 70.2 | 42.3 | 60.6 | 7.65 |
| Interval-2 | 28.9 | 27.8 | 30.6 | 27.2 | 70.8 | 39.3 | 61.8 | 8.25 |
| Full-layer | 26.0 | 23.3 | 30.5 | 25.3 | 76.6 | 58.8 | 61.9 | 7.60 |
Key Findings¶
- Dual-modality compression is non-negotiable for long-context training: Compressing only images or only text achieves limited expansion (reaching 50K and 40K tokens, respectively), as the uncompressed modality quickly triggers out-of-memory errors; unifying both modalities expands trainable context to 120K tokens (a 6x improvement).
- Resilience and emergence at extreme lengths: On LongVLBench at sequences beyond 128k tokens, competing baselines degrade sharply to 0.16 or crash from OOM, whereas VLZip achieves a leading score of 7.60, demonstrating that multi-layer injection preserves narrative cohesion when long sequences overwhelm conventional models.
- Full-layer injection underpins multi-shot in-context learning: While early-layer injection slightly favors local retrieval matching (VRAG/NIAH), full-layer injection dominates on complex sequence synthesis, boosting ICL-128k to 58.8 (outperforming the 7B baseline at 44.0) and overall LongVLBench average to 61.9.
- Chunk granularity shifts the balance between retrieval and synthesis: Smaller chunks (\(C=25\)) retain finer localized cues for short-range retrieval, while larger chunks (\(C=100, M=4\)) capture coherent global semantic abstractions essential for extreme-context reasoning.
Highlights & Insights¶
- Single-pass multi-layer prefix generation via query expansion: By scaling Q-Former learnable queries to \(L \times M\) dimensions, VLZip extracts layer-specific soft prefixes across all decoder layers simultaneously in one forward pass, avoiding iterative layer-by-layer compression overhead.
- Drastic sequence reduction within a pure Transformer: Condensing attention length by 25x allows a 3B model to execute 280K token inference on a single 80GB A100 GPU at 23.6x speedup, while prefilling peak memory curves remain virtually flat up to 2M tokens.
- LongVLBench benchmark eliminates artificial noise: Constructed from video narratives using CLIP keyframe clustering and multi-tiered captioning, LongVLBench provides chronologically grounded, interleaved sequences that measure holistic narrative reasoning rather than artificial needle-in-a-haystack memorization.
Limitations & Future Work¶
- Performance trade-off on short-context visual perception: On standard short-context VQA benchmarks (e.g., GQA, AI2D), fixed chunk compression induces a modest 5%–10% drop relative to uncompressed baselines, suggesting a need for dynamic resolution routing or skip-connections for fine-grained localized details.
- Symmetric compression ratio across modalities: The framework adopts identical chunk sizes and token counts (\(C=100, M=4\)) for both images and text, despite marked differences in information density between vision and natural language. Exploring asymmetric compression rates could further refine the Pareto frontier.
- Heuristic bypass for short text fragments: Text sequences under 100 characters are currently bypassed via a fixed heuristic threshold rather than an adaptive, data-driven gating mechanism.
Related Work & Insights¶
- vs Input Pruning (GlimpsePrune, VisionZip): Pruning methods permanently drop tokens during early stages, incurring irreversible fine-grained information loss and ignoring textual expansion; VLZip preserves high-order semantics across all chunks and compresses both modalities symmetrically.
- vs Non-Transformer Architectures (LongLLaVA, LongMamba): Linear attention and state space models reduce sequence costs at the expense of non-local associative reasoning capacity; VLZip preserves the full expressive power of a standard Transformer decoder.
- vs Text Soft Prompt Compression (GIST, ICA, DeepStack): Prior soft prompt techniques focus solely on language and inject prompts at the input layer; DeepStack explores multi-layer visual stacking but omits interleaved text. VLZip is the first to unify dual-modality compression with layer-wise residual injection for interleaved contexts.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant formulation of dual-modality chunk distillation combined with layer-wise soft prefix injection inside a standard Transformer.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Exhaustive evaluation spanning memory scaling, latency, standard benchmarks, and the custom narrative-level LongVLBench.
- Writing Quality: ⭐⭐⭐⭐⭐ Clearly articulated motivation, rigorous mathematical formulation, and well-structured empirical ablation analysis.
- Value: ⭐⭐⭐⭐⭐ Highly practical blueprint for enabling 100K+ token multimodal training and inference without altering core Transformer architectures.