G2TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models¶
Conference: NeurIPS2026
arXiv: 2605.12309
Code: https://github.com/lijunxian111/G2TR
Area: VLM Efficiency (vlm_efficiency)
Keywords: unified multimodal models, visual token compression, generation guidance, VAE latent anchors, training-free
TL;DR¶
G2TR uses VAE latent anchors to select and merge understanding-side visual tokens before LLM prefill in separate-encoder unified multimodal models; retaining 50% of the tokens yields 99.0% relative-average understanding performance and 98.0% editing performance on BAGEL, with approximately 1.94ร lower prefill FLOPs, but not lossless performance on every task.
Background & Motivation¶
Unified multimodal models (UMMs) both answer image questions and edit images according to instructions. Separate-encoder architectures provide different entry points for these demands: a ViT understanding encoder extracts semantic features, while a VAE generation encoder extracts latent representations for reconstruction and generation. Decoupling encoding mitigates the conflicting demands of abstract semantics and low-level appearance details, but does not eliminate the cost of processing dense visual tokens in the subsequent Transformer. BAGEL and InternVL-U are the architectures studied here, not representatives of every possible UMM design.
Most existing visual token compression methods target vision-language models whose final output is a textual answer. FastV prunes using internal LLM attention, VSCAN combines visual attention with image-text similarity, and IVC-Prune incorporates RoPE and image-text signals. Yet understanding-side tokens also condition image editing: colors, textures, or background layout that appear unimportant for an answer can still determine whether an edit preserves source-image details. The paper observes limited differences between several text-centric rules on some understanding benchmarks at a fixed budget, alongside local-detail degradation in editing. This motivates a different guidance signal, but does not establish that all understanding tasks are insensitive to pruning rules.
The generation branch already contains representations learned through reconstruction and generation objectives, avoiding the need to train another importance predictor. The method uses VAE encodings of the same input image as local references for ViT features, while encouraging spatial distribution and recovering unselected features. Core idea: constrain understanding-side visual token compression with generation-side latent representations, so compressed semantic conditions retain editing-relevant visual details rather than only tokens useful for text prediction.
Method¶
Overall Architecture¶
The input is a source image and a textual instruction; the output is a shortened understanding-side sequence compatible with the original UMM. Both the original ViT and VAE encoders remain unchanged: full ViT features enter Latent-Anchor Scoring, guided by VAE features; Balanced Budget Selection, Similar-Token Merging, and Sequence Layout Rebuilding then produce the input for LLM prefill.
Generation guidance does not mean generating an image before selecting tokens: only VAE encoding of the input image is required. The compressed objects are always understanding-side ViT tokens, not the VAE generation latents. Editing continues through the original generation pathway; pure text-to-image generation has no source-image understanding tokens and is not directly affected by this compression.
The diagram depicts inference data flow only, with no additional training supervision or optimization loop. For understanding tasks, VAE encoding provides the scoring signal; for editing tasks, it also belongs to the model's existing source-image processing pathway.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Source image<br/>and instruction"] --> V["Original ViT encoding<br/>full understanding tokens"]
I --> Z["Original VAE encoding<br/>generation-side features"]
V --> A["Latent-Anchor Scoring"]
Z --> A
A --> B["Balanced Budget Selection"]
B --> C["Similar-Token Merging"]
C --> D["Sequence Layout Rebuilding"]
I -->|preserve text and image boundaries| D
D --> O["Original LLM prefill<br/>answer or editing conditions"]
Key Designs¶
1. Latent-Anchor Scoring: measure consistency with local generation representations
Raw ViT tokens and VAE latents differ in feature dimension and spatial granularity, so their raw vectors cannot simply be compared. The model's existing branch-specific projectors map them into a common hidden dimension; Section 4.2 subsequently uses the understanding-token and latent notation for these aligned features. This relies on the pretrained model's existing representations and projectors, not additional alignment training introduced by G2TR.
After restoring both feature sequences to two-dimensional grids, a mean-pooling-like operation combines each group of four adjacent VAE latents into one latent anchor. Let the ViT grid be \(H_u\times W_u\) and the pooled anchor grid be \(H_g\times W_g\). A ViT token at \((p_i,q_i)\) is proportionally mapped to an anchor and scored by cosine similarity:
The anchor grid dimensions refer to the grid after pooling. Flooring applies to spatial mapping, not the token budget. Each comparison uses the spatially corresponding local anchor, rather than matching against every VAE token or a single global image vector. Scores therefore relate to local structure and can incorporate backgrounds or object boundaries that matter for editing despite low textual salience.
A high similarity is only a surrogate for consistency with generation representations. Appendix D explicitly states that it is not an exact mutual-information estimate, nor an optimal importance measure with a reconstruction-error guarantee. Perturbing spatial correspondence tests whether the surrogate genuinely depends on location rather than merely borrowing an arbitrary collection of VAE vectors.
2. Balanced Budget Selection: find regional representatives before truncating or filling the actual budget
Global top-K alone may concentrate tokens in a few high-scoring regions. G2TR first chooses the highest-scoring ViT token associated with each nonempty anchor, forming a candidate set, then selects representatives by score up to the budget. The budget is the keep ratio times the original understanding-token count, rounded to the nearest integer and bounded by a minimum count and the original token count. The default keep ratio is 0.5 and the minimum count is 1.
Two cases must be distinguished. If the budget is smaller than the candidate-set size, only its K highest-scoring candidates are retained; coverage of every anchor is not guaranteed. If the budget exceeds the candidate-set size, all regional representatives are retained, and the remaining slots are filled once with the highest-scoring unselected ViT tokens. This is neither iterative equal allocation across regions nor mandatory retention of low-scoring anchors to guarantee full-image coverage.
The design makes consideration of each region's best representative a preference, not an absolute constraint. It reduces the tendency of direct global ranking to ignore regional distribution, while allowing some regions to be omitted under a tight budget. This also exposes a limitation: lower keep ratios can still compromise spatial coverage, and no mechanism guarantees preservation of every local detail.
3. Similar-Token Merging: shorten the sequence without directly discarding all remaining features
After selection, unselected tokens lose their independent sequence positions, but their information can enter retained representatives. Each unselected token finds the most similar retained token by cosine similarity. This matching occurs in understanding-feature space: it does not require membership in the same anchor and is not nearest-neighbor matching by two-dimensional coordinates.
Let \(\mathcal{S}\) be the retained set and \(\mathcal{A}_j\) the unselected tokens whose cosine nearest neighbor is retained token \(j\). The merging mechanism in the paper can be written as:
This is an average with a weight on the original representative, not similarity-weighted averaging of individual merged tokens. A representative with no assigned tokens remains unchanged; one with multiple assignments carries their averaged feature representation. Sequence length therefore stays at the budget K, without adding positions to recover information.
Averaging mitigates the feature loss of direct deletion, but can also dilute fine-grained distinctions. The component analysis reports benchmark-dependent effects when merging is used alone, with the strongest overall results obtained by combining it with balanced selection. Merging should not be interpreted as an operation that necessarily improves accuracy at every budget.
4. Sequence Layout Rebuilding: preserve the original model interface with a short visual sequence
Compressed features cannot simply be inserted into an <IMG_CONTEXT> span still sized for the original token count. G2TR forms the short sequence in the original order of retained indices, rather than sorting by importance. It replaces the understanding-side visual span, preserves text and image start/end markers, rebuilds input ids, attention masks, and position ids, and repads the batch.
Preserving original index order and rebuilding position ids operate at different levels: the former preserves the ordering of visual content, while the latter aligns text, visual spans, masks, and positions with the new total length. The paper does not require copying every old position id and does not modify the UMM attention operator.
All these operations occur after original ViT encoding and before LLM prefill. The LLM receives a shorter but valid input and need not expose internal attention maps for pruning. The original Flash-Attention kernel stays unchanged, and layer-internal dynamic pruning does not introduce additional KV-cache rearrangement. This explains the deployment advantage over some implementations that reconstruct explicit attention scores; it does not claim that other methods can never support Flash-Attention.
A Worked Example¶
The following numbers illustrate algorithmic branches, not an experimental sample from the paper. Suppose an image has 12 understanding tokens, with 4 nonempty pooled anchors corresponding to 3 tokens each, and a budget of 6. Choosing the highest-scoring token per anchor gives 4 representatives; a single fill step selects the 2 highest-scoring tokens from the remaining 8, producing 6 retained tokens.
The other 6 tokens each find their cosine-nearest retained representative and participate in averaging. Only 6 understanding tokens enter the rebuilt visual span; the instruction and image boundaries are preserved, and VAE generation latents follow the original pathway. If the budget becomes 2, the method selects the 2 highest-scoring representatives from the 4 regional candidates, rather than covering all 4 anchors with 2 tokens.
Loss & Training¶
G2TR introduces no additional loss, back-propagation, or fine-tuning; learned VAE generation representations supply an inference-time signal. Defaults are \(\rho=0.5\), \(K_{\min}=1\), and \(\lambda=1.0\). Baselines follow their original repositories' default parameters and pruning layers. Appendix C fixes the random seed at 42; maximum response length is 20 for non-thinking understanding tasks and 2048 for BAGEL thinking mode.
Appendix D presents preservation of generation-related mutual information as an ideal objective, while the actual method uses the similarity surrogate and merging. Under a locally Lipschitz decoder assumption, it discusses output perturbations controlled by token reconstruction error. This is not a measured lossless bound for the real models and cannot override observed editing-score reductions.
Key Experimental Results¶
Main Results¶
Experiments use BAGEL-7B-MoT, with 7B active and 14B total parameters, and InternVL-U, with 4B parameters, evaluating each once on a single A6000 48G GPU. Understanding covers MME, MMBench dev, MMVP, and RealWorldQA. Editing covers English queries from GEdit-Bench and IntelligentBench, with RISE-Bench additionally evaluated in thinking mode. The following table excerpts source Table 2; omitted baselines were also evaluated.
Relative average is the arithmetic mean of each benchmark score divided by the same model's Vanilla score, expressed as a percentage. Raw scores with different units cannot be directly averaged. This measure is also not the fraction of overall sample-level accuracy retained.
| Model | Method | Average visual tokens | MME | MMBench dev | MMVP | RWQA | Relative average |
|---|---|---|---|---|---|---|---|
| InternVL-U | Vanilla | 100% | 2057 | 80.5 | 46.0 | 48.8 | 100% |
| InternVL-U | IVC-Prune | 50% | 2020 | 78.4 | 39.0 | 46.0 | 93.7% |
| InternVL-U | G2TR | 50% | 2034 | 78.8 | 40.0 | 46.8 | 94.9% |
| BAGEL-7B | Vanilla | 100% | 2388 | 88.5 | 69.3 | 60.2 | 100% |
| BAGEL-7B | IVC-Prune | 50% | 2332 | 88.4 | 70.0 | 57.8 | 98.6% |
| BAGEL-7B | G2TR | 50% | 2318 | 88.5 | 70.0 | 59.0 | 99.0% |
On BAGEL, MMBench matches Vanilla and MMVP increases from 69.3 to 70.0, but MME drops from 2388 to 2318 and RWQA from 60.2 to 59.0. All four raw InternVL-U scores are below Vanilla. The paper's statement that the method โdoes not sacrifice understanding abilityโ should therefore be read as retaining strong capabilities, not eliminating accuracy loss.
Source Table 4 reports editing results: BAGEL Vanilla scores 7.36/6.83/6.52 on GEdit's G_SC/G_PQ/G_O, compared with 7.18/6.67/6.42 for G2TR; IntelligentBench is 44.9 versus 44.0. G2TR achieves a 98.0% editing relative average, better than the other compressed methods, but all these raw metrics decline. The source metric names are retained without guessing expansions or evaluation definitions absent from the cache.
Thinking results appear in Table 5: Vanilla scores 64.7/63.4/11.9/55.3 on MMVP/RWQA/RISE/IntelligentBench, compared with 66.0/63.4/9.8/54.3 for G2TR, whose relative average is 95.6%. The RISE drop shows that complex editing can remain more sensitive to compression. This is not the same evaluation combination as the non-thinking editing average of 98.0%.
The following table excerpts the raw efficiency measurements from source Table 3; all rows use BAGEL-7B. FLOPs ratio measures computation, not timed speedup.
| Method | KV cache (MB) | Average visual tokens | FLOPs ratio | Prefill FLOPs (T) | Decode latency (ms/token) |
|---|---|---|---|---|---|
| Vanilla | 42.02 | 100% | 100.00% | 4.1325 | 49.60 |
| FastV | 22.04 | 54% | 55.10% | 2.2769 | 48.04 |
| IVC-Prune | 22.06 | 50% | 74.09% | 3.0619 | 49.30 |
| G2TR | 22.06 | 50% | 51.64% | 2.1341 | 47.88 |
G2TR reduces KV cache from 42.02 to 22.06 MB, approximately the source table's 1.90ร reduction, and prefill computation from 4.1325 to 2.1341 T, approximately a 1.94ร reduction. Decode latency changes only from 49.60 to 47.88 ms/token, reported as approximately 1.04ร in the source table. The 1.94ร figure must not be interpreted as measured end-to-end or prefill-latency speedup.
Some baseline multiplier annotations in source Table 3 are inconsistent: for example, IVC-Prune is annotated as 1.37ร in the FLOPs-ratio column and 1.35ร in the prefill-FLOPs column. The table above preserves the raw measurements without silently harmonizing the authors' multipliers or deriving stronger speed claims from them.
Ablation Study¶
The following corresponds to source Table 7. Guidance signals are compared after ViT encoding, reducing pruning-location confounds in the comparison of guidance sources.
| Guidance source | MMVP | RWQA | IntelligentBench |
|---|---|---|---|
| Random | 65.3 | 48.2 | 40.6 |
| Attention-based | 68.7 | 55.8 | 42.6 |
| Text similarity | 68.0 | 55.2 | 41.7 |
| VAE latent (G2TR) | 70.0 | 59.0 | 44.0 |
VAE guidance exceeds random guidance by 4.7, 10.8, and 3.4 points across the three columns. This supports generation-side representations as a relevant additional signal, but the table alone does not establish independent contributions of individual components.
Source Table 6 further reduces the budget: at keep ratios of 50%, 37.5%, and 25%, G2TR scores 70.0, 63.3, and 60.0 on MMVP, and 44.0, 42.5, and 42.0 on IntelligentBench. At 25%, FastV/IVC-Prune score 54.7/56.0 on MMVP and 40.3/41.0 on IntelligentBench. G2TR retains a relative advantage, but absolute scores decline as the budget tightens.
Figure 5 also analyzes high-/low-score deletion, spatial-mapping perturbations, balanced selection, and merging. The cache contains its caption and textual conclusions, but not reliably readable complete component-bar-chart values; no numerical component-removal table is invented here.
Key Findings¶
- At the same post-encoding insertion point, VAE guidance outperforms attention, text similarity, and random guidance on all three understanding/editing metrics; the importance signal is a central part of the evidence.
- Shuffling or shifting token-to-anchor correspondence degrades performance, supporting local spatial consistency without establishing cosine scores as equivalent to information content.
- Compression before prefill reduces subsequent input-processing computation and explains efficiency differences better than final keep ratio alone. IVC-Prune and G2TR do not have identical prefill FLOPs despite both retaining 50% of tokens.
Highlights & Insights¶
- The generation branch becomes a reference for retaining understanding features, rather than only a component that outputs images. Reusing learned representations is lighter than training a scorer and extends the compression objective beyond textual answers to editing-condition quality.
- Regional representative selection and feature merging address spatial concentration and information deletion separately. They are not redundant: merging is not restricted to spatial partitions, so the combination provides both a coverage preference and cross-region redundancy recovery.
- Deployment friendliness follows from the insertion point. Rebuilding a short sequence before calling the original LLM preserves existing inference implementations more readily than requesting internal attention maps or dynamically rearranging caches.
Limitations & Future Work¶
- Applicability depends on separate encoders and comparable hidden features from both pathways. Unified encoders, video inputs, and more native cross-pathway interactions are not validated here. Additional VAE encoding for understanding tasks should also be included in deployment-specific cost accounting.
- A fixed budget does not distinguish simple question answering from complex editing; there is no full-image coverage guarantee when the budget is smaller than the anchor count. Task- or detail-adaptive allocation is a future direction, not an implemented module.
- Experiments use one GPU, a fixed seed, and one evaluation run, with no error bars or significance tests. Small MMVP gains and decode-latency changes should not be interpreted as established statistical advantages.
- Appendix D's conditional perturbation argument relies on a local Lipschitz assumption and an expanded reconstruction representation. It supplies neither measured decoder constants nor an empirical guarantee of lossless editing with shortened sequences.
- The BAGEL editing-hyperparameter sentence in cached Appendix C.3 is incomplete. Section 5.2 points quantitative editing results to Table 5, although the corresponding non-thinking data are actually in Table 4. This note distinguishes the tables by their contents and does not guess missing parameters.
- The abstract supplies a code URL, while checklist item 5 still says the code will be released after acceptance. The paper-provided URL is retained without online verification of its current release status.
Related Work & Insights¶
- vs FastV / W-FastV / PDrop: These primarily select understanding tokens using internal attention or its variants; G2TR uses local VAE consistency and completes compression before the LLM. Differences include both guidance objectives and insertion points, so efficiency differences cannot all be attributed to scoring quality.
- vs VSCAN / IVC-Prune: These use visual/image-text signals and RoPE/image-text signals, respectively, whereas G2TR incorporates generation-side information. IVC-Prune already ties G2TR at 70.0 on BAGEL MMVP in Table 2; the advantage concerns the multi-task combination rather than universal per-metric superiority.
- vs architectural sparsification such as UniMoD: Architectural methods adjust network depth or width, while G2TR preserves the network and shortens its input. They may be complementary, but combined quality, compute, and implementation costs require separate validation.
Rating¶
- Novelty: 4/5 โ Uses local generation-side representations to guide understanding-side compression, addressing UMM needs beyond question-answering-only pruning.
- Experimental Thoroughness: 3/5 โ Two models, multiple tasks, and guidance-source ablations provide useful coverage, but quantitative editing focuses on BAGEL and repeated runs and end-to-end timing are absent.
- Writing Quality: 3/5 โ The main pipeline is clear, with remaining table-reference errors, an incomplete hyperparameter sentence, and inconsistent multiplier annotations.
- Value: 4/5 โ Training-free compression preserves the original attention interface and provides a lightweight efficiency baseline for separate-encoder UMMs, subject to task-dependent accuracy loss.