Skip to content

Beyond Selection: Token Parameterization for Extreme Visual Token Compression

Conference: NeurIPS2026 Spotlight (according to the arXiv comment)
arXiv: 2609.35232
Area: VLM Efficiency
Keywords: visual token compression, discrete cosine transform, basis-coordinate embedding, orthogonal reparameterization, spatial residuals

TL;DR

Rather than selecting a few original patches, Braco constructs a compact visual interface using a DCT low-frequency backbone, basis-coordinate embeddings, budget-dependent coordinate organization, and sparse-pooled spatial residuals; at 576โ†’16 tokens, it retains a Vanilla-normalized aggregate score of 94.0 while reducing full-pipeline prefill latency to 40.59 ms.

Background & Motivation

Vision-language models (VLMs) typically encode an image into hundreds or thousands of patch tokens before passing them to a large language model (LLM). At extremely small budgets, otherwise effective pruning or merging becomes brittle: retaining a few local patches can discard text, small objects, or relational evidence in a low-scoring region. Methods such as PruMerge and DivPrune are lightweight, but importance ranking alone cannot guarantee that all task evidence survives the 4โ€“25-token regime tested here. Another approach uses queries or cross-attention to aggregate the full image into a few latent tokens, as in MQT-LLaVA, QueCC, and TokenPacker. These interfaces can improve scores, but their own computation, parameters, and alignment training can introduce new overhead.

The authors separate two questions: whether the retained coordinates contain useful information, and whether downstream models can easily learn to read them. An orthogonal transform followed by truncation changes the retained subspace; an orthogonal rotation within that subspace adds or removes no information, yet can change token correlations, scale distribution, and spatial organization. These are not interchangeable forms of โ€œbetter token selection.โ€ For example, DCT coefficients and the coarse grid obtained by applying inverse DCT to the same low-frequency block are information-equivalent, but projector and LLM training need not be invariant to this token-axis rotation.

The paper therefore does more than transfer JPEG-style low-pass coding to VLMs, and does not claim that low frequencies solve every visual task. A low-frequency backbone carries global structure, but transformed tokens no longer correspond to original positions, and truncation can remove local details. Coordinate identity, ease of readout, and detail compensation must be addressed together. Core Idea: treat extreme visual compression as an interface-design problem of choosing a retained subspace and then organizing its coordinates, using a structured DCT backbone for global information, basis-coordinate embeddings and budget-dependent orthogonal organization for alignment, and a small spatial residual branch for local evidence.

Method

Overall Architecture

Braco (Backboneโ€“residual + basis + coordinate) sits after the vision encoder and before the multimodal projector. It receives the original two-dimensional patch-feature grid and outputs compressed visual tokens under a strict fixed budget; the compressor itself does not read the user question or text prompt. The main branch passes through the โ€œDCT Low-Frequency Backbone,โ€ โ€œBasis-Coordinate Embedding,โ€ and โ€œBudget-Dependent Coordinate Organization.โ€ The parallel โ€œSpatial Residualsโ€ branch reads the original grid directly; the branches are concatenated, normalized, and projected to the LLM hidden dimension.

The total budget is divided between a square low-frequency backbone and several learnable residual pooling slots. The residual branch does not explicitly subtract a reconstructed low-frequency signal to obtain a high-frequency error; it learns to aggregate complementary evidence from the original spatial tokens. During training, answer-token cross-entropy updates the trainable interface. At inference time, the same visual data path produces the compressed interface without labels, diagnostic probes, or an additional teacher.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    X["Encoded image<br/>original patch grid"] --> B["DCT Low-Frequency<br/>Backbone"]
    B --> E["Basis-Coordinate<br/>Embedding"]
    E --> A["Budget-Dependent<br/>Coordinate Organization"]
    X --> R["Spatial Residuals"]
    A --> Z["Concatenate, normalize, project<br/>fixed-budget visual interface"]
    R --> Z
    Z --> L["LLM + text prompt"]
    T["Training answer supervision"] -.->|Cross-entropy| L

Key Designs

1. DCT Low-Frequency Backbone: concentrate information before fixed truncation

The parameterization matters more than treating DCT as another importance scorer. Each row of the original token matrix corresponds to a spatial location. An orthogonal basis re-expresses the entire grid along the token axis; a structured selection matrix retains a few transform coordinates; a final orthogonal matrix only reorganizes those retained coordinates. The central relation is:

\[ \mathbf{Z}=\mathbf{A}\mathbf{P}_{\mathcal{S}}\mathbf{U}_{\mathbf{B}}\mathbf{X}. \]

The basis and truncation set jointly determine the retained subspace, while the orthogonal matrix leaves it unchanged. An identity basis with fixed truncation merely keeps a few spatial positions; DCT with the same-shaped truncation retains global cosine modes. Braco applies two-dimensional DCT to each feature channel and keeps the top-left square low-frequency block. A backbone token is consequently a global weighted combination of the grid, not an original patch. The rule is fixed and image-independent: it is neither top-k selection of original patches nor per-image selection of the largest-magnitude frequency coefficients.

The rationale is that encoded visual features often exhibit short-range correlations, allowing DCT to concentrate substantial energy at low frequencies. However, the authors distinguish energy retention from task-information retention: a background can carry considerable energy, while small text can have little energy yet determine the answer. Energy retention is the expected squared feature norm after truncation divided by that of the original features. Task readability, under a locally linear task model, measures the fraction of weighted task-direction power remaining after projection onto the retained subspace. โ€œReadabilityโ€ here means accessible task information, not fluent generated text.

The two are combined into a compressibility diagnostic:

\[ \mathcal{C}(\mathbf{B};\mathcal{S})=\lambda\mathcal{E}(\mathbf{B};\mathcal{S})+(1-\lambda)\mathcal{R}(\mathbf{B};\mathcal{S}). \]

Energy measures generic signal distortion, while readability assesses whether the retained subspace matches task directions; the weight controls their relative importance. Frozen-feature linear probes on CelebA provide a controlled readability experiment, not the final VLM's multimodal accuracy. The energy advantage of DCT/Haar under structured truncation largely shrinks when per-image magnitude selection is allowed, showing that the basis and deployment rule must be considered together. KLT is only an oracle reference fitted to the diagnostic distribution. Deployed Braco does not use KLT or add these diagnostic functionals as end-to-end training losses.

2. Basis-Coordinate Embedding: give global coefficients stable identities, not fictitious patch positions

A DCT coefficient aggregates multiple spatial locations, so its frequency index cannot simply be interpreted as an original patch position. Flattening low-frequency coefficients directly into a sequence forces the downstream projector to learn which mode each entry represents. Braco adds a basis-coordinate embedding (BCE) on the retained transform lattice, providing consistent identity cues for the same frequency location across images. It is independent of the input image and text question, and identifies a basis coordinate rather than image content.

The implementation uses Polar Fourier features: frequency indices are converted into a normalized radius and angle, encoded with multiscale sine and cosine features, mapped to the visual feature dimension by a shared linear layer, and added to the coefficient through a learnable scalar gate. The radius is normalized using the original transform-grid scale, not residual locations. Input-independent does not mean parameter-free: the coordinate features are fixed, but their mapping and magnitude gate are trainable. Compared with an unconstrained learned position table, this encoding explicitly represents frequency direction and scale; compared with ordinary two-dimensional spatial sinusoidal encoding, it addresses frequency lattice points rather than original pixel positions.

3. Budget-Dependent Coordinate Organization: equal information does not imply equal readout or optimization

After BCE, Braco chooses an organization within the retained backbone. Vanilla uses coefficients directly as tokens; idct applies an orthogonal inverse DCT at the retained block size to obtain a coarse spatial grid; randrot is a random orthogonal control. This inverse DCT neither restores the original resolution nor recovers discarded high frequencies. It changes coordinates within the same low-frequency subspace. The original input information represented by the backbone remains unchanged, while cross-token correlations and geometric layout can change.

Learnability therefore differs from compressibility: with the subspace fixed, it concerns convergence and cross-modal alignment. The authors penalize off-diagonal energy in the token Gram matrix to discourage correlations, penalize diagonal deviations from the mean to discourage scale imbalance, and add geometric distance from the preferred coarse-grid organization. The statistical term is normalized by the squared backbone budget. Comparing coefficient coordinates with the coarse grid gives:

\[ \mathcal{L}_{\mathrm{learn}}(\mathbf{U}_{C};K_{b})-\mathcal{L}_{\mathrm{learn}}(\mathbf{I};K_{b})=\Delta_{\mathrm{st}}/K_{b}^{2}-\beta\rho_{C}. \]

The statistical gap measures the additional statistical penalty incurred by the coarse grid; the geometric mismatch measures the penalty for remaining in coefficient coordinates rather than that grid. A positive difference favors coefficient coordinates in the diagnostic, while a negative difference favors the coarse grid. This is a design heuristic comparing two candidates, not a proof of global optimality over all orthogonal rotations, nor a data-independent identity guaranteeing higher accuracy.

In the main setting, Braco uses vanilla below 16 backbone tokens and idct at 16. This is the backbone budget, not the total budget: c3s7 has 16 tokens in total but only 9 backbone tokens, so it still uses vanilla; c4s9 has 25 tokens in total and 16 backbone tokens, so it uses idct. A new encoder or distribution can use a small unlabeled calibration sample to estimate the statistical gap and choose the crossover. The appendix reports median LLaVA-Pretrain crossover budgets of 16, 64, and 36 for CLIP, SigLIP, and SigLIP2, respectively, demonstrating that โ€œ16โ€ is not universal.

4. Spatial Residuals: spend the compensation budget where local evidence is concentrated

The low-frequency backbone captures global structure but can miss local evidence. The argument is not that all high frequencies matter: a spatially sparse signal can spread over many coefficients in a transform basis incoherent with the spatial basis, so retaining a few more frequency coefficients may recover a small region inefficiently. Braco instead allocates additional budget to the spatial domain, where learnable pooling aggregates original patch features directly. The theoretical sparse-residual energy bound supports this motivation, but does not guarantee spatially sparse errors for every image or task.

A lightweight TokenLearner-style scorer produces logits over the entire original grid for each residual slot. Sparsemax converts them into weights with sparse support and unit sum, and each residual token is a weighted sum of original spatial tokens. Temperature is annealed during training to encourage sharper selection; the implementation also allows a normalized grid-gradient-energy bias when enabled. These slots produce learned aggregate tokens, not necessarily individual selected patches, and do not rank and copy original patches by top-k. The appendix does not specify a separate diversity constraint, so disjoint attention regions across slots cannot be assumed.

Spatial residuals are appended to backbone tokens, normalized to harmonize their scales, and passed to the projector. They are not free: the c3s7 stage breakdown reports approximately 1236.293 MFLOPs for the residual stage, compared with 56.623 MFLOPs for DCT and 38.928 MFLOPs for embeddings. The point is that the branch remains cheaper than a heavy query compressor while complementing the backbone. Performance gains cannot be attributed entirely to fixed DCT, and the complete Braco interface is not training-free.

A Worked Example

Consider the main 576โ†’16-token setting. The vision encoder outputs a 24ร—24 patch-feature grid. The main branch applies two-dimensional DCT, keeps only 3ร—3 low-frequency coefficients, and adds their Polar Fourier basis-coordinate embeddings. Since the backbone has only 9 tokens, it does not apply inverse DCT and retains vanilla coefficient organization.

In parallel, the spatial branch computes 7 sparse weight maps over the full 576-token grid, producing 7 residual tokens. The 9 backbone tokens and 7 residual tokens form a 16-token visual interface, which is projected and fed to the LLM with text. If the image contains a small label, the backbone might describe the scene and overall layout while a residual slot might aggregate the label region. This is a mechanism illustration, not a specific successful case reported by the paper, and it does not guarantee that the text will be captured.

Loss & Training

The authors retrain and evaluate all main-table compressors in a shared LLaVA-1.5-7B harness rather than combining published numbers from different papers. Each method and budget uses a budget-specific checkpoint and the two-stage LLaVA pipeline: pretraining on blip_laion_cc_sbu_558k and instruction tuning on llava_v1_5_mix665k, with one epoch per stage. The optimizer, learning-rate schedule, and weight decay follow the public LLaVA recipe.

Pretraining uses a total batch size of 256 for 2181 steps, training the compressor/interface and projector while freezing the vision encoder and LLM. Instruction tuning uses a total batch size of 128 for 5198 steps, training the interface, projector, and LLM. The learnability probe uses only the pretraining objective, a 0.99/0.01 train/dev split, and the first step at which dev cross-entropy falls below 2.50; it must be distinguished from the complete two-stage main-table experiment. Compressibility, readability, and learnability primarily diagnose and guide interface choices, and should not be described as three jointly optimized supervision losses.

Key Experimental Results

Main Results

The main table covers eight benchmark scores: GQA, MMBench EN/CN, MME All, POPE F1, ScienceQA, TextVQA, and MMVet. โ€œAcc.โ€ averages each benchmark score divided by its 576-token Vanilla score, then multiplies by 100. It is not ordinary classification accuracy and does not mean that 94% of questions were answered correctly. The following results retain this normalized-score convention:

\[ \mathrm{Acc.}=100|\mathcal{B}|^{-1}\sum_{b\in\mathcal{B}}s_{b}/s_{b}^{\mathrm{Vanilla}}. \]

Full-pipeline cost includes the vision encoder, compression/projector interface, and LLM prefill. It is neither standalone compressor cost nor full autoregressive answer-generation time. The main measurements use one GPU, one image, and a fixed-length prompt, averaging 100 runs after 20 warmup iterations on A100.

Total visual tokens Method Acc. Full-pipeline prefill FLOPs (T) Prefill latency (ms)
576 Vanilla 100.0 8.67 67.25
25 TokenPacker 94.2 1.38 38.40
25 Braco 95.2 1.37 41.03
16 QueCC 93.9 1.36 63.22
16 TokenPacker 93.7 1.26 37.87
16 Braco 94.0 1.25 40.59
9 QueCC 93.1 1.26 62.79
9 Braco 93.2 1.15 40.22
4 QueCC 91.4 1.20 64.17
4 Braco 91.2 1.09 40.97

Source: Table 1. Braco has the highest aggregate score among evaluated methods at 25/16/9 tokens, but does not win every task or minimize latency: TokenPacker is faster at 25/16 tokens. QueCC's native grid does not support a 5ร—5 output on the fixed 24ร—24 input, so its missing 25-token row does not establish that Braco beats it at that budget.

At 16 tokens, prefill latency falls from QueCC's 63.22 to 40.59 ms, approximately a 36% reduction. Relative to the uncompressed model, FLOPs fall from 8.67 to 1.25T. At 4 tokens, Braco is merely close to QueCC and scores 0.2 lower in the main table; the three-seed results are also 91.28ยฑ0.15 versus 91.35ยฑ0.16, not a win at every budget.

Ablation Study

Configuration Total tokens Acc. Compressor FLOPs (G) Compressor latency (ms)
c2s5: hybrid backbone and residuals 9 93.2 1.382 1.067
c3s0: backbone only 9 89.8 0.152 0.289
c0s9: residuals only 9 92.0 1.396 1.077
c2s0: smaller backbone only 4 85.7 0.152 0.288
c0s5: smaller residual-only interface 5 88.2 1.382 1.091

Source: Table 6. The number after c is the low-frequency block side length, not the backbone token count. The first three rows are matched 9-token comparisons: the hybrid scores 3.4 higher than backbone-only and 1.2 higher than residual-only. The last two rows help analyze branch complementarity but use different total budgets, so they are not matched-budget gains. Backbone-only is indeed cheaper; the observation that residual-only is not cheaper must not be generalized to all single-branch alternatives.

16-token method Measurement boundary Acc. Compressor latency (ms) Compressor FLOPs (G)
QueCC Pre-projector 93.9 17.864 109.504
Braco Pre-projector 94.0 1.073 1.389
TokenPacker Post-projector 93.7 1.004 15.271
Braco Post-projector 94.0 1.270 2.060

Source: Table 2. Only the matched pre-projector comparison gives Braco's 16.6-fold compressor latency and 78.8-fold FLOP advantages over QueCC; these factors do not apply to the full pipeline. Post-projector Braco has fewer FLOPs than TokenPacker but slightly higher latency, again showing that minimum FLOPs need not mean minimum implementation latency.

Key Findings

  • Coordinate organization is tested through convergence within a fixed subspace: at backbone budget 4, only vanilla reaches the dev cross-entropy threshold, at 1718 steps; at budget 16, idct/vanilla require 707/860 steps, and at budget 64 they require 237/590. These backbone budgets must not be conflated with the main table's total token budgets.
  • At c3s7, forcing idct gives final Acc. of 90.9 versus 94.0 for the selected vanilla organization; at c4s9, idct gives 95.2 versus vanilla's 93.2. Replacing BCE with learned or two-dimensional sineโ€“cosine embeddings lowers the score by 0.8โ€“0.9 points according to the main text.
  • Larger inputs do not uniformly retain 16 tokens: AnyRes multiview Vicuna-7B at 2880โ†’80 tokens reaches 98.1 with full-pipeline cost of 3.45T/44.36 ms, versus 40.57T/261.08 ms uncompressed. Qwen2.5-3B at 729โ†’16, 1024โ†’16, and 3645โ†’20 reaches 93.1, 92.6, and 90.8; QueCC scores 68.9 in the last setting. Each row is normalized to its corresponding uncompressed model, preventing interpretation as cross-model absolute accuracy.

Highlights & Insights

  • Separating what is retained from how it is presented explains more than replacing an importance-ranking criterion. Coordinate controls within the same low-frequency subspace prevent optimization improvements from being misattributed to additional information retention.
  • The global backbone and local aggregation are complementary rather than forcing one domain to carry all evidence. Matched 9-token ablations rule out both โ€œsimply more tokensโ€ and โ€œresidual-only is sufficientโ€ as complete explanations.
  • The diagnostic functionals have empirical support: the appendix reports average within-budget Spearman correlations of 0.93 between compressibility and MLP scores, and 0.88 between learnability and dev-loss AUC. They are experimentally supported design tools, not universal performance guarantees.

Limitations & Future Work

  • Evaluation focuses on LLaVA-family vision-encoderโ†’LLM pipelines and standard single-image benchmarks. Video, dense localization, irregular region features, and dynamic tokens require reassessing the structured basis, backbone/residual allocation, and coordinate crossover.
  • Complementarity between low frequencies and sparse local evidence may be insufficient for OCR-heavy or fine-grained localization tasks. DocVQA calibration demonstrates distribution-sensitive crossover points, not comprehensive end-to-end validation on those tasks.
  • Braco is a budget-specifically trained interface, not a training-free plugin. Key QueCC/TokenPacker comparisons use three seeds, while some others use one checkpoint. Small aggregate-score advantages, boundary-specific module costs, and fixed-prompt prefill measurements must retain their experimental conditions.
  • vs PruMerge / DivPrune: These methods prune or merge spatial tokens; Braco retains global transform modes in its backbone and supplements details with learned pooling. The trade-off is interface retraining rather than only lightweight inference-time selection.
  • vs QueCC / MQT-LLaVA / TokenPacker: These methods use learned interfaces, some with query conditioning; Braco emphasizes a prompt-independent fixed backbone and lightweight residuals. QueCC scores slightly higher at 4 tokens and TokenPacker has lower latency at some budgets, so learned interfaces are not universally dominated.
  • vs Fourier-VLM: Both use DCT low-frequency compression. Fourier-VLM presents an inverse-DCT coarse-grid interface, whereas Braco explicitly separates coordinate organization and adds BCE and spatial residuals. A transferable lesson is to test whether downstream models can read compressed coordinates easily, rather than measuring reconstruction error alone.

Rating

  • Novelty: 4/5 โ€” DCT is not new; the contribution is separating and validating retained subspaces, coordinate organization, and residual interfaces.
  • Experimental Thoroughness: 4/5 โ€” Shared retraining, module costs, ablations, larger inputs, and key three-seed results are provided, but architecture and task coverage remain limited.
  • Writing Quality: 4/5 โ€” The mechanism is coherent, but the Acc. name and backbone/total budget distinction require careful appendix reading.
  • Value: 4/5 โ€” A low-overhead, calibratable interface for extreme visual budgets; deployment gains depend on hardware and the prefill measurement scope.