Skip to content

Stabilizing Ultra-Low-Bit Quantization of Multimodal LLMs via Global Bit Allocation

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Model Compression
Keywords: Multimodal LLM Quantization, Mixed-Precision Quantization, Block-Level Anisotropy, Global Bit Allocation, Post-Training Quantization (PTQ)

TL;DR

To prevent severe accuracy collapse in multimodal LLMs at ultra-low bit-widths (2–3 bits), GloBitQ uncovers a two-level output-channel anisotropy and outlier-sensitivity resonance at the Transformer block level, deriving a training-free global bit allocation and affine smoothing framework that matches full-precision performance within 2.1% at 3.01 bits while remaining completely stable at 2.01 bits.

Background & Motivation

Large vision–language models (VLMs) have demonstrated impressive multimodal reasoning and visual perception capabilities, but their immense parameter footprint and memory bandwidth demands impose severe bottlenecks on resource-constrained edge deployment. Post-training weight-only quantization (PTQ) provides an attractive compression avenue by directly converting floating-point weights into low-bit integers without expensive end-to-end retraining, effectively alleviating memory-bound constraints during autoregressive token generation. However, when pushed into the ultra-low-bit regime (2–3 bits), conventional uniform-precision PTQ methods suffer catastrophic accuracy collapse, frequently generating invalid tokens or diverging into non-numerical values.

The fundamental culprit behind this collapse lies in the severe anisotropy of the quantization loss landscape. Prevailing PTQ frameworks are universally built upon layer-wise reconstruction objectives, optimizing each linear layer in isolation. Under this layer-wise formulation, all output channels within the same layer share identical input covariance \(H_l = \mathbb{E}[X_l^\top X_l]\), leaving the objective mathematically blind to sensitivity differences across output channels. Uniform quantization consequently wastes precious bit budget on insensitive channels while leaving critically vulnerable channels under-protected. Although heuristic mixed-precision schemes attempt to distribute variable bit-widths, they rely on ad-hoc local proxies that lack global optimality across layers and modalities.

Crucially, when the reconstruction objective is elevated from an isolated layer to the complete Transformer block, downstream attention routing, SwiGLU gated activations, and layer normalizations break the layer-wise Kronecker symmetry, causing different output channels to induce vastly disparate block-level output errors. Furthermore, the few persistent input activation outliers multiplicatively interact with downstream channel propagation matrices—a phenomenon termed outlier–sensitivity resonance—greatly amplifying cross-channel sensitivity gaps. Core idea: Grounded in block-level reconstruction analysis, we uncover a two-level anisotropy across layers and output channels amplified by outlier resonance, formulating GloBitQ—a training-free framework that unifies power-law layer balance, rank–value hybrid channel scoring, and coupled affine smoothing into a closed-form global bit allocation scheme.

Method

Overall Architecture

GloBitQ operates within a training-free calibration pipeline under a constrained global bit budget (e.g., an average of 2.01 or 3.01 bits per weight parameter). It assigns each output channel of every linear layer across the entire model either a base precision \(b_\ell\) or a promoted precision \(b_h = b_\ell + 1\).

The overall pipeline proceeds from the unquantized VLM and a small set of unlabelled multimodal calibration data. First, the block-level Hessian propagation is analyzed to formulate the two-level sensitivity criterion. Next, per-channel mean absolute weight magnitude is leveraged as an \(O(d)\) proxy for Hessian trace computation. Using a single global balance factor \(\alpha\), GloBitQ applies power-law weighting across layers and rank–value hybrid normalization within layers to produce a cross-layer comparable global importance score. Sorting these scores globally yields a Top-\(\tau\) promoted bit map. Finally, this discrete bit allocation is frozen and combined with lightweight, learnable affine smoothing to reduce feature condition numbers, producing a hardware-compatible, stable low-bit VLM.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Full-precision VLM weights & calibration samples"] --> B["Block-Level Anisotropy & Outlier Resonance Analysis<br/>Symmetry breaking · Jacobian propagation · Outlier amplification"]
    B --> C["Unified Single-Hyperparameter Global Bit Allocation<br/>Power-law layer weighting · Rank-value hybrid scoring · Global Top-τ"]
    C --> D["Coupled Mixed-Precision Affine Smoothing<br/>Freeze discrete bit map · Joint affine optimization · Condition flattening"]
    D --> E["Output: Hardware-compatible ultra-low-bit VLM (2.01/3.01-bit)"]

Key Designs

1. Block-Level Anisotropy & Outlier Resonance Analysis: Breaking layer-wise output channel symmetry

Conventional PTQ relies on layer-wise mean-squared error \(\mathbb{E}[\|X_l \Delta W_l^\top\|_F^2]\), where each output channel weight vector \(w_{l,i}\) incurs an expected distortion \(\mathbb{E}[\Delta w_{l,i}^\top H_l \Delta w_{l,i}]\) governed by the same input covariance \(H_l = \mathbb{E}[X_l^\top X_l]\). This shared covariance renders all output channels within a layer indistinguishable to the optimization objective. GloBitQ lifts the objective to the entire Transformer block \(\mathcal{E}_{\mathrm{block}}(\Delta\theta) = \mathbb{E}_{X}[\|\mathcal{B}(X; \theta + \Delta\theta) - \mathcal{B}(X; \theta)\|^2]\), incorporating self-attention, SwiGLU gating, and LayerNorm. Under a Gauss-Newton first-order expansion, the perturbation of output channel \(i\) propagates through downstream layers via Jacobian \(M_l \in \mathbb{R}^{Td \times T m_l}\), yielding an effective channel-specific Hessian \(H_{l,i}^{\mathrm{block}} = X_l^\top [\Phi_l]_{ii} X_l\), where \([\Phi_l]_{ii} \in \mathbb{R}^{T \times T}\) represents the diagonal block of \(\Phi_l = M_l^\top M_l\). Because attention heads, nonlinear activations, and normalization paths treat channels unequally, \([\Phi_l]_{ii}\) varies dramatically across channel indices \(i\), shattering the layer-wise symmetry and producing heterogeneous channel sensitivities \(\sigma_{l,i}^{\mathrm{block}} = \mathrm{tr}(H_{l,i}^{\mathrm{block}})\).

Furthermore, the activation outlier phenomenon—where a tiny fraction of activation channels exhibit extreme magnitudes (\([H_l]_{jj} \ge \alpha_l \gg \beta_l\))—interacts multiplicatively with downstream propagation. The ratio between maximum and minimum channel sensitivities is amplified by an alignment factor \(\kappa_\Phi \gg 1\): $\(R_l^{\mathrm{out}} = \kappa_\Phi \cdot R_l^{\mathrm{iso}}, \quad R_l^{\mathrm{iso}} = \frac{\max_i \mathrm{tr}([\Phi_l]_{ii})}{\min_i \mathrm{tr}([\Phi_l]_{ii})}\)$ This outlier–sensitivity resonance establishes a two-level anisotropy: inter-layer variance in average sensitivity and severe intra-layer inter-channel variance, establishing the first rigorous theoretical foundation for per-output-channel mixed-precision quantization.

2. Unified Single-Hyperparameter Global Bit Allocation: Power-law layer weighting and rank–value hybrid channel scoring

Minimizing block reconstruction error under independent channel quantization dictates that channels with the largest sensitivity-range product \(\sigma_{l,i}^{\mathrm{block}} \cdot R_{l,i}^2\) should be promoted to \(b_h\). However, computing exact Hessian traces across millions of parameters is computationally prohibitive. GloBitQ discovers that the mean absolute weight magnitude \(I_{l,i} = \frac{1}{d_l} \sum_{j=1}^{d_l} |w_{l,i,j}|\) exhibits a Spearman rank correlation exceeding 0.8 with \(\sigma_{l,i}^{\mathrm{block}}\) across all Transformer layers, serving as an efficient \(O(d)\) proxy. Because magnitude distributions vary widely across layers (from uniform to heavy-tailed), a raw cross-layer comparison would be biased. GloBitQ develops a two-stage balancing formulation governed by a single scalar \(\alpha \in (0, 1]\).

At the inter-layer level, motivated by rate–distortion theory's sub-linear bit assignment on variance, a concave power-law transformation normalizes layer weights: $\(\omega_l(\alpha) = \frac{\mu_l^\alpha}{\sum_{k=1}^L \mu_k^\alpha}, \quad \mu_l = \frac{1}{m_l}\sum_{i=1}^{m_l} I_{l,i}\)$ At the intra-layer level, a rank–value hybrid score strikes an optimal bias–variance trade-off between distribution-free CDF ranking and magnitude-proportional sensitivity: $\(S_{l,i}(\alpha) = (1 - \alpha) \frac{\mathrm{rank}(I_{l,i})}{n_l} + \alpha \frac{I_{l,i} - I_l^{\min}}{I_l^{\max} - I_l^{\min}}\)$ The composite score \(G_{l,i}(\alpha) = \omega_l(\alpha) \cdot S_{l,i}(\alpha)\) is ranked across all channels in the entire model. Channels falling into the global Top-\(\tau\) quantile are assigned bit-width \(b_h\), while the remainder receive \(b_\ell\), provably approximating the block-level optimum.

3. Coupled Mixed-Precision Affine Smoothing: Bridging theoretical approximations and flattening conditioning

The global allocation assumes i.i.d. quantization noise across input dimensions and proxy monotonicity. To resolve residual coupling and suppress ill-conditioned curvature without complicating the search space, GloBitQ adopts a decoupled two-stage design: the discrete bit map \(\{b_{l,i}\}\) is computed once from original weights and locked; subsequently, a set of learnable per-layer invertible affine transformations is optimized with the mixed-precision quantizer embedded in the loop.

Optimizing affine transformations over 15 epochs on 512 calibration samples jointly suppresses the condition numbers of weight matrices and activation covariance. Because the precision assignment is fixed, the optimization avoids discrete-continuous combinatorial instability, and the learned affine transformations are absorbed into adjacent linear weights before deployment, incurring zero additional inference latency or storage overhead.

Loss & Training

  • Calibration Dataset: 512 randomly sampled image-text instruction tuning pairs from the COCO subset of ShareGPT4V.
  • Optimization Strategy: 15 epochs using Adam with a learning rate of \(5 \times 10^{-3}\), optimizing mean squared error (MSE) on block output activations.
  • Hyperparameter \(\alpha\): Set to 0.7 for Qwen2-VL-7B and Qwen2.5-VL-7B, and 0.5 for LLaVA-OneVision.
  • Mixed-Precision Ratios: 3.01-bit denotes 99% 3-bit and 1% 4-bit channels; 2.01-bit denotes 99% 2-bit and 1% 3-bit channels (\(\tau = 0.01\)).

Key Experimental Results

Main Results

GloBitQ is evaluated across six multimodal benchmarks: ChartQA, DocVQA, OCRBench, ScienceQA, SeedBench 2 Plus, and TextVQA on three state-of-the-art architectures (Qwen2-VL-7B, Qwen2.5-VL-7B, and LLaVA-OneVision) via VLMEvalKit.

Table 1: 3-bit Post-Training Quantization comparison across six multimodal benchmarks

Model & Method Granularity Bit ChartQA DocVQA OCRBench ScienceQA Seed2Plus TextVQA AVG
Qwen2-VL-7B
Full Precision (FP16) - - 83.16 93.80 86.50 85.70 69.10 84.30 83.76
GPTQ Per-Channel 3 58.84 74.55 61.50 65.13 50.37 73.49 63.98
GPTAQ Per-Channel 3 58.76 74.03 61.10 64.25 52.17 73.18 63.91
VLMQ Per-Channel 3 59.44 76.59 61.90 63.05 52.92 74.54 64.74
GloBitQ (Ours) Per-Channel 3.01 79.28 92.82 82.70 84.33 67.32 83.76 81.70
Qwen2.5-VL-7B
Full Precision (FP16) - - 86.32 94.91 88.40 72.68 69.43 85.31 82.84
GPTQ Per-Channel 3 62.44 89.33 77.70 79.39 66.58 78.38 75.63
GPTAQ Per-Channel 3 63.68 90.41 79.60 80.52 66.40 77.94 76.42
VLMQ Per-Channel 3 65.32 91.36 80.00 81.70 67.32 79.28 77.49
GloBitQ (Ours) Per-Channel 3.01 82.00 93.77 86.30 59.34 64.82 84.08 78.30
LLaVA-OneVision
Full Precision (FP16) - - 80.12 87.05 62.30 95.33 65.43 75.93 77.69
GPTQ Per-Channel 3 74.24 77.82 56.60 85.26 60.65 71.90 71.07
GPTAQ Per-Channel 3 75.72 78.07 58.30 85.97 61.84 72.69 72.09
VLMQ Per-Channel 3 74.56 78.21 59.20 86.37 61.05 73.66 72.17
GloBitQ (Ours) Per-Channel 3.01 78.52 85.17 62.10 94.69 64.86 74.65 76.66

Table 2: 2-bit Ultra-low-bit Post-Training Quantization comparison across six multimodal benchmarks

Model & Method Granularity Bit ChartQA DocVQA OCRBench ScienceQA Seed2Plus TextVQA AVG
Qwen2-VL-7B
Full Precision (FP16) - - 83.16 93.80 86.50 85.70 69.10 84.30 83.70
GPTQ (Per-Channel) Per-Channel 2.01 - - - - - - NAN
GPTQ (Group-128) Group-128 2 56.44 74.90 62.60 50.37 51.78 71.93 61.30
GPTAQ (Group-128) Group-128 2 56.08 72.57 61.30 59.70 51.25 67.98 61.40
VLMQ (Group-128) Group-128 2 55.32 75.76 62.50 62.32 53.80 74.80 64.00
GloBitQ (Ours) Per-Channel 2.01 59.44 79.15 66.60 68.22 53.22 73.35 66.60
Qwen2.5-VL-7B
GPTQ (Per-Channel) Per-Channel 2.01 - - - - - - NAN
GPTQ (Group-128) Group-128 2 57.72 78.79 70.80 55.77 48.40 74.36 64.30
VLMQ (Group-128) Group-128 2 59.68 73.20 69.30 62.72 57.27 65.88 64.60
GloBitQ (Ours) Per-Channel 2.01 64.80 81.17 74.10 48.78 54.89 74.69 66.40
LLaVA-OneVision
GPTQ (Per-Channel) Per-Channel 2.01 - - - - - - NAN
VLMQ (Group-128) Group-128 2 62.76 64.82 51.40 69.04 53.32 68.20 61.50
GloBitQ (Ours) Per-Channel 2.01 59.48 64.06 52.10 73.72 53.53 61.94 60.80
GloBitQ (Ours) Per-Channel 2.05 62.63 68.03 52.70 77.29 55.59 62.83 63.10

Ablation Study

Table 3: Ablation study on 2-bit quantization components for Qwen2-VL-7B-Instruct

Method Config Granularity Bit ChartQA Seed2Plus DocVQA OCRBench ScienceQA TextVQA AVG
GPTQ Group-128 2 52.64 58.04 56.90 54.73 40.79 68.31 53.50
FlatQuant Per-Channel 2 56.48 58.39 62.90 52.40 46.20 59.47 51.10
GloBitQ (Full) Per-Channel 2.01 59.44 79.15 66.60 68.22 53.22 73.35 66.60

Key Findings

  • Enormous Gains at Under 0.1% Memory Overhead: Allocating a mere 1% of channels to a higher bit-width (\(b_\ell + 1\)) elevates Qwen2-VL-7B average performance by +15.5% over uniform 2-bit FlatQuant (66.60 vs. 51.10), preventing complete collapse while adding negligible memory cost.
  • Near-Lossless 3-bit Quantization: On Qwen2-VL-7B, GloBitQ 3.01-bit achieves an average score of 81.70, staying within 2.06 points (2.1%) of the full-precision baseline (83.76) and outperforming GPTQ by +17.72% and VLMQ by +16.96%.
  • Sensitivity of Balance Factor \(\alpha\): As shown in the ablation curve on LLaVA-OneVision, low \(\alpha\) (0.2–0.3) leads to uniform, suboptimal precision distribution, whereas excessive \(\alpha\) (0.7–0.8) over-concentrates precision on extreme outlier channels at the expense of general reasoning; a moderate value (\(\alpha = 0.5 \sim 0.7\)) achieves the optimal trade-off.

Highlights & Insights

  • Theoretical Flaw in Layer-Wise PTQ Uncovered: Rigorously demonstrated that layer-wise reconstruction loss treats all output channels identically due to shared input covariance, proving that channel-differentiated mixed-precision is only justifiable when viewed through block-level error propagation and outlier resonance.
  • Decoupled Global Ranking with \(O(d)\) Proxy: Replaced intractable full-model Hessian trace computation with per-channel mean absolute weight magnitude (Spearman rank correlation \(> 0.8\)), uniting inter-layer power-law weighting and intra-layer rank–value scoring under a single hyperparameter.
  • High Practical Value for Edge VLM Deployment: Resolves numerical divergence (NAN) in per-channel 2-bit quantization without retraining, outperforming group-wise baselines that demand complex hardware grouping indices.

Limitations & Future Work

  • Lack of Modality-Specific Feature Granularity: The block-level analysis treats vision and language tokens uniformly within the sequence dimension without explicitly modeling the dynamic token density differences introduced by high-resolution visual encoders (e.g., SigLIP).
  • Linear Accumulation Assumption across Depth: Inter-layer error accumulation is approximated additively via power-law weighting; highly deep architectures (70B+) might exhibit non-linear error compounding across cascading blocks.
  • Future Directions: Extending the global allocation framework to weight-activation joint quantization (e.g., W2A4/W3A8) and engineering custom mixed-precision GEMM kernels to realize actual inference acceleration on edge hardware.
  • vs GPTQ / GPTAQ: Layer-wise greedy solvers treat output channels homogeneously and completely collapse (NAN) at 2-bit per-channel precision; GloBitQ provides stable, competitive 2-bit performance by globally reallocating bits to sensitive channels.
  • vs VLMQ: VLMQ introduces token-level weighting for multimodal models but retains uniform bit assignments; GloBitQ exploits cross-layer and cross-channel anisotropy to achieve double-digit gains.
  • vs FlatQuant: FlatQuant utilizes affine smoothing to reshape activations but suffers at 2 bits under uniform allocation; GloBitQ couples affine smoothing with global mixed-precision to recover +15.5% accuracy.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the first rigorous block-level output-channel anisotropy and outlier resonance theory, deriving an elegant, closed-form global allocation scheme.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation on three modern VLMs across six standard multimodal benchmarks, including extensive 2-bit/3-bit comparisons and ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical derivations, coherent narrative progression from layer-wise limitations to block-level formulation, and clean visualizations.
  • Value: ⭐⭐⭐⭐⭐ Solves a critical bottleneck in ultra-low-bit VLM deployment without requiring backpropagation or architectural modifications.