Skip to content

Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/jhjangjh/TRAG
Area: Multimodal VLM / Model Compression
Keywords: Multimodal LLMs, Knowledge Distillation, Attention Guidance, Cross-modal Alignment, Entropy-driven Adaptive Weighting

TL;DR

Addressing the limitations of conventional multimodal distillation that rely on prompt-to-vision attention with uniform objectives, TRAG re-targets cross-modal attention distillation to the response generation phase (Response-to-Vision) and dynamically balances forward and reverse KL divergences based on teacher attention entropy, significantly outperforming prior baselines on VQA and compositional reasoning.

Background & Motivation

As multimodal large language models (MLLMs) rapidly scale up in both model size and training data, their prohibitive computational footprint creates an urgent demand for lightweight compression. Knowledge distillation (KD) serves as a primary paradigm for transferring capability from large teacher models to compact students. In text-only LLMs, logit-level output distribution matching works reliably; however, in multimodal scenarios, supervising solely on output token distributions fails to explicitly transfer how visual evidence is grounded during decoding. Because output tokens are inherently a byproduct of the model attending to visual inputs, an unguided student easily diverges in its internal cross-modal evidence allocation, inducing severe hallucinations and reasoning failures.

To provide direct internal supervisory signals, recent methods have begun exploring cross-modal attention distillation (e.g., Align-KD, CompoDistill). Nevertheless, existing frameworks almost universally supervise visual attention during the prompt processing phase (Prompt-to-Vision) and invest substantial manual engineering into aligning network layers across teachers and students with differing depths. Rigorous empirical analysis reveals that teacher-student attention similarity in the Prompt-to-Vision phase shows negligible correlation with downstream benchmark performance; in fact, standard SFT models often exhibit higher Prompt-to-Vision similarity to the teacher than distilled models despite trailing far behind in actual task capability. Conversely, attention similarity during the Response-to-Vision phase correlates strongly and positively with downstream performance. Furthermore, attention patterns across adjacent intermediate layers are highly redundant, whereas they display extreme variance across response tokens: function words attend diffusely to broad scene context, while entity and action words sharply focus on localized objects. Imposing a uniform global distance metric across all tokens is therefore fundamentally suboptimal.

The core angle of attack in this work is to anchor cross-modal supervision directly into the autoregressive decoding phase with token-level adaptive precision. The core idea is to shift cross-modal attention distillation entirely to the response generation phase (Response-to-Vision), bypass brittle layer-matching heuristics via intermediate-layer aggregation, and dynamically modulate forward KL (mean-seeking for broad coverage) and reverse KL (mode-seeking for sharp localization) based on the teacher's attention entropy, achieving token-specific cross-modal visual grounding alignment.

Method

Overall Architecture

TRAG is built upon a standard autoregressive MLLM distillation pipeline to enable a compact student model to faithfully mimic a large teacher's internal visual attention allocation at each decoding step. The framework first extracts the cross-modal attention matrix generated by the causal Transformer decoder, isolates the response token submatrix, averages attention over the most informative intermediate layers, and renormalizes it over the visual token span. Subsequently, for each response token, the Shannon entropy of the teacher's normalized attention distribution is computed to dynamically assign an adaptive weighting coefficient, blending asymmetric KL divergences for precise token-level supervision.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multimodal Input Sequence<br/>Visual Tokens + Prompt + Response"] --> B["Response-to-Vision Attention Isolation<br/>Extract Submatrix Restricted to Response Range"]
    B --> C["Intermediate Layer Aggregation & Renormalization<br/>Aggregate Depths in 0.3-0.6 & Project to Probability Simplex"]
    C --> D["Entropy-Driven Adaptive KL Regulation<br/>Dynamically Blend Forward & Reverse KL via Teacher Entropy"]
    D --> E["End-to-End Distillation Objective<br/>Joint Optimization with LM & Uni-modal Losses"]

Key Designs

1. Response-to-Vision Attention Isolation: Shifting Distillation Focus from Prompt to Response Generation

Prior attention distillation works (such as Align-KD and CompoDistill) supervise cross-modal attention predominantly over prompt tokens \(I_P = \{N+1, \dots, N+L_P\}\), capturing \(A_{P \to V}\). However, prompt processing only reflects initial comprehension and does not determine how visual cues are actively utilized during generative decoding. TRAG explicitly redirects the query index range to the response tokens \(I_R = \{N+L_P+1, \dots, N+L\}\), directly supervising the response-to-vision attention vector \(a_{l,i} = \{a_{l,i,j}\}_{j=1}^N\) for each generation step. This isolation ensures that distillation directly regularizes the student's active visual grounding during autoregressive token generation, directly mitigating visual detachment and hallucination.

2. Intermediate Layer Aggregation & Renormalization: Bypassing Brittle Layer Mapping and Eliminating Attention Mass Leakage

In heterogeneous teacher-student architectures with different layer depths, designing layer-wise mapping or grouping schemes is notoriously fragile. The authors observe that intermediate layers within relative depths \([0.3, 0.6]\) exhibit high internal redundancy while hosting the most intensive cross-modal grounding. Consequently, TRAG foregoes complex layer-selection heuristics and computes the arithmetic mean across the designated intermediate layers \(\mathcal{M}_S\) and \(\mathcal{M}_T\): $\(\bar{\mathbf{a}}_i^S = \frac{1}{|\mathcal{M}_S|}\sum_{l\in \mathcal{M}_S} \mathbf{a}_{l,i}^S, \qquad \bar{\mathbf{a}}_i^T = \frac{1}{|\mathcal{M}_T|}\sum_{l\in \mathcal{M}_T} \mathbf{a}_{l,i}^T\)$ Because self-attention also distributes probability mass over preceding textual tokens, the extracted visual vectors do not sum to one. To formulate a mathematically rigorous probability distribution for information-theoretic divergence, TRAG locally renormalizes the vectors over the visual span \(N\): $\(\tilde{\mathbf{a}}_i^S = \frac{\bar{\mathbf{a}}_i^S}{\sum_{j=1}^N \bar{a}_{i,j}^S + \epsilon}, \qquad \tilde{\mathbf{a}}_i^T = \frac{\bar{\mathbf{a}}_i^T}{\sum_{j=1}^N \bar{a}_{i,j}^T + \epsilon}\)$ This maps both teacher and student attention onto the same probability simplex for fine-grained distributional comparison.

3. Entropy-Driven Adaptive KL Regulation: Dynamically Balancing Coverage and Sharp Mode-Seeking via Asymmetric KL

Different response tokens serve disparate grounding roles: function words (e.g., "The") exhibit diffuse attention across the entire image to capture broad background context, whereas content tokens (e.g., "throwing" or "mound") exhibit sharply localized focus on specific regions. A fixed forward KL \(D_{KL}(\tilde{\mathbf{a}}_i^T \parallel \tilde{\mathbf{a}}_i^S)\) is mean-seeking and aggressively penalizes zero student probability on non-zero teacher regions, tending to over-smooth the distribution and dilute sharp evidence. Conversely, a reverse KL \(D_{KL}(\tilde{\mathbf{a}}_i^S \parallel \tilde{\mathbf{a}}_i^T)\) is mode-seeking and heavily penalizes student mass outside teacher modes, enforcing sharp concentration but risking neglect of broader contextual cues. TRAG resolves this by first calculating the Shannon entropy of the teacher's normalized attention distribution: $\(H(\tilde{\mathbf{a}}_i^T) = -\sum_{j=1}^N \tilde{a}_{i,j}^T \log \tilde{a}_{i,j}^T\)$ Using running lower and upper bounds \(H_{\min}\) and \(H_{\max}\) tracked via exponential moving averages (momentum \(\beta=0.995\)), the entropy is converted via min-max normalization into a bounded coefficient \(\lambda_i \in [0, 1]\): $\(\lambda_i = \mathrm{clip}\left(\frac{H(\tilde{\mathbf{a}}_i^T) - H_{\min}}{H_{\max} - H_{\min} + \epsilon}\right)\)$ This yields the final per-token adaptive cross-modal distillation objective: $\(\mathcal{L}_{\text{cross-modal}}^{(\text{TRAG})} = \sum_{i\in I_R} \left[ \lambda_i D_{KL}(\tilde{\mathbf{a}}_i^T \parallel \tilde{\mathbf{a}}_i^S) + (1-\lambda_i) D_{KL}(\tilde{\mathbf{a}}_i^S \parallel \tilde{\mathbf{a}}_i^T) \right]\)$ When attention entropy is high (diffuse context), \(\lambda_i \to 1\) and forward KL dominates to ensure comprehensive coverage; when entropy is low (sharp entity focus), \(\lambda_i \to 0\) and reverse KL dominates to enforce precise mode alignment.

Loss & Training

The framework follows a three-stage training schedule: 1. Distillation Pre-training (DPT): On LLaVA-Pretrain-558K, the vision encoder and student LLM backbone are frozen, and only the 2-layer GELU-MLP projection module is trained with the composite loss \(\mathcal{L} = \mathcal{L}_{\text{LM}} + \mathcal{L}_{\text{uni-modal}} + \mathcal{L}_{\text{cross-modal}}^{(\text{TRAG})}\); 2. Supervised Fine-Tuning (SFT): On LLaVA-Instruct-665K, the student LLM backbone is unfrozen, optimizing solely the standard language modeling objective \(\mathcal{L}_{\text{LM}}\); 3. Distillation Fine-Tuning (DFT): On LLaVA-Instruct-665K, the vision encoder remains frozen while both the projector and the student LLM are fully fine-tuned using the joint objective \(\mathcal{L}_{\text{LM}} + \mathcal{L}_{\text{uni-modal}} + \mathcal{L}_{\text{cross-modal}}^{(\text{TRAG})}\), where \(\mathcal{L}_{\text{uni-modal}}\) integrates the unimodal vision logit and relation distillation terms from LLaVA-KD.

Key Experimental Results

Main Results

Evaluation spans general Visual Question Answering (8 benchmarks) and fine-grained compositional reasoning (3 benchmarks) across diverse student backbones and model scales.

General VQA Benchmark Comparison (Avg6 excludes MMBCN and MMMU; Avg8 averages all 8 tasks):

Size Method LLM Backbone Samples GQA ScienceQA TextVQA MME MMB MMBCN POPE MMMU Avg6 Avg8
≤4B (Teacher) TinyLLaVA Qwen2.5-3B 1.2M 63.0 76.1 60.1 71.6 73.1 70.4 87.8 41.3 71.9 67.9
≤2B TinyLLaVA (SFT) Qwen2.5-1.5B 1.2M 61.7 70.9 58.4 70.4 70.6 64.0 86.0 39.2 69.7 65.1
≤2B LLaVA-KD Qwen2.5-1.5B 1.2M 62.5 71.6 59.7 70.0 71.0 66.6 86.7 35.8 70.2 65.4
≤2B TRAG (Ours) Qwen2.5-1.5B 1.2M 62.6 73.2 60.0 71.2 70.8 69.3 87.2 38.7 70.8 66.6
≤2B TinyLLaVA (SFT) Qwen1.5-1.8B 1.2M 60.9 65.2 47.7 61.3 57.8 56.2 83.4 34.8 62.7 58.4
≤2B CompoDistill Qwen1.5-1.8B 1.2M 61.2 66.5 53.5 67.0 64.5 63.0 85.5 34.1 66.4 61.9
≤2B LLaVA-KD Qwen1.5-1.8B 1.2M 62.3 64.7 53.4 69.1 64.0 63.7 86.3 33.6 66.6 62.1
≤2B TRAG (Ours) Qwen1.5-1.8B 1.2M 61.7 66.9 55.6 67.0 66.6 63.3 86.8 35.2 67.4 62.9
≤0.5B TinyLLaVA (SFT) Qwen2.5-0.5B 1.2M 58.6 61.2 47.9 62.3 59.3 52.4 85.1 30.4 62.4 57.1
≤0.5B LLaVA-KD Qwen2.5-0.5B 1.2M 59.8 60.6 52.0 64.7 61.3 57.0 86.4 28.3 64.1 58.7
≤0.5B TRAG (Ours) Qwen2.5-0.5B 1.2M 60.2 62.7 52.0 68.1 61.7 57.7 87.6 31.3 65.3 60.2
≤0.5B TinyLLaVA (SFT) Qwen1.5-0.5B 1.2M 56.7 60.9 46.6 59.4 55.2 51.8 83.9 31.9 60.4 55.8
≤0.5B CompoDistill Qwen1.5-0.5B 1.2M 59.1 61.1 48.2 63.8 59.6 54.9 85.7 34.0 62.9 58.3
≤0.5B LLaVA-KD Qwen1.5-0.5B 1.2M 59.6 60.6 49.9 64.5 60.1 55.5 85.9 30.2 63.4 58.2
≤0.5B TRAG (Ours) Qwen1.5-0.5B 1.2M 59.5 62.1 48.9 65.2 60.9 54.6 86.5 30.9 63.8 58.5

Compositional Reasoning (CR) Benchmark Comparison:

Backbone Model Size Method SugarCrepe BiVLC Winoground CR Avg
Qwen1.5 4B (Teacher) TinyLLaVA 87.3 93.2 70.1 83.5
Qwen1.5 1.8B TinyLLaVA (SFT) 74.8 83.8 62.6 73.7
Qwen1.5 1.8B KD (Logits) 76.1 85.3 61.5 74.3
Qwen1.5 1.8B LLaVA-KD 76.6 85.5 60.9 74.3
Qwen1.5 1.8B CompoDistill 81.7 86.5 62.2 76.8
Qwen1.5 1.8B TRAG (Ours) 83.1 89.1 67.2 79.8
Qwen1.5 0.5B TinyLLaVA (SFT) 52.8 56.5 50.2 53.2
Qwen1.5 0.5B LLaVA-KD 66.9 74.5 51.7 64.4
Qwen1.5 0.5B CompoDistill 72.1 78.6 56.6 69.1
Qwen1.5 0.5B TRAG (Ours) 73.4 80.2 57.3 70.3
Qwen2.5 3B (Teacher) TinyLLaVA 88.1 94.6 75.0 85.9
Qwen2.5 1.5B TinyLLaVA (SFT) 87.9 93.3 70.3 83.8
Qwen2.5 1.5B TRAG (Ours) 87.3 93.8 73.9 85.0
Qwen2.5 0.5B TinyLLaVA (SFT) 75.1 85.1 57.0 72.4
Qwen2.5 0.5B TRAG (Ours) 78.9 87.1 59.6 75.2

Ablation Study

1. Query Token Scope Ablation (Qwen2.5-0.5B):

Supervised Query Token Type VQA Avg (Avg8) CR Avg Insights
Prompt tokens only 58.6 69.2 Omits generative visual guidance, negligible gain on reasoning
Prompt + Response mixed 58.4 71.2 Forcing alignment on prompt tokens degrades VQA output
Response tokens only (TRAG Ours) 60.2 75.2 The critical alignment signal lies purely in decoding phase

2. Attention Matching Objectives Comparison:

Matching Objective Formulation VQA Avg (Avg8) CR Avg Performance Analysis
Cosine Similarity 58.7 72.5 Aligns direction but ignores local density concentration
MSE 58.2 74.4 Penalizes raw value gaps, decent on CR but suboptimal on VQA
Forward KL only 57.6 73.9 Mean-seeking induces over-smoothing and dilutes localized modes
Reverse KL only 58.1 73.4 Mode-seeking causes premature collapse, missing broad context
JSD (Jensen-Shannon Divergence) 57.9 71.6 Static symmetric blending fails to handle token heterogeneity
Entropy-Weighted KL (TRAG Ours) 60.2 75.2 Dynamically adapts to token roles, reaching clear superiority

Key Findings

  • Generative attention is the primary signal for MLLM distillation: Switching attention supervision from prompt to response tokens surges CR Avg from 69.2 to 75.2 (+6.0 points), disproving the common assumption that prompt-to-vision alignment is essential.
  • Adaptive asymmetric KL resolves token-wise grounding conflict: Single static metrics inevitably stumble due to token heterogeneity. Modulating via teacher attention entropy allows the student to reach an empirical layer-wise attention fidelity of 0.72 (cosine similarity at 30%-60% relative depth), drastically outperforming SFT (0.18) and CompoDistill (0.41).
  • Pronounced improvements in low-capacity regimes with positive scaling: Improvements are largest in the ultra-compact \(\le 0.5\text{B}\) scale (Qwen2.5-0.5B gains +3.1 in Avg8 and +2.8 in CR Avg over SFT). As teacher capacity grows (1.8B \(\to\) 4B), TRAG scales smoothly, effectively transferring compositional grounding capabilities.

Highlights & Insights

  • Counter-intuitive empirical debunking: The work challenges the conventional reliance on prompt-to-vision attention by rigorously proving that downstream performance correlates strongly with response-time attention similarity, but negligibly with prompt-time attention.
  • Elegant information-theoretic adaptation: Rather than adding extra neural prediction heads or tuning ad-hoc threshold hyper-parameters, TRAG directly leverages the teacher's intrinsic Shannon entropy to dynamically modulate between mean-seeking and mode-seeking KL regimes.
  • Reusable intermediate-layer aggregation: By computing arithmetic averages over the normalized 0.3-0.6 relative depth window, TRAG eliminates brittle manual layer-matching across different network depths.

Limitations & Future Work

  • Authors' admitted limitations: Extracting dense causal attention matrices across extensive sequence lengths increases peak GPU memory overhead during the distillation fine-tuning phase.
  • Identified scope constraints: Evaluations were performed primarily on single-image benchmarks. Dynamic attention guidance across multiple high-resolution image crops or continuous video frames remains unexamined.
  • Future directions: Extending token-adaptive asymmetric attention guidance to temporal video understanding, or coupling attention guidance with test-time compute (TTC) search algorithms.
  • vs LLaVA-KD: LLaVA-KD relies heavily on unimodal distillation (visual logits and relation matrices) without explicit cross-modal attention supervision; TRAG provides the missing response-time cross-modal guidance, offering synergistic orthogonal improvements.
  • vs CompoDistill: CompoDistill relies on complex layer grouping and fixed cosine similarity on prompt tokens; TRAG simplifies layer matching via intermediate depth aggregation and refines token-wise supervision using entropy-adaptive KL on response tokens.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Overturns the standard prompt-attention paradigm and introduces an elegant entropy-driven asymmetric KL objective.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive coverage of 8 general VQA and 3 compositional benchmarks across multiple Qwen model generations and scales, complemented by attention fidelity analyses.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous motivation, logical empirical deductions, and clean mathematical formulation.
  • Value: ⭐⭐⭐⭐⭐ Delivers an actionable, effective, and reproducible blueprint for compact MLLM distillation and internal attention alignment.