Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?¶
Conference: NeurIPS2026
arXiv: 2609.31140
Area: Vision-Language Model Reasoning
Keywords: reasoning vectors, activation intervention, cross-modal transfer, representation matching, chain-of-thought
TL;DR¶
LIFT extracts the hidden-state difference at identical answer tokens with and without a reasoning trace from a base large language model, then injects it into the corresponding vision-language model's language layers without updating the backbone, raising six-benchmark average accuracy from 66.6/67.8 to 68.3/69.0 with static intervention and 69.1/69.4 with vector adaptation.
Background & Motivation¶
A vision-language model (VLM) typically adds a vision encoder and cross-modal projector to a large language model (LLM), followed by multimodal training. Although the language backbone comes from an existing LLM, its reasoning behavior may not remain intact: a model can formulate the right relationships but later copy a number incorrectly, omit a condition, or answer a different target. The paper observes that the base LLM can be more reliable than its corresponding VLM on some text-only questions. This motivates a more specific question than simply training on more reasoning data: can reasoning-related representations already present in the base LLM become useful again inside the VLM?
Existing multimodal chain-of-thought (CoT) supervision, verification, and reinforcement learning mainly change the target VLM's training or decoding behavior. This paper instead uses activation intervention to compress the internal representation change induced by explicit reasoning into a reusable direction. Simply averaging reasoning-token states could mix question content, trace length, and output format, making it difficult to distinguish representing an answer from organizing that answer after reasoning. The authors therefore hold the question and final answer fixed, vary whether a reasoning trace precedes the answer, and compare states at the same answer tokens.
This comparison does not encode an entire CoT as a decodable program. It uses a task-level average difference as an intervention tool and tests whether that direction reliably improves subsequent generation. Core Idea: extract the reasoning-conditioned shift in answer representations from successful base-LLM examples and transfer it to the corresponding VLM's language layers; when needed, learn only the vector rather than retraining the backbone.
Method¶
Overall Architecture¶
LIFT has two kinds of input: offline support examples containing questions, model-generated reasoning traces, and correct answers; and ordinary text-only or image-text questions at inference time. It first performs answer-state difference extraction, then either directly uses static vectors or applies learnable vector adaptation, and finally changes the target VLM's autoregressive generation through layer-selected language intervention.
The static variant requires no parameter optimization, but still uses correctness-based support filtering and development-set layer selection, so it is not entirely label-free. The learnable variant likewise updates neither the VLM, vision encoder, nor projector; the training variable is only a vector injected into a language layer. Neither variant appends support examples to test prompts, and both retain the baseline's zero-shot CoT inference setup.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Offline support examples<br/>question, trace, correct answer"] --> B["Answer-state<br/>difference extraction"]
B -->|Learnable variant| C["Learnable<br/>vector adaptation"]
B -->|Static variant| D["Layer-selected<br/>language intervention"]
C --> D
E["Development-set accuracy"] -.->|Offline layer selection| D
F["Inference-time text<br/>or image-text question"] --> D
D --> G["Generated reasoning<br/>and final answer"]
Key Designs¶
1. Answer-state difference extraction: isolate the representation change induced by reasoning while fixing the answer
For each support example, the source model first generates a trace under a step-by-step reasoning prompt, and only examples with correct final answers are retained. Two teacher-forced forward passes follow: the Reasoner path contains the question, reasoning trace, and final answer; the Solver path contains only the question and the same final answer. Teacher forcing means both paths receive predetermined tokens rather than freely generating potentially different answers. Extraction aligns answer tokens, not the absolute positions of the two sequences; neither question tokens nor reasoning tokens enter the vector average.
At each language layer, differences are first averaged across an example's answer positions and then across support examples. Longer answers therefore do not receive greater weight solely because they contain more tokens. The following equation captures this two-stage average:
Here, \(\mathcal{C}_{R,i}=(q_i,t_i,a_i)\), \(\mathcal{C}_{S,i}=(q_i,a_i)\), \(\mathcal{A}_i\) denotes the final-answer token set, and \(N\) is the number of support examples included in the average. Each layer independently yields a task-level vector. It is neither a reasoning trace specific to a test question nor a set of reasoning tokens inserted into the model.
The authors compare three sources. LLM-derived vectors come from successful text-only examples of the base LLM; VLM-derived vectors come from successful text-only examples of the target VLM; VLM-MM vectors come from image-text support examples of the target VLM. All three use language-side answer states, so a multimodal source does not mean extracting or injecting vectors inside the vision encoder.
The source model and data distribution must be distinguished. Text-only mathematical vectors come from GSM8K and serve GSM8K, MathVista, and MathVision; CommonsenseQA vectors also serve ScienceQA, while StrategyQA uses its own training data. VLM-MM instead uses support data from the corresponding multimodal task. Although LLM-derived and VLM-derived follow the same extraction protocol, their correctly answered subsets can differ, leaving sample composition as a factor in interpreting the source-model advantage.
Fixing the answer avoids mistaking a difference between two answers for a reasoning direction, but does not completely remove effects of context length, positional changes, or reasoning phrasing. The extracted vector is therefore an approximation of a reasoning-conditioned representation shift, not a proven pure, universal, or independently executable reasoning capability.
2. Learnable vector adaptation: retain the directional initialization and optimize only the injected vector
Static transfer assumes that a base-LLM representation direction remains usable in its corresponding VLM. Multimodal training may nevertheless change the appropriate location and magnitude. LIFT initializes a learnable vector from the extracted direction and adapts it on the support set: the Reasoner supplies reasoning-conditioned answer representations, while the Solver receives the vector at the selected layer, and training brings the resulting representations closer.
The objective matches support-set averaged answer representations. It neither requires the Solver to reproduce the full trace token by token nor uses language-model cross-entropy to learn new answers. The trainable object is a vector parameter, not backbone weights or an additional reasoning network. Because optimization also learns its magnitude, the learnable variant does not use a separate static scaling coefficient.
\(\boldsymbol{z}_{l}^{\mathcal{R}}\) is the Reasoner's support-set averaged answer representation, and \(\boldsymbol{z}_{l}^{\mathcal{S}}\) is the corresponding Solver average after vector injection. This loss constrains average representation proximity; it does not guarantee correct reasoning on every question. Final accuracy and control experiments must establish its practical usefulness.
Meaningful initialization is central to this design. LLM-derived initialization outperforms random initialization, whereas VLM-derived initialization can even bring InternVL2.5's average below baseline. Adding a small number of trainable parameters alone is therefore insufficient to explain the gains, and adaptation cannot be assumed to improve vectors from every source.
3. Layer-selected language intervention: apply the offline vector as a persistent shift during generation
An activation direction has different effects at different depths. The authors test candidate language layers individually on a development subset of 100 randomly sampled training examples for each benchmark, select the best-performing layer, and fix it for full evaluation. All sources and static or learnable settings follow the same selection protocol. This neither chooses layers separately using each test answer nor injects the entire layer-wise collection simultaneously.
At the selected layer, the static variant adds the extracted vector to language-token hidden states, while the learnable variant uses the optimized vector. The VLM still generates its own CoT and answer; it does not call the base LLM to construct a new vector for every test question or retrieve support demonstrations into the prompt:
The left expression is static intervention, and the right expression is learnable intervention. The main experiments fix static strength at \(\mu=1.0\). Intervention occurs on the language side during autoregressive generation rather than editing the output after the final answer has been produced. Extraction reads only answer tokens, but deployment acts during generation, allowing the intervention to affect intermediate reasoning rather than merely the final answer.
A Worked Example¶
Consider an image-text mathematics question from MathVision. Offline extraction can first build task-level vectors from GSM8K examples correctly solved by the base LLM. The two teacher-forced paths for each example receive the same answer and differ only in whether the generated reasoning trace precedes it. Averaging produces layer-wise directions, and a MathVision development subset selects the injection layer.
At test time, the VLM reads the actual image and question, without support examples in its prompt. The static variant adds the mathematical vector at the selected layer during generation; the learnable variant adds the vector adapted through support-set representation matching. In the paper's target-score case, the LLM-derived intervention preserves repeated-hit possibilities and distinguishes scoring combinations from distinct total scores, producing 9. The VLM-derived intervention still treats the number of combinations as the number of distinct scores. This is case-level process evidence, not a claim that every visual error is repaired.
Loss & Training¶
The default procedure samples 500 examples per support source and then retains those correctly answered by the source model, so the final extraction count should not be assumed to be exactly 500. The corresponding source model generates the traces. Support data is used only for offline extraction or adaptation, not in the final test set or as in-context demonstrations.
Learnable vectors are trained with AdamW for 6 epochs at learning rate \(1\times10^{-4}\), batch size 1, and gradient accumulation over 4 steps. Weight decay is \(1\times10^{-3}\), warmup ratio is 0.1, gradient clipping has maximum norm 1.0, and training uses 16-bit mixed precision. The sole objective is representation matching; the entire VLM backbone and vision encoder remain frozen.
The baseline and LIFT share zero-shot CoT prompts, answer parsers, and deterministic decoding: beam size 1 with sampling disabled, at most 512 new tokens for text-only tasks and 4096 for multimodal tasks. The default seed for support sampling, layer-selection sampling, and adaptation is 42. Freezing the backbone reduces training requirements, but extraction, development-set layer selection, and vector learning still incur offline costs.
Key Experimental Results¶
Main Results¶
The targets are Qwen2.5-VL-7B-Instruct and InternVL2.5-8B, paired with Qwen2.5-7B-Instruct and InternLM2.5-7B-Chat as their base LLMs. The table below reports accuracy (%) from the paper's Table 1. Averages are arithmetic means across six benchmarks, not weighted by example count. MVista, MVision, CSQA, SQA, and SciQA denote MathVista, MathVision, CommonsenseQA, StrategyQA, and ScienceQA, respectively.
| Model | Config | GSM8K | MVista | MVision | CSQA | SQA | SciQA | Avg. |
|---|---|---|---|---|---|---|---|---|
| Qwen2.5-VL | Baseline | 82.2 | 68.5 | 25.0 | 77.3 | 58.6 | 87.8 | 66.6 |
| Qwen2.5-VL | LLM-derived | 84.2 | 69.1 | 26.6 | 78.0 | 63.5 | 88.1 | 68.3 |
| Qwen2.5-VL | VLM-derived | 83.0 | 70.7 | 25.7 | 77.7 | 60.3 | 87.8 | 67.5 |
| Qwen2.5-VL | LLM-learnable | 85.6 | 71.8 | 27.3 | 78.3 | 62.4 | 88.9 | 69.1 |
| InternVL2.5 | Baseline | 77.1 | 62.8 | 22.7 | 82.3 | 65.6 | 96.5 | 67.8 |
| InternVL2.5 | LLM-derived | 78.1 | 64.0 | 25.0 | 83.1 | 67.1 | 96.8 | 69.0 |
| InternVL2.5 | VLM-derived | 77.9 | 63.0 | 23.4 | 82.6 | 65.9 | 96.7 | 68.3 |
| InternVL2.5 | LLM-learnable | 78.0 | 64.6 | 25.3 | 83.2 | 67.6 | 97.4 | 69.4 |
Static LLM-derived intervention raises the two model averages by 1.7 and 1.2 percentage points, while learnable intervention raises them by 2.5 and 1.6 points over baseline. These are matched-protocol intervention comparisons, not claims of SOTA against all existing multimodal reasoning methods. The corresponding 4-shot CoT averages are 67.4/68.5, and prompt-changing few-shot methods also have different resource requirements from LIFT.
Ablation Study¶
This table combines the paper's Table 3 and Appendix Table 11, retaining six-benchmark averages that distinguish useful directions from gains attributable merely to more parameters or perturbations. Random initialization still undergoes vector learning; same-norm random vectors instead replace the static direction without learning. These are different controls.
| Config | Qwen2.5-VL average accuracy (%) | InternVL2.5 average accuracy (%) | Note |
|---|---|---|---|
| Baseline | 66.6 | 67.8 | No vector intervention |
| Same-norm random vector | 66.4 | 66.1 | Preserve static perturbation magnitude, replace direction |
| LLM-derived | 68.3 | 69.0 | Static LLM-derived direction |
| LLM-learnable | 69.1 | 69.4 | Adaptation from LLM-derived initialization |
| Random initialization, LLM-learnable control | 66.7 | 68.3 | Random initialization under the corresponding learning setting |
| VLM-learnable | 67.2 | 66.6 | Adaptation from VLM-derived initialization |
Key Findings¶
- The LLM-derived advantage is primarily an across-task average trend, not a universal per-task win. On Qwen2.5-VL's MathVista, static VLM-derived reaches 70.7 versus 69.1 for LLM-derived; on InternVL2.5's GSM8K, static LLM-derived at 78.1 also slightly exceeds the learnable variant at 78.0.
- VLM-MM is reported only in a two-task comparison covering MathVision and ScienceQA. Its averages of 57.0/60.6 fall between LLM-derived at 57.4/60.9 and VLM-derived at 56.8/60.1, and cannot be compared directly with six-task averages. InternVL2.5's VLM-MM ScienceQA score is 96.4, below the baseline's 96.5, so average improvement does not imply improvement everywhere.
- Appendix paired bootstrap analysis uses 10,000 resamples. Across modelโbenchmark pairs, static LLM-derived gains 1.45 percentage points over baseline with a 95% CI of [0.45, 2.45], and 0.75 over VLM-derived with a CI of [0.15, 1.35]. This does not establish separate significance on every dataset.
- On GSM8K, average diagonal cosine similarity between cross-layer vectors is 0.866 for InternVL2.5 and 0.531 for Qwen2.5-VL. This supports partial structural correspondence at matching depths, but does not identify a unique reasoning mechanism represented by the vectors.
Highlights & Insights¶
- Aligning answer tokens rather than averaging reasoning tokens gives the representation difference a precise operational definition. It makes the influence of reasoning context on answer representations controllable while retaining the qualification that this is an approximation.
- Text-only mathematics and commonsense support data can improve image-text tasks, indicating that some benefits arise from language-side computation and condition preservation. This does not require an annotated trace for every test image, nor does it imply that perception bottlenecks have disappeared.
- Same-norm random controls and random-initialization controls separately test direction and learning initialization. Together they strengthen evidence for intervention usefulness, but do not establish multimodal training as the sole cause of reasoning degradation.
Limitations & Future Work¶
- The experiments cover only two model families with corresponding base LLMs. Applicability to heterogeneous language backbones, other alignment procedures, or stronger reasoning models remains unclear.
- Correctness filtering and development-set layer selection depend on answer information. The 500 examples are a sampling budget, not the actual correctly answered support count for every source. Future work should control shared correct-example subsets, support size, and trace length to separate source-model effects from data composition.
- MathVision has only 304 evaluation examples; the other sizes are GSM8K 1319, MathVista 1000, CommonsenseQA 1221, StrategyQA 687, and ScienceQA 4241. Macro-average confidence intervals and a trend under one additional seed cannot replace sufficient stability analysis for each task.
- Appendix A's single-head attention decomposition motivates an additive shift; it is not a capability-recovery theorem for a full multilayer Transformer. The explanation that alignment changes reasoning directions should be treated as a hypothesis, not causal proof.
- Intervention does not directly modify the vision encoder and cannot guarantee correction of recognition or visual grounding errors. Combining grounding verification with condition-preservation diagnostics is a possible extension. Appendix F in the cached text mainly retains case observations, so it does not support a claim that the complete traces inside the figures were checked.
Related Work & Insights¶
- vs Task Arithmetic: Task arithmetic generally operates on parameter differences, whereas LIFT operates in activation space without merging base-LLM and VLM weights. It suits lightweight intervention and source comparison, but does not directly establish parameter-level transfer effects.
- vs Contrastive Activation Addition / latent CoT vectors: These methods also steer behavior with activation directions. LIFT emphasizes fixed-answer ReasonerโSolver extraction and transfer from a base LLM to its corresponding VLM, rather than a general claim that vector addition enables reasoning.
- vs multimodal CoT supervision / MM-Verify / reasoning-oriented RL: These methods train or verify explicit reasoning behavior. LIFT constructs an internal shift using offline support data while retaining the inference prompt. The paper offers no comprehensive matched-training-budget comparison; the approaches are more plausibly complementary than interchangeable.
Rating¶
- Novelty: 4/5. Answer-state differences are used for base-LLM-to-VLM reasoning transfer with a focused research question, although activation steering has precedents.
- Experimental Thoroughness: 4/5. Six tasks, two models, random controls, and confidence intervals provide useful coverage; source-sample confounding and per-task stability need further study.
- Writing Quality: 4/5. Extraction and evaluation protocols are clear, but the representation-change interpretation exceeds the causal evidence, and case-figure details require consulting the original paper.
- Value: 4/5. The method offers lightweight reasoning enhancement with a frozen backbone and an experimental handle on language-representation changes before and after multimodal training.