The Alignment Illusion in Multimodal Large Language Models¶
Conference: NeurIPS2026
arXiv: 2609.30210
Area: Multimodal VLM
Keywords: Representation Similarity, Weight-Induced Alignment, Principal-Angle Gap, Visual Intervention, Task Relevance
TL;DR¶
Using norm-matched noise and irrelevant images in 13 multimodal large language models, this paper shows that shared MLP weights can produce high similarity and introduces the principal-angle gap to distinguish one-directional collapse from multi-directional structure, while emphasizing that geometry alone cannot establish the use of question-relevant visual content.
Background & Motivation¶
A common approach to analyzing multimodal large language models (MLLMs) compares visual-token and text-token hidden states layer by layer: rising CKA, SVCCA, or subspace similarity is interpreted as visual information becoming integrated into language representations. These tools were originally used to compare representations from different networks. Projector-LLM architectures instead concatenate visual and text tokens into one sequence and process them with the same attention and MLP weights. Similar representations can therefore arise either from meaningful cross-modal interaction or from the same weights imposing a common geometric bias on arbitrary inputs.
Ordinary image-question evaluation does not cleanly distinguish these explanations because its images usually have both natural structure and task relevance. The paper first replaces visual tokens at the projector output with norm-matched Gaussian noise: if similarity reflects content use, removing content should damage both accuracy and alignment. Accuracy does fall sharply, but several scalar measures remain high. A structured yet irrelevant image then separates the preservation of visual structure from the preservation of task-relevant content, testing whether even a more structure-sensitive geometric measure can be misread as evidence of understanding.
This is an audit of the interpretive validity of internal alignment measures, not a new training algorithm for improving question-answering accuracy. Core idea: combine controlled visual interventions, shared-weight mechanism analysis, and the principal-angle spectrum to identify the source of similarity, then report geometry and task performance separately rather than treating high alignment as proof of content-level interaction.
Method¶
Overall Architecture¶
The study examines released MLLMs consisting of a vision encoder, a projector, and a language model. The encoder extracts image features, the projector maps them into the language-model embedding space, and the resulting tokens pass through Transformers together with the question text. The authors extract the two token groups at each layer and compare their geometry with final multiple-choice accuracy.
The analysis proceeds through Visual-Stream Interventions, Shared-Weight Attribution, Principal-Angle Gap, and Task-Evidence Calibration: remove or mismatch content, locate the source of high similarity, and establish what the spectrum-based diagnostic can and cannot identify. The diagram represents the experimental analysis, not a new network. All main experiments perform inference with existing weights; linear probes are trained separately with category labels and do not provide training supervision to the MLLM.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Image and question<br/>Released MLLM"] --> B["Visual-Stream<br/>Interventions"]
B --> C["Shared-Weight<br/>Attribution"]
C --> D["Principal-Angle Gap"]
D --> E["Task-Evidence<br/>Calibration"]
E --> F["Interpret geometry and<br/>accuracy separately"]
G["Category labels<br/>Independent probe training only"] -.-> E
Key Designs¶
1. Visual-Stream Interventions: control visual content, structure, and scale separately
Orig uses the original projector output of the image paired with the current question. Noise independently samples isotropic Gaussian vectors and rescales each to the norm of the corresponding Orig token. It removes visual content and directional structure rather than merely reducing input magnitude. Intervening after the projector preserves the visual-token interface and avoids conflating vision-encoder failure with geometry inside the language model. Irr recomputes the projector output from another question's image, using the same fixed mismatched pairing across models. It retains natural image structure but does not support the current question. Irr norms are empirically close to Orig norms, not forcibly rescaled token by token as in Noise.
Two auxiliary conditions constrain the interpretation. Shuf permutes Orig's visual-token order while retaining token content. Text removes all visual tokens and provides a language-prior baseline. The same 1,000 MMBench questions are used throughout all models and conditions, so comparisons concern interventions on identical questions rather than dataset changes. Crucially, Noise and Irr are different kinds of corrupted visual input: the former lacks natural visual structure, whereas the latter has structure without task relevance.
The authors also gradually restore original tokens by linear interpolation, checking whether accuracy and the measures change together:
Here \(Z\) is the norm-matched noise described above. The default grid has 11 points at intervals of 0.1, but Appendix Table 11 explicitly reports only 6 points at intervals of 0.2 for a subset of models. The complete 11-point experiment should not be attributed to every model. Norm matching at the noise endpoint also does not imply that intermediate linear mixtures preserve the original norm.
2. Shared-Weight Attribution: locate common output directions behind high similarity
The attribution does not jump directly from shared weights to a conclusion; it progressively tests alternative explanations. Comparing fixed projector visual outputs with layer-wise text states yields only moderate leading principal-angle cosines, close to the random reference. Near-unity visual-text similarity requires both streams to pass through the language model. The authors then bypass either the MLP or attention at matched target layers, record the maximum downstream similarity change, and average over target layers. The MLP effect is larger in all 13 models. This localizes the main source of similarity amplification, not a claim that attention is unimportant for visual question answering.
The analysis next examines the singular values and left singular directions of the MLP down-projection \(W_{\mathrm{out}}\). Trained matrices exhibit stronger directional anisotropy than matched random matrices. Since both visual and text streams pass through this matrix, both can concentrate along the same output axes. Empirically, the MLP output subspaces under Orig and Noise retain a high leading principal-angle cosine, and the visual-side principal-angle basis preferentially overlaps the down-projection's dominant output subspace. Random-initialization controls reduce similarity under Noise, indicating dependence on trained weight structure rather than interface shape alone.
Proposition 1 supplies a conditional geometric bound. If the leading directions of two outputs are close to the same \(U_1\), with respective angular bounds \(\theta_X\) and \(\theta_Y\), their leading output-subspace principal-angle cosine satisfies
The bound requires no content coupling between \(X\) and \(Y\), so a common direction can produce high subspace similarity. However, proximity to that common direction is an assumption, not a theorem that every shared MLP necessarily satisfies. If the angle sum exceeds \(\pi/2\), the right-hand side is nonpositive and no longer provides a useful guarantee of high similarity. The result explains why a high leading cosine can be weight-induced alignment; it does not establish that all alignment lacks content contributions.
3. Principal-Angle Gap: check whether the second direction follows the first
At each layer, PCA is applied separately to the visual-token and text-token matrices, retaining \(k=30\) directions by default to form orthonormal bases \(U_V,U_H\). The singular values of \(U_V^{\top}U_H\) yield the descending principal-angle cosines \(\sigma_1,\sigma_2,\ldots\). These measure overlap between subspaces along different directions, not elementwise cosine similarity between token pairs. The conventional leading cosine only asks whether one pair of directions is close, allowing one common weight-imposed direction to inflate the measure.
Under Noise, the first direction remains strong while secondary directions weaken. Orig preserves stronger secondary overlap. This motivates the principal-angle gap, or PA gap:
A large gap indicates that the first direction dominates, consistent with one-directional weight-induced collapse. A small gap means that at least the second direction approaches the first and, under this paper's controls, accompanies richer structure. The difference does not identify objects, relations, or answers, nor does it measure effective dimensionality across all directions. In particular, if both leading cosines are low, a small gap cannot unconditionally mean strong, useful alignment. Spectrum shape and task controls remain necessary for interpretation.
The full-layer mean in the definition must be distinguished from layer trimming in experiments. Ranking tables and the appendix correlation analysis generally use the inner 80% of layers, excluding approximately the first and last 10% to reduce input-geometry and output-head effects. Under graded noise, the expected relation between PA gap and accuracy is negative: restoring visual content raises accuracy while reducing the dominance of one direction.
4. Task-Evidence Calibration: propagation of structure is not proof of answer relevance
Irr makes this distinction testable. A model that fully ignores irrelevant images should perform similarly to Text. If natural visual structure is processed without being appropriately filtered, Irr can preserve multi-directional geometry yet reduce question-answering accuracy. The authors train independent 20-class logistic-regression probes on mean-pooled projector-token features using 5-fold stratified cross-validation. These probes recognize the attached irrelevant image's own category but barely recognize the image category required by the current question. This supports continued encoding of the irrelevant input by the projector, not understanding of the question.
Inside the language model, the authors construct a principal-angle basis at each layer from Orig and inject either in-band or out-of-band noise into visual-token residual states. In-band noise lies in the subspace spanned by that basis; out-of-band noise lies in its orthogonal complement. Both are normalized and scaled by the local token norm with the same relative magnitude \(\varepsilon\), while text tokens are not directly perturbed. Comparing accuracy drops then tests whether the shared subspace is more behaviorally sensitive at matched perturbation strength, rather than inferring content propagation from similarity alone.
At \(\varepsilon=1.0\), in-band perturbations cause larger accuracy drops in all 13 models. At \(\varepsilon=0.3\), the direction remains consistent, but every difference is below 1 percentage point. Such interventions support behavioral influence of the subspace, yet do not turn PA gap into evidence of content use on an individual question. Perturbing important directions can affect an answer without establishing that relevant facts in the original image were correctly selected, reasoned over, and used.
A Worked Example¶
For Qwen2.5-VL-7B in the appendix, Orig accuracy is 85.4% and the mean PA gap over the inner 80% of layers is 0.166. Replacing projector tokens with norm-matched Noise gives 42.0% accuracy and a gap of 0.294. The leading principal-angle cosine, however, changes only from 0.709 to 0.690, making the loss of visual content easy to underestimate if this measure is used alone.
Under Irr, the gap is 0.262, below Noise's 0.294, indicating less one-directional geometry. Accuracy nevertheless falls further to 36.1%, below the 42.0% obtained with both Noise and Text. A single model can thus exhibit more structure that is more harmful to the current question. These values come from Appendix Tables 8 and 18, not estimates from figures or cross-model medians.
Loss & Training¶
The paper introduces no MLLM training loss, performs no fine-tuning of evaluated models, and reports no accuracy improvement from optimizing PA gap. Main experiments use released weights; random initialization is only a mechanism control. The only additional training is for independent linear probes, using multiclass logistic regression with \(\ell_2\) regularization, \(C=1.0\), and balanced class weights.
Key Experimental Results¶
Main Results¶
Evaluation covers five familiesโLLaVA-OV, LLaVA-OV-1.5, Qwen2-VL, Qwen2.5-VL, and InternVL3โwith 13 models spanning 0.5Bโ72B. Each condition uses the same 1,000 MMBench multiple-choice questions. This is not a new-model-versus-SOTA competition, but a behavioral comparison of the same models under controlled interventions.
The following table reproduces the informative contrasts from Table 2. Units are percentage points, and brackets give cross-model interquartile ranges. Each paired difference is computed within a model before taking the group median; subtracting medians of two accuracy distributions is not equivalent.
| Model group | NoiseโOrig | ShufโOrig | IrrโText | NoiseโText | IrrโNoise |
|---|---|---|---|---|---|
| All, 13 models | โ45.2 [โ46.6, โ43.2] | โ0.8 [โ1.9, โ0.1] | โ5.0 [โ5.9, โ2.9] | โ2.1 [โ3.7, +0.0] | โ2.5 [โ4.6, โ1.0] |
| <3B, 3 models | โ42.4 [โ42.8, โ40.1] | โ0.4 [โ0.5, +0.0] | โ2.9 [โ3.1, โ1.9] | โ5.3 [โ6.4, โ4.0] | +1.9 [+1.8, +3.3] |
| โฅ3B, 10 models | โ46.5 [โ47.5, โ44.1] | โ1.4 [โ2.5, โ0.3] | โ5.4 [โ5.9, โ4.7] | โ0.4 [โ3.0, +0.9] | โ3.2 [โ5.6, โ2.4] |
The graded-degradation correlation summary below corresponds to Table 1 and preserves the reported values. Appendix Table 12 specifies varying \(\alpha\) and computing Pearson \(r\) separately at each layer, averaging \(\lvert r\rvert\) over the inner 80% of layers, and then summarizing across models. This should not be described as computing one correlation directly from layer-averaged measures, nor as one pooled correlation across all 13 models.
| Metric | Mean \(\lvert r\rvert\) | Median \(\lvert r\rvert\) | Minimum \(\lvert r\rvert\) | Models with \(\lvert r\rvert>0.80\) | Models with expected sign |
|---|---|---|---|---|---|
| Leading PA cosine \(\sigma_1\) | 0.730 | 0.765 | 0.441 | 5/13 | 4/13 |
| PR | 0.775 | 0.767 | 0.527 | 6/13 | 9/13 |
| Spectral entropy | 0.825 | 0.847 | 0.622 | 9/13 | 12/13 |
| CKA | 0.752 | 0.765 | 0.443 | 6/13 | 4/13 |
| SVCCA | 0.736 | 0.788 | 0.508 | 6/13 | 3/13 |
| MIR | 0.760 | 0.777 | 0.569 | 4/13 | 11/13 |
| PA gap | 0.894 | 0.917 | 0.655 | 12/13 | 13/13 |
PR is the squared sum of principal-angle cosines divided by their sum of squares. Spectral entropy normalizes the cosines into weights summing to 1 and computes Shannon entropy. Both measure dispersion across the retained spectrum and serve as spectral baselines for PA gap. Strong absolute correlation still requires checking its sign, because the same measure can vary in opposite directions across models.
Ablation Study¶
The following table corresponds to Table 3 and reports cross-model medians of layer-mean scores. Arrows indicate the authors' preferred direction. An ordering requires a difference of at least 5% of the corresponding per-model metric range. Full ordering means Orig, Irr, Noise in the metric-preferred order, not their accuracy order.
| Metric | Orig | Irr | Noise | Orig ahead of Noise | Orig ahead of Irr | Full ordering |
|---|---|---|---|---|---|---|
| \(\sigma_1\uparrow\) | 0.705 | 0.664 | 0.830 | 3/13 | 11/13 | 2/13 |
| CKAโ | 0.100 | 0.098 | 0.202 | 4/13 | 4/13 | 4/13 |
| SVCCAโ | 0.496 | 0.474 | 0.579 | 3/13 | 13/13 | 2/13 |
| MIRโ | 9.29 | 9.50 | 9.10 | 5/13 | 5/13 | 5/13 |
| PA gapโ | 0.130 | 0.213 | 0.324 | 13/13 | 13/13 | 12/13 |
Mechanistic controls further support this interpretation. Cross-model median bypass effects are 0.036 for MLP and 0.010 for attention, with a reported ratio of 3.5. The displayed values are rounded and should not be used to recompute and replace that ratio. The median leading principal-angle cosine between Orig and Noise MLP-output subspaces is 0.996. Projection energy of the principal-angle basis into the down-projection's top 10 left singular directions is 2.770โ13.574 times the random baseline, with median 9.972.
At \(\varepsilon=1.0\), the in-band-minus-out-of-band accuracy-drop difference is 1.12โ5.48 percentage points, with median 2.38. Probes recognize the irrelevant image's own category at 74.4%โ78.5%, with median 76.4%, versus 5.5%โ7.2%, with median 6.4%, for the question-required category. The 20-class chance level is 5%. These results test sensitive directions inside the language model and image information retained by the projector, respectively; they are not the same accuracy measure.
Key Findings¶
- Reported cross-model median accuracies for Orig, Irr, and Noise are 85.4%, 36.1%, and 39.0%, respectively, while median PA gaps are 0.130, 0.213, and 0.324. Irrelevant images have more structured geometry than noise without being more useful.
- PA gap places Orig ahead of both Noise and Irr in all 13 models, but achieves the complete three-condition ordering in only 12/13. A universal strict ordering should not be claimed.
- Median IrrโNoise is +1.9 percentage points in the small-model group and โ3.2 in the larger-model group. Whether irrelevant images are more harmful than noise depends on scale rather than following a universal per-model rule.
- Median ShufโOrig is small, but some models show larger decreases. The authors did not store Shuf per-sample predictions, so that condition has no bootstrap interval or paired McNemar test; it cannot establish that shuffling has no effect.
Highlights & Insights¶
- Identify the diagnostic target: The distinction between representation structure and task-relevant information is the central contribution. PA gap helps avoid errors caused by a dominant first direction rather than replacing task evaluation.
- Counterfactuals strengthen interpretation: Norm-matched noise tests missing content, irrelevant images separate structure from relevance, and text-only baselines test the net benefit of visual input. This combination can audit other multimodal representation measures.
- Mechanisms and behavior constrain each other: Bypasses, weight spectra, subspace perturbations, and probes answer different questions. Combining them is more reliable than inferring image understanding directly from one high score.
Limitations & Future Work¶
- The scope is limited to projector-LLM architectures, a fixed MMBench multiple-choice subset, and the associated visual interventions. Open-ended generation, video, documents, and agentic tasks are not directly validated.
- PA gap uses only the first two principal-angle cosines and does not guarantee detection of every multi-dimensional structure or relevance to the answer. Future work could analyze more of the spectrum conditional on the task and add causal validation of specific visual facts.
- PCA-dimension sensitivity covers only two models, and random initialization only four. These auxiliary controls should not be presented as comprehensive robustness evidence across every family.
- Statistical descriptions are inconsistent: Table 1 wording suggests correlations involving layer averages, whereas Appendix Table 12 explicitly aggregates per-layer correlations. This note follows the latter procedural description while retaining the reported numbers.
- The down-projection-spectrum appendix simultaneously states that 12 models have available spectra and that all 13 values exceed the reference, while the main text says 13 models. Coverage is internally inconsistent. The main text gives an approximate leading singular-value ratio of 1.2โ1.6, whereas the appendix gives 1.166โ1.635; neither statement resolves the model-count conflict.
- The main text states that Irr is below Text in every model, but Appendix Tables 6 and 8 give Irr 38.9% and Text 38.7% for OV-1.5-8B, a difference of +0.2 percentage points. This note reports the overall negative paired contrast without repeating the universal claim.
- Noise endpoint accuracies differ across appendix experiments. For Qwen2.5-VL-7B, Table 8 gives 42.0%, Table 11 gives 41.2%, and Table 15 gives 42.1%. The paper does not fully explain these differences, and this note does not force distinct runs into one numerical record.
Related Work & Insights¶
- vs CKA, SVCCA, and classical principal-angle analysis: These tools compare representations; this paper audits content-level interpretations when weights are shared. The issue is not that the mathematical quantities are meaningless, but that geometric similarity alone lacks the evidence needed to establish visual content use.
- vs MIR and multimodal representation-alignment studies: Earlier work uses internal measures to study modality integration or shared task representations. This paper adds removal and mismatch controls, showing that high scores can mix in weight-induced components. Training diagnostics should place alignment curves alongside controlled behavioral contrasts.
- vs rogue dimensions and similarity-reliability studies: Unimodal analyses already show that a few dominant directions can obscure representation quality. This paper examines two modalities sharing a down-projection and supplements the first direction with information from the second principal angle.
- vs visual-dependence evaluation and causal tracing: Language-prior and irrelevant-image studies primarily test behavioral visual dependence, while causal tracing investigates how content affects outputs. This paper connects these questions to internal geometry without claiming that a geometric measure replaces causal content validation.
Rating¶
- Novelty: 4/5 โ Counterfactual interventions audit common alignment interpretations and connect a one-direction mechanism to a spectral diagnostic.
- Experimental Thoroughness: 4/5 โ The 13-model study and complementary controls are substantial, but task scope and auxiliary-control coverage remain limited.
- Writing Quality: 3/5 โ The structure and diagnostic boundary are clear, but coverage counts, aggregation descriptions, and some per-model claims conflict.
- Value: 5/5 โ A reusable hierarchy of evidence for multimodal interpretability prevents internal similarity from being mistaken for proof of visual understanding.