Skip to content

Preventing Expert Collapse in MoE-dVLMs via Modality-Wise Norm Alignment

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/Weixin-AI/TowerAlign
Area: Multimodal VLM
Keywords: Discrete Diffusion Vision-Language Models, Mixture-of-Experts, Expert Collapse, Modality-Wise Norm Alignment, Structure-Preserving Rescaling

TL;DR

Addressing catastrophic expert selection collapse on visual tokens in sparse mixture-of-experts discrete diffusion vision-language models, this paper uncovers an upstream geometric mismatch where a 167x norm disparity causes Pre-Norm attention drowning and representational inertia, and proposes TowerAlign to globally rescale visual token norms while preserving internal ViT sink topologies with negligible overhead.

Background & Motivation

Discrete masked diffusion language models (dLLMs) have emerged as a formidable alternative to autoregressive generation, offering bidirectional context reasoning and flexible non-monotonic decoding. Naturally, extending discrete diffusion backbones to multimodal architectures (dVLMs) by incorporating sparse Mixture-of-Experts (MoE) provides an appealing pathway toward scaling parameter capacity while keeping per-token inference activation costs low. However, when transitioning from dense diffusion models to MoE variants (MoE-dVLMs), the established two-stage training paradigm fails catastrophically, suffering a steep performance collapse on benchmarks such as ChartQA and DocVQA.

The direct manifestation of this breakdown is severe expert selection collapse on visual tokens: although the router's softmax entropy remains comparable to or even higher than text tokens across layers, greedy top-k routing assigns the vast majority of visual tokens to only a small subset of experts. Conventional routing-side interventionsโ€”such as scaling auxiliary load-balancing losses up to 10x or injecting strong Gumbel-Softmax perturbationsโ€”completely fail to resolve the collapse and degrade downstream task optimization even further. Singular value decomposition of the hidden states reveals the genuine underlying culprit: in a 2048-dimensional embedding space, the effective rank of visual tokens drops to an astonishingly low 2.1, in stark contrast to 1673.3 for textual tokens. Visual representations become confined to a degenerate, low-dimensional cone where routers cannot geometrically differentiate distinct visual tokens.

This severe representational homogenization stems from a massive L2 norm disparity across the modality boundary. Vision encoders output visual tokens with an average L2 norm of 60.47, whereas text embeddings have a mean norm of only 0.36โ€”a staggering 167-fold gap. Inside the Pre-Norm Transformer architecture, where sublayer outputs are normalized and calibrated to small text-scale updates while residual paths bypass normalization, injecting a text-scale update into a massive visual residual causes severe attention drowning and representational inertia. Visual tokens traverse deep layers along nearly straight trajectories with an effective angular velocity of only about 1 degree. The core idea of this paper is to bypass futile downstream router regularizations and directly introduce TowerAlign immediately after the visual projectorโ€”a structure-preserving global affine compression that rescales visual token norms to match text statistics, eliminating attention drowning and restoring router distinguishability.

Method

Overall Architecture

To eliminate geometric degradation in Pre-Norm residual dynamics, TowerAlign incorporates a statistical norm calibration pipeline immediately at the visual projector output. The multimodal inputsโ€”comprising raw images and interleaved text prompts with masked response sequencesโ€”are processed by a deep ViT encoder and projected via a two-layer MLP. Before visual tokens enter the Pre-Norm language backbone, TowerAlign evaluates the global average norm across the batch and applies a single structure-preserving scaling factor. The rescaled visual tokens and textual embeddings are then fed into the Pre-Norm Transformer layers, ensuring balanced residual updates in both self-attention and MoE feed-forward blocks under the discrete masked diffusion training objective.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image & Text Prompt"] --> B["Visual Feature Extraction & Projection<br/>SigLIP2 Encoder + 2-layer MLP"]
    B --> C["Global Affine Compression<br/>Scalar factor structure-preserving rescaling"]
    C --> D["Bidirectional Mask Concatenation<br/>Joint visual & text corrupted sequence"]
    D --> E["Pre-Norm Language Tower Balanced Updates<br/>Eliminate attention drowning & inertia"]
    E --> F["Top-k Differentiated Expert Routing<br/>Activate top-8 of 64 experts & denoise"]

Key Designs

1. Global Affine Compression: Structure-Preserving Radial Rescaling

Directly enforcing per-token normalization, such as token-wise RMSNorm, suppresses absolute norm discrepancies but completely destroys critical relative semantic hierarchies within visual representations. Deep vision transformers inherently produce a small set of ViT sink tokens (typically 3 to 5 tokens per image) characterized by exceptionally high L2 norms (frequently exceeding 100) that serve as global information aggregation hubs, whereas non-sink tokens average around 60 and capture fine-grained local details. Individually dividing each token by its own norm erases the magnitude cues distinguishing global context from local patches. TowerAlign resolves this by calculating a single global scalar across the entire batch and sequence dimensions:

\[\tilde{\mathbf{h}}_{v,b,i}^{(0)} = \lambda \cdot \mathbf{h}_{v,b,i}^{(0)}, \quad \lambda = \frac{\mu_{\text{target}}}{\frac{1}{BL_v}\sum_{b=1}^{B}\sum_{i=1}^{L_v}\|\mathbf{h}_{v,b,i}^{(0)}\|_2}\]

where \(\mu_{\text{target}} \approx 0.36\) aligns with the textual embedding norm. Because every token is scaled by the identical scalar \(\lambda\), the relative norm ratio \(\|\tilde{\mathbf{h}}_{v,i}^{(0)}\|_2 / \|\tilde{\mathbf{h}}_{v,j}^{(0)}\|_2 = \|\mathbf{h}_{v,i}^{(0)}\|_2 / \|\mathbf{h}_{v,j}^{(0)}\|_2\) remains strictly invariant. ViT sink tokens retain their relative salience in dot-product routing calculations, allowing the router to naturally assign global contextual reasoning and local feature processing to distinct specialized experts.

2. Mitigating Pre-Norm Attention Drowning and Representational Inertia

The core of this design is restoring the effective angular velocity of self-attention and FFN updates within Pre-Norm Transformer layers. In a Pre-Norm block \(\mathbf{u}^{(l)} = \mathbf{h}^{(l)} + \text{SelfAttn}(\text{RMSNorm}(\mathbf{h}^{(l)}))\), the normalized sublayer output yields a bounded, modality-agnostic update scale \(C^{(l)}\). When unscaled visual residual states carry norms \(\alpha_v \approx 60.47\), the angular change relative to the residual vector is suppressed:

\[\tan\theta_{\text{eff}, v}^{(l)} = \frac{C^{(l)}\sin(\theta_{v}^{(l)})}{\|\mathbf{h}_{v}^{(l)}\|_2 + C^{(l)}\cos(\theta_{v}^{(l)})} \approx \frac{C^{(l)}}{\alpha_v} \approx 1^\circ\]

This minuscule directional update traps visual representations in an inertial trajectory across dozens of layers, preventing angular dispersion and causing them to converge toward an isotropic dot product of 1.0 on a degenerate low-dimensional manifold. By compressing the visual residual baseline to \(\alpha_v \approx 0.36\), TowerAlign amplifies the relative update contribution by two orders of magnitude, allowing visual features to diversify across hidden dimensions and recovering an effective rank that enables clean top-k expert discrimination.

3. Joint Multi-Turn Bidirectional Masked Training Paradigm

During visual instruction tuning for discrete diffusion models, practitioners must choose between treating multi-turn dialogues as independent single-turn samples or packing complete conversations into a single sequence. Because discrete diffusion employs bidirectional attention rather than causal masking, joint multi-turn training theoretically allows early turns to perceive future conversational context during masked denoising. However, empirical evaluation confirms that joint multi-turn training substantially outperforms per-turn fine-tuning. Denoising under joint multi-turn sequences forces the model to capture dialogue-level consistency under a shared image representation while amortizing repetitive vision tower forward passes, boosting both token efficiency and long-context multimodal reasoning.

Loss & Training

The framework follows the standard LLaVA two-stage training scheme: 1. Stage 1 (Feature Alignment Pretraining): Only the two-layer MLP projector and TowerAlign scaling factor are trained for 1 epoch on the LCS-558K dataset with a learning rate of \(1 \times 10^{-3}\), cosine decay, and a 0.03 warmup ratio. 2. Stage 2 (Visual Instruction Tuning): Full end-to-end instruction tuning on the 780K single-image multi-turn conversations of LLaVA-NeXT, optimizing the vision encoder, projector, and language tower. The learning rate is set to \(1 \times 10^{-5}\) for the language model and \(2 \times 10^{-6}\) for the vision encoder. The model is optimized using the discrete masked diffusion negative log-likelihood objective:

\[\mathcal{L}_{\text{SFT}}(\theta) = -\mathbb{E}_{\tau \sim \mathcal{U}[0, 1], (\mathbf{x}_v, \mathbf{p}, \mathbf{y})}\left[\sum_{i=1}^{L_y} \log p_\theta\left(y^i \mid \mathbf{x}_v, \mathbf{p}, \mathbf{y}_\tau\right)\right]\]

where \(\mathbf{y}_\tau\) represents the masked corrupted sequence sampled with noise level \(\tau\). No aggressive auxiliary router loss is required to achieve well-balanced expert utilization.

Key Experimental Results

Main Results

Under matched data budgets, TowerAlign (LLaDA-MoE-7B-A1B, activating only 1B parameters per token) is evaluated against dense diffusion baselines and representative autoregressive models.

Benchmark Metric TowerAlign (7B-A1B) LLaDA-VLM (8B Dense) LLaVA-1.6 (7B AR) Analysis
ChartQA Accuracy 59.0 57.4 54.8 +1.6 over dense diffusion; substantially outperforms autoregressive baseline
DocVQA (val) ANLS 67.7 61.2 74.4 +6.5 over dense diffusion, markedly improving fine-grained text understanding
InfoVQA (val) ANLS 25.8 25.5 37.1 Competitive with dense diffusion baseline (+0.3)
RealWorldQA Accuracy 58.9 57.4 - Real-world perception exceeds dense diffusion by +1.5
MMStar Accuracy 46.7 46.0 - Comprehensive multimodal performance gains +0.7
AI2D Accuracy 70.5 70.2 66.6 Scientific diagram reasoning reaches 70.5 (+3.9 over LLaVA-1.6)
MMBench (EN) Score 73.3 71.9 54.6 Outperforms autoregressive 7B by +18.7 and dense diffusion by +1.4
MMMU (val) Accuracy 40.3 40.1 35.1 Multi-discipline reasoning surpasses both dense 8B and AR 7B
MMMU-Pro (standard) Accuracy 25.2 25.4 - Maintains parity with dense diffusion (-0.2) on complex tasks
MMMU-Pro (vision) Accuracy 15.1 15.7 - Slight variance on pure vision reasoning (-0.6)

Ablation Study

The ablation suite isolates the impact of norm alignment operators, router-side regularizers, and instruction-tuning sequence packaging.

Configuration ChartQA DocVQA RealWorldQA AI2D MMBench MMMU Note
TowerAlign (Full) 59.0 67.7 58.9 70.5 73.3 40.3 Global affine compression preserving relative token topology
w/o TowerAlign 4.7 13.8 30.6 58.7 61.4 36.7 Omitting alignment triggers severe collapse and performance drop
TokRMS 19.4 26.7 36.7 48.2 58.0 31.2 Token-wise RMSNorm erases global ViT sink token signals
Strong LB Loss (10x) 9.1 9.3 28.9 34.9 43.9 31.5 Router constraints fail to fix geometry and damage task loss
Gumbel Noise (ฯ„=2.0) 8.1 6.8 16.7 28.9 48.2 25.4 Stochastic perturbation disrupts expert specialization
Per-turn SFT 43.8 55.3 55.6 67.7 72.1 37.3 Lacks dialogue-level consistency and incurs redundant encoding overhead

Key Findings

  • Removing TowerAlign causes catastrophic collapse on document and chart reading tasks (ChartQA plunges from 59.0 to 4.7, DocVQA from 67.7 to 13.8), confirming that upstream norm alignment is an essential precondition for training MoE-dVLMs.
  • Per-token RMSNorm (TokRMS) degrades DocVQA and ChartQA by 41.0 and 39.6 points compared to TowerAlign, empirically validating the necessity of preserving ViT sink token relative magnitudes as global contextual beacons.
  • Conventional routing balance mechanisms (10x auxiliary load-balancing loss, Gumbel noise) completely fail to escape geometric collapse and actively harm model performance across all benchmarks.

Highlights & Insights

  • Incisive geometric diagnosis: Shifts the MoE collapse paradigm from router-level loss optimization to upstream modality-wise geometric discrepancies, identifying attention drowning and low-rank degeneracy as the root cause.
  • Elegant, zero-parameter intervention: Global affine compression introduces zero trainable parameters, restoring cross-modal dynamics while perfectly preserving internal visual semantic saliency.
  • Counter-intuitive empirical finding: Demonstrates that when representations suffer from low-rank geometric collapse, adding router regularization actually accelerates the degradation of primary task learning.

Limitations & Future Work

  • The target norm scale \(\mu_{\text{target}}\) remains fixed to the empirical text norm mean (0.36), without exploring dynamic or adaptive adjustments tailored to different language backbones.
  • Evaluations are focused on 7B-A1B models under single-image multi-turn scenarios; scaling behavior across higher resolutions, multi-image sequences, and video frames remains to be verified.
  • Sampling latency in iterative diffusion decoding over long chain-of-thought responses poses practical inference speed trade-offs compared to single-pass autoregressive systems.
  • vs LLaDA-V / LaViDa: While prior diffusion VLMs rely entirely on dense backbones, this work pioneers the integration of sparse MoE with discrete diffusion models, delineating unique failure modes and remedies.
  • vs Conventional MoE Routing Regularization (Switch Transformer / DeepSeekMoE): Traditional methods assume isotropic representation distributions and optimize router logits; this paper shows that Pre-Norm representational inertia renders router-side losses completely ineffective.
  • vs Dense Modality Gap Studies: Prior studies noted norm gaps in dense VLMs, but this work demonstrates how discrete top-k routing acts as a non-linear amplifier that turns mild representation flattening into catastrophic expert collapse.

Rating

  • Novelty: โญโญโญโญโญ Formulates and mathematically demonstrates the upstream geometric mechanism behind expert collapse in MoE-dVLMs.
  • Experimental Thoroughness: โญโญโญโญโญ Rigorously validated across 10 benchmarks, with deep SVD rank analyses, routing entropy evaluations, and ablations.
  • Writing Quality: โญโญโญโญโญ Logically seamless progression from empirical failure to mathematical proof and lightweight engineering resolution.
  • Value: โญโญโญโญโญ Provides foundational theoretical insight and an essential architectural guideline for scaling sparse diffusion vision-language models.