Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE¶
Conference: NeurIPS 2026 Spotlight
arXiv: 2609.38140
Paper: Project page
Area: Video Generation
Keywords: video diffusion models, mixture-of-experts, semantic routing, prototype guidance, sparse scaling
TL;DR¶
SplitMoE divides the sparse experts of a video diffusion Transformer into semantic and generic groups, guides semantic routing with clean VAE features, and improves video quality over same-source baselines with 27B total and 14B activated parameters, rather than forcing every expert to process an equal number of tokens.
Background & Motivation¶
Video generation must preserve object shapes, motion across frames, and local textures simultaneously, while simply enlarging a DiT feed-forward network increases the cost of every denoising step. Mixture-of-Experts (MoE) accesses larger total capacity through a small number of activated experts, but conventional implementations inherit token-wise routing from language models and use load balancing to give experts approximately equal token counts. Video tokens are distributed differently: backgrounds such as sky and walls can occupy large regions, while neighboring human or object patches are strongly correlated. Uniform expert utilization does not imply distinct visual responsibilities.
The paper's โuniformity trapโ concerns the conflict between this statistical objective and video structure. A single object may be split among experts to satisfy capacity or balancing constraints, and similar regions may switch experts between adjacent frames, producing spatial striping, semantic fragmentation, and temporal jitter. The authors use segmentation masks to analyze within-region routing consistency and between-class expert distinctiveness, together with routing visualizations. These analyses support a structural mismatch in standard routing; they do not prove that every video MoE should abandon load control.
Removing all balancing is not necessarily stable either, because redundant backgrounds can still attract most of the capacity. The paper instead asks which experts should acquire nonuniform specialization following natural visual distributions, and which should distribute generic reconstruction work. Core Idea: guide semantic experts with VAE prototypes to preserve regional coherence, while generic experts handle visual residuals under load control, avoiding the simultaneous demands of semantic clustering and uniform allocation within one homogeneous pool.
Method¶
Overall Architecture¶
The model receives text conditioning, the current noisy video latent, and a diffusion timestep, and produces predictions used for flow-matching denoising. SplitMoE upcycles the low-noise checkpoint of Wan 2.2, replacing only the FFNs at even-numbered layers between layers 15โ35 while leaving other components unchanged. Each sparse layer retains one continuously active shared expert and introduces 100 routed experts.
At inference time, each DiT token selects semantic and generic experts separately, and both groups' weighted outputs contribute to generation. Training adds a teacher path: clean-video VAE features and learnable prototypes produce soft assignments that supervise semantic routing from noisy DiT tokens. The prototypes are optimized through attraction to current and historical VAE features and repulsion between prototypes. Clean VAE features are not ground-truth video inputs additionally required at inference time.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Text, timestep,<br/>noisy video tokens"] --> B["Split-role expert partitioning"]
B --> C["Prototype-guided semantic routing"]
T["Clean VAE features<br/>training teacher only"] -.->|Soft-target supervision| C
T -.-> D["Prototype attraction and repulsion"]
D -.->|Update teacher prototypes| C
C --> E["Group-aware load control"]
E --> F["Aggregate expert outputs<br/>flow prediction and denoising"]
Solid edges indicate the generation path; dashed edges indicate training-only supervision or parameter updates. Prototype attraction and repulsion are not an additional network stage traversed at every inference step. The shared expert is also outside the Top-8 routed-expert budget below.
Key Designs¶
1. Split-role expert partitioning: fix the activation budget without forcing different visual information into one pool
The 100 routed experts comprise 20 semantic experts and 80 generic experts, with each token activating 2 semantic and 6 generic experts. The semantic group targets objects, layouts, and other high-level structure; the generic group retains capacity for local appearance and visual residuals. These are functional roles encouraged by routing supervision, not a strict low-/high-frequency decomposition of the video. Nor are foreground tokens sent only to the semantic group and background tokens only to the generic group: every token accesses both groups.
The router first applies a linear projection and GELU to each token, then computes logits through cosine similarity between normalized features and expert vectors, followed by independent sigmoid affinities. Independent sigmoids avoid a global softmax that makes the two groups compete for probability mass, allowing one token to match both a structural expert and a texture expert. Top-K selection is performed within each group, and output weights are normalized separately using the original, unbiased scores before the group contributions are summed. This is different from one normalization over all 8 selected experts.
The fixed \(K_s=2\), \(K_g=6\), and \(K=K_s+K_g=8\) constrain the number of activated experts. Expert widths are also adjusted to match the dense baseline's activated intermediate dimension, so each forward pass activates 14B out of 27B total parameters. The method expands accessible capacity rather than claiming the per-token computation of a 27B dense model; the shared expert continues to preserve general pretrained capabilities.
2. Prototype-guided semantic routing: turn visual similarity in reconstruction space into supervision for noisy tokens
Partitioning experts alone does not ensure that โsemantic expertsโ acquire semantic responsibilities. The authors associate one learnable prototype with each of the 20 semantic experts, placing prototypes in the clean VAE feature space. During training, each token's corresponding clean VAE feature is compared with all prototypes, and temperature-scaled cosine similarities yield a teacher soft assignment. The DiT router predicts a semantic distribution from the current noisy token. A KL-divergence objective aligns these distributions:
Here \(N\) is the token count, \(\mathbf{q}_i\) is the teacher assignment induced by VAE features and prototypes, and \(\mathbf{r}_i\) is the router's semantic prediction. The teacher target is stop-gradient, preventing the router from moving teacher prototypes merely to reduce KL loss; prototypes are optimized with their own objective. Semantic coherence is thus grounded in relatively stable clean visual features rather than clustering continually changing noisy DiT features alone.
Because the VAE already belongs to latent-space video generation, this teacher does not require an additional external representation model. However, VAE features are reconstructive representations, and prototypes are not manually defined semantic categories. A prototype may cover related appearance or structural patterns, so each semantic expert should not be interpreted as a fixed, nameable object class. At inference time, the router operates directly on DiT tokens without querying a clean target video.
3. Prototype attraction and repulsion: prevent background collapse without pushing diverse prototypes away from data
Attraction alone can draw many prototypes toward repetitive backgrounds, creating redundant centers. Repulsion alone can move prototypes to locations unsupported by actual tokens, leaving experts unused. The pull mechanism uses both the current mini-batch and recent historical VAE features stored in a circular token bank. Current features encourage coverage of the current samples, while historical features keep each prototype near visual patterns that have actually occurred, reducing the impact of rare patterns disappearing from a small batch.
The push mechanism penalizes prototype pairs whose cosine similarity exceeds the margin \(\delta\), preventing overlapping centers. Together they form \(\mathcal{L}_{\rm proto}\), which seeks data-supported, nonredundant centers on the visual feature manifold rather than equal token counts. The original prototype objective includes smooth coverage of current features, a nearest-similar-feature term over the historical bank, and pairwise margin-based repulsion. This note retains the mechanism explanation rather than replacing the complex objective with an unverified simplified formula.
4. Group-aware load control: constrain utilization in the generic group without disrupting semantic coherence
Semantic experts receive no explicit load balancing, allowing common and rare visual patterns to produce different utilization rates; prototype supervision and attractionโrepulsion mitigate collapse. Generic experts use loss-free load control inspired by DeepSeek-V3, with non-gradient biases affecting discrete expert selection. Biases change which experts are selected but do not enter the final mixture weights. Selected experts remain weighted by their original affinity scores, avoiding the interpretation of load corrections as genuine tokenโexpert compatibility.
The design does not โeliminate balancingโ; it restricts balancing to the group providing generic reconstruction capacity. Fixed 2+6 activation slots also mean that coarse-to-fine generation does not dynamically transfer more Top-K slots to semantic experts. The paper separately observes timestep-dependent semantic routing weights, but this is a routing diagnostic rather than an explicit timestep-gating rule. The within-group normalization equations alone cannot be interpreted as the ratio of the two branches' total output energy.
A Worked Example¶
Consider a generated clip in which a horse transitions from standing to running. For neighboring tokens on the same object, standard balanced routing may choose different experts to distribute load and change those choices across frames. During SplitMoE training, the teacher uses clean VAE features at these positions to provide similar prototype soft targets. The router learns similar semantic assignments from noisy states, while the generic group handles fur, backgrounds, and local appearance differences.
Each token still selects 2 of the 20 semantic candidates and 6 of the 80 generic candidates, alongside the continuously active shared expert. At test time, no real horse video is supplied as a teacher: the model predicts routing from its current latent and iteratively denoises. This example illustrates cooperation between modules; it neither requires every horse token to use one expert nor adds segmentation-mask inputs to the model.
Loss & Training¶
The full training objective is:
\(\mathcal{L}_{\rm flow}\) is the flow-matching loss, while \(\lambda_{\rm align}\) and \(\lambda_{\rm proto}\) control the two auxiliary training objectives. Generic-group load control adds no auxiliary optimization loss and updates only non-gradient selection biases. Routed experts are initialized from the dense FFN to retain pretrained generative priors. Same-source models all train for 80k steps with the same low-noise initialization, training data, and inference settings.
The cached text does not provide a reproducible training-data scale, numerical values for the two loss weights, or all optimization hyperparameters, and it contains no separate appendix section. These settings therefore cannot be supplied from the available material. Validation-loss curves support faster convergence in training steps: the authors report comparable loss at roughly 70% of the steps, which should not be restated as a 30% reduction in total wall-clock training time.
Key Experimental Results¶
Main Results¶
VBench-2.0 evaluates Creativity, Commonsense, Controllability, Human Fidelity, and Physics. T2V-CompBench evaluates Consistent Attribute, Interaction, and Numeracy. The table preserves the original percentages, with higher being better for every metric. Videos contain 81 frames, but independent models use their own default inference settings, and some baseline scores come from prior papers or the benchmark.
| Model | Creativity | Commonsense | Controllability | Human Fidelity | Physics | Consistent Attribute | Interaction | Numeracy |
|---|---|---|---|---|---|---|---|---|
| Official Wan2.2 | 58.35% | 62.45% | 48.73% | 75.33% | 69.07% | 83.19% | 73.29% | 16.96% |
| LongCat-Video | 54.73% | 70.94% | 44.79% | 80.20% | 59.92% | 73.38% | 61.76% | 47.87% |
| LTX-2 | 50.39% | 67.09% | 38.09% | 73.69% | 76.71% | 69.93% | 47.98% | 25.83% |
| Full SplitMoE | 58.46% | 64.89% | 45.72% | 84.47% | 69.41% | 84.63% | 69.98% | 44.01% |
Relative to official Wan2.2, Human Fidelity improves by 9.14 percentage points, but Controllability is lower by 3.01 points and Interaction by 3.31 points. SplitMoE also trails LongCat-Video in Commonsense and Numeracy and LTX-2 in Physics. These results do not establish state of the art on every dimension.
Official Wan2.2 uses a high-/low-noise dual-model ensemble, whereas the paper's same-source dense fine-tuning and MoE variants start from the low-noise checkpoint. The table describes system capabilities, but differences from the official model cannot be attributed entirely to SplitMoE routing. The same-source ablations below provide a more controlled comparison of architectural contributions.
Ablation Study¶
The table selects four dimensions most relevant to structure, appearance, and composition from the same-source experiments. All models use the same 80k-step training setup, and MoE models activate 14B parameters.
| Config | Human Fidelity | Physics | Consistent Attribute | Numeracy |
|---|---|---|---|---|
| Dense Wan2.2-FT | 73.08% | 54.66% | 77.51% | 36.71% |
| Wan2.2-MoE | 73.49% | 63.32% | 78.82% | 38.55% |
| SplitMoE w/o PG | 74.04% | 55.89% | 78.46% | 38.57% |
| SplitMoE w/o Push | 73.60% | 56.44% | 77.29% | 37.15% |
| SplitMoE w/o Pull | 73.81% | 56.83% | 78.14% | 38.32% |
| Full SplitMoE | 84.47% | 69.41% | 84.63% | 44.01% |
The w/o PG variant retains the 20/80 partition and dual-track routing but removes the VAE teacher, prototypes, and both auxiliary objectives; it does not merely delete one KL term. Relative to standard MoE, the full model improves Human Fidelity by 10.98 percentage points and Physics by 6.09 points. Relative to partition-only w/o PG, Physics improves by 13.52 points, showing that naming expert roles is insufficient to create useful specialization.
Costs below come from the original Table 2 and are measured in seconds per iteration, not total generation time for a video:
| Config | Total Parameters | Activated Parameters | Training Seconds/Iteration | Inference Seconds/Iteration |
|---|---|---|---|---|
| Full SplitMoE | 27B | 14B | 6.84 | 6.08 |
| Standard MoE | 27B | 14B | 6.69 | 6.12 |
| Dense Model | 14B | 14B | 5.66 | 5.76 |
SplitMoE training iterations are approximately 20.85% slower than the dense model and 2.24% slower than standard MoE; inference iterations are approximately 5.56% slower than the dense model. Matched activation budgets therefore do not imply matched latency, let alone faster sparse-model training iterations.
Key Findings¶
- The full model exceeds dense fine-tuning, standard MoE, and the listed ablations on all eight dimensions of the same-source Table 1. This conclusion should not be extended to all dimensions of independently developed models.
- Attraction and repulsion must work together. Removing push lets prototypes cluster in dominant background regions; removing pull can move them away from the actual feature manifold. Both variants score below standard MoE's 63.32% Physics result, so incomplete auxiliary priors can be harmful.
- Over a 50-step denoising trajectory, the authors observe semantic weights peaking around steps 10โ15 and then declining, supporting a coarse-to-fine change in responsibilities. The cached text does not fully define this diagnostic weight, so it is not interpreted as a change to the fixed 2+6 activation slots.
Highlights & Insights¶
- Making balancing role-dependent rather than universal is more consequential than merely enlarging the expert pool. The semantic group accommodates natural long tails while the generic group retains utilization control, avoiding the misconception that rejecting uniformity means rejecting load management.
- The teacher resides in the generation model's existing VAE space instead of requiring another large visual encoder. Stop-gradient targets and separate prototype optimization distinguish learning the router from constructing teacher centers, reducing instability from mutual chasing.
- Pull and push address opposite failure modes rather than duplicate regularization. Their ablations show that either direction alone does not yield reliable benefits, demonstrating the need for both data support and collapse prevention.
Limitations & Future Work¶
- The authors acknowledge dependence on frozen visual features and sparse-routing overhead. Visual similarity in VAE prototypes does not guarantee coverage of action causality or physical attributes, and reliability under domain shift remains untested.
- The setup fixes the 20/80 expert partition and 2+6 activation slots without a ratio sweep or comprehensive scaling trends across model sizes. A single 27B total-parameter configuration does not establish equally sustained gains for larger video models.
- Results primarily concern 81-frame text-to-video generation and benchmark evaluation, not long videos, interactive world models, or other generation tasks. Multiple-seed uncertainty and end-to-end distributed throughput evidence are also absent.
- Noise-stage-adaptive activation budgets are a possible research direction, but must jointly control latency, generic-group loads, and rare semantic-expert coverage. Weight curves alone do not establish that dynamic allocation would improve performance.
Related Work & Insights¶
- vs Wan2.2: The official model separates high- and low-noise stages at a coarse model level; this paper introduces token-wise semantic/generic roles inside a model initialized from the low-noise checkpoint. These splits operate at different levels and should not be conflated as one MoE strategy.
- vs ProMoE: Both use prototypes, but SplitMoE connects clean VAE representations with DiT routing predictions and partitions experts around video semantics and visual residuals, rather than only separating conditioning types or directly matching prototypes within DiT space.
- vs DeepSeek-V3: SplitMoE borrows loss-free load control but applies it only to generic experts. The transferable lesson is to distinguish system-capacity utilization from task-structural coherence instead of importing language-model routing constraints unchanged.
Rating¶
- Novelty: 4/5, role partitioning and VAE prototype supervision form a clear video-specific combination, although the individual components have precedents.
- Experimental Thoroughness: 4/5, same-source ablations, two benchmarks, and routing analyses are substantial, but ratio sweeps, uncertainty estimates, and broader tasks remain missing.
- Writing Quality: 4/5, failure modes map clearly to module designs, while timestep-weight diagnostics and some reproduction details need more precise specification.
- Value: 4/5, a concrete, testable routing design for sparse video scaling, whose practical benefits still require evaluation alongside memory, latency, and distributed costs.