SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization¶
Conference: NeurIPS 2026
arXiv: 2608.04084
Code: https://github.com/Beryex/SpecDrop
Area: Mixture of Experts (MoE)
Keywords: category-conditioned routing, modular specialization, fixed soft gating, shared expert, training-signal granularity
TL;DR¶
SpecDrop shapes modular specialization through a fixed category-to-branch assignment, nonzero cross-category soft weights, and fixed-denominator merging, outperforming label-free matched controls on vision tasks with trusted inference-time category labels, while label-aware dense logit masking is more accurate and routing gains on fuzzy text partitions are near zero.
Background & Motivation¶
Mixture-of-experts (MoE) models and multi-branch networks distribute capacity across modules without necessarily creating distinct expertise: several experts may remain generalists handling many kinds of input. Common interventions learn a router, add load-balancing or mutual-information losses, or impose random masks to encourage division of labor. However, if the training signal reaching a module already mixes several categories, a more elaborate gating algorithm may not turn that signal into clear expertise. Rather than attributing the problem simply to router instability, this paper asks whether routing-signal granularity matches the category structure the model is meant to learn.
Two meanings of alignment must be separated. Data-side partition alignment concerns whether a training unit carries one clean category label; model-side branchโcategory alignment concerns whether a category mainly depends on its assigned branch after training. An image can be assigned a superclass derived from its fine class, so the tag corresponds to the whole image. A document-source domain or instruction-task domain does not necessarily describe each token's content or required skill. Even when domain routing makes text branches specialize, aggregate language modeling may not improve. Four vision and language settings therefore separate the formation of structure from performance gains.
SpecDrop treats the category label as available external information rather than asking an additional network to infer routing. The vision margins consequently include the label's value: the test-time CIFAR superclass and BREEDS supercategory are derived from target fine labels and are not standard label-free classification results. Core Idea: use fixed, category-asymmetric soft gating to change the proportions of training gradients reaching each branch, then test when category granularity translates specialization into performance through matched-architecture controls, label-aware masking, and fuzzy-partition comparisons.
Method¶
Overall Architecture¶
The model receives a category tag as well as an image or text input. Shared components produce features, and parallel branches process the same input at each routed layer; Category Assignment & Soft Gating identifies a preferred branch, Fixed-Denominator Merge & Shared Expert combines all branch outputs, and Constant-Sum Routing Schedule gradually strengthens category differences during training. Each input's branch weights are reused across routed layers, with a shared classification head or language-model output producing the prediction.
CIFAR routes at the three layer groups of ResNet-110; ImageNet routes at all 12 ViT MLPs; language modeling routes at each layer's FFN; and SuperNI routes at the seven adapted linear projections in each Llama-3.2-1B block. The main experiments use Soft SpecDrop: both training and inference apply deterministic soft weights, every branch has nonzero weight, and computation is not reduced by executing only a few experts. This is not stochastic dropout.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
X["Input and shared features"] --> A["Category Assignment &<br/>Soft Gating"]
C["Category tag<br/>required in training and inference"] --> A
A --> B["Fixed-Denominator Merge &<br/>Shared Expert"]
B --> O["Prediction"]
B -.->|training schedule| D["Constant-Sum<br/>Routing Schedule"]
D -.->|update soft weights| A
O -.->|training only| L["Task loss and back-propagation"]
Y["Target label<br/>training supervision only"] -.-> L
The category tag is not merely a training-supervision input: it still enters gating along the solid path at inference. Dashed paths indicate training schedules and loss supervision, not an additional learning round during inference. Although vision category tags are coarsenings of target labels, deployment must still obtain them reliably from another source.
Key Designs¶
1. Category Assignment & Soft Gating: fix a preference without isolating other branches
Each category is assigned a preferred branch through a round-robin rule; when there are more categories than branches, one branch can receive several categories. The assignment is not updated through gradient descent and does not depend on input features. Thus, โparameter-freeโ describes only the routing rule, not a network with no trainable parameters. The four deployed settings use 20, 46, 7, and 20 branches, respectively, matching their category counts.
The preferred branch receives a larger weight, while other branches receive smaller but nonzero leakage weights. Uniform merging makes every branch process each category at equal strength, whereas hard category routing cuts off knowledge pathways from other categories. Soft SpecDrop lets a branch learn more from its assigned category while preserving cross-category gradient flow. An assignment matrix and two weights express this mechanism:
Here \(A_{ck}\) indicates whether the category is assigned to the branch, and \(p_{\mathrm{a}}\) and \(p_{\mathrm{i}}\) are weights directly applied by the Soft variant. They retain the stochastic variant's โactivation probabilityโ notation, but the main model samples no Bernoulli masks. A small weight does not imply that the branch is skipped.
Back-propagation explains why gating may induce specialization: each branch's base gradient is multiplied by its merge coefficient. However, the theorem's assigned-versus-unassigned gradient ratio of \(p_{\mathrm{a}}/p_{\mathrm{i}}\) additionally requires conditional mask-independence of the base gradient and equal expected base-gradient magnitudes across categories. Soft removes mask randomness without removing category differences in base gradients. The theorem therefore does not unconditionally guarantee exclusive expertise in the final branches.
2. Fixed-Denominator Merge & Shared Expert: keep the routed mixture at single-branch scale
If changing category preferences also changes merged-output magnitude, downstream components may interpret gating differences as global scale changes. SpecDrop divides the weighted sum by a fixed total weight. Because each category prefers exactly one branch, this total is category-independent. An optional shared expert is added with unit coefficient after normalization rather than following category-dependent weights.
Here \(h_k\) is a branch output and \(h_{\mathrm{SE}}\) is the always-on shared expert's output. The routed coefficients sum to one, preserving convex-combination scale, while the shared expert handles capabilities common to categories. The final CIFAR configuration uses no shared expert; ImageNet, SlimPajama, and SuperNI retain one. Adding a shared expert requires reallocating capacity, so its architectural benefit must not be counted entirely as a routing benefit.
For the appendix's Bernoulli variant, a fixed denominator also matches the expectation of a merge to deterministic test-time merging, provided branch outputs do not depend on the current mask. Because one mask is reused across multiple nonlinear layer groups, this result cannot be extended to exact equality between the entire stochastic network's training and inference outputs. The main Soft variant has no stochastic-denominator bias to remove; scale calibration and composition with the shared expert are the denominator's main roles there.
3. Constant-Sum Routing Schedule: strengthen category differences without changing the merge total
Training does not immediately impose a strong category preference on narrow branches. It starts with uniform merging and raises the preferred weight through a cosine schedule; other weights follow an inverse coupling that preserves the total. This does not progressively reduce the number of computed experts. It changes each category's contribution to branch training. Vision schedules operate per epoch, while language and LoRA schedules operate per step; the final schedule spans the whole training run.
For imbalanced category frequencies, the method additionally amplifies the preferred-minus-leakage gap for rare categories, then redistributes the two weights around a common average. The following expression combines the source's gap definition and its constant-total transformation:
Here \(\pi_c\) is category frequency, \(M\) is the category count, and \(\beta\) controls amplification. The weighted total remains \(S\); balanced data or \(\beta=0\) recovers the scalar configuration. This extension requires valid weight ranges and does not permit arbitrarily large exponents. The appendix gives validity bounds and describes runtime checks for clipping that would break the total. Its empirical benefit is weak, so it should not be presented as a proven key solution to category imbalance.
A Worked Example¶
For a CIFAR input from one superclass, the fixed assignment selects its preferred branch. The final configuration uses \(K=20\), \(p_{\mathrm{a}}=0.7\), and \(p_{\mathrm{i}}=0.3\), giving \(S=6.4\). The preferred branch's actual merge coefficient is approximately 0.109, and each other branch's is approximately 0.047. This is neither โa 70% chance of using one expertโ nor single-branch computation.
These coefficients are reused at all three ResNet layer groups. Each group executes all narrow branches, merges their outputs, and passes the result to the next group. CIFAR adds no shared expert. The final shared classifier still predicts 100 fine classes, not only the superclass's five classes; output restriction emerges through training rather than explicit inference-time masking of other classes.
During training, the same input's fine label supervises cross-entropy. Inference omits the loss but still requires the superclass. If another classifier must predict that tag first, its errors affect routing, and an additional approximately dense-scale forward pass is required.
Loss & Training¶
Training uses only the original task objective: classification cross-entropy, next-token cross-entropy, or supervised instruction tuning. There are no routing auxiliary losses or learned router parameters. CIFAR uses SGD with learning rate 0.1, batch 128, and 200 epochs; ImageNet uses AdamW with learning rate \(2.5\times10^{-4}\), batch 256, and 100 epochs under a short DeiT recipe without EMA or RepeatedAugmentation.
SlimPajama uses an approximately 30M Transformer with six layers and hidden width 384, trained for 10 epochs on 500M unique tokens with batches of 32 sequences of 512 tokens. SuperNI freezes the approximately 1.24B Llama-3.2-1B base and trains approximately 222M multi-branch LoRA parameters with effective batch 128 for 3 epochs. This is not a conventional LoRA regime with less than 1% trainable parameters.
The final preferred weights are 0.7, 0.6, 0.6, and 0.8, respectively; leakage weights are complementary in the selection experiments. Different preferred-weight candidates can have different fixed totals, whereas warmup within one configuration preserves its own total. ImageNet, SlimPajama, and SuperNI use imbalance exponents 1, 4, and 1 and shared-expert capacities relative to one branch of 2, 0.5, and 1, respectively; CIFAR uses no shared expert.
Key Experimental Results¶
Main Results¶
The following values come from main-text Tables 1โ4 and are three-seed means with standard deviations. Top-1 is in percent, lower PPL is better, and ROUGE-L is F1. PPL standard deviations use the population convention; other main-table standard deviations use the sample convention. No-Routing means uniform merging on the same branch architecture, and these vision main-table controls do not consume inference-time category labels.
| Setting / Method | Parameters | Implemented compute | Primary metric | Align (%) |
|---|---|---|---|---|
| CIFAR / dense ResNet-110 | 1.737M | 255.3M MACs/image | 74.48 ยฑ 0.13 | โ |
| CIFAR / No-Routing | 1.721M | 287.6M MACs/image | 63.08 ยฑ 0.04 | 3.3 ยฑ 2.9 |
| CIFAR / HardCategory | 1.721M | 287.6M MACs/image | 57.67 ยฑ 3.48 | 100.0 ยฑ 0.0 |
| CIFAR / Soft SpecDrop | 1.721M | 287.6M MACs/image | 79.23 ยฑ 0.17 | 68.3 ยฑ 5.8 |
| ImageNet / dense ViT-S/16 | 22.051M | 4.25G MACs/image | 76.38 ยฑ 0.15 | โ |
| ImageNet / tuned Soft MoE | 22.196M | 2.91G MACs/image | 76.69 ยฑ 0.70 | โ |
| ImageNet / compute-matched Soft MoE | 481.2M | 4.26G MACs/image | 66.72 ยฑ 0.63 | โ |
| ImageNet / No-Routing+SE | 22.263M | 4.25G MACs/image | 73.36 ยฑ 0.29 | 2.9 ยฑ 2.5 |
| ImageNet / Soft SpecDrop | 22.263M | 4.25G MACs/image | 79.89 ยฑ 0.18 | 100.0 ยฑ 0.0 |
| SlimPajama / dense Transformer | 30.143M | Within 15.93โ15.97G MACs/sequence | 44.80 ยฑ 0.05 | โ |
| SlimPajama / No-Routing+SE | 30.168M | Same range | 45.28 ยฑ 0.10 | 5.6 ยฑ 9.6 |
| SlimPajama / Soft SpecDrop | 30.168M | Same range | 45.38 ยฑ 0.02 | 94.4 ยฑ 9.6 |
| SuperNI / HydraLoRA | 226.56M | Adapter add-on over base: +18.1% | 0.5153 ยฑ 0.003 | 0.0 ยฑ 0.0 |
| SuperNI / No-Routing+SE | 221.92M | Adapter add-on over base: +18.0% | 0.5094 ยฑ 0.007 | 0.0 ยฑ 0.0 |
| SuperNI / Soft SpecDrop | 221.92M | Adapter add-on over base: +18.0% | 0.5106 ยฑ 0.003 | 22.2 ยฑ 15.4 |
Compute values are from Appendix F.3; the entire table is not strictly FLOPs-matched. CIFAR multi-branch models use 12.7% more MACs than dense. The two Soft MoE configurations match parameters or approximately match compute, not both simultaneously. ImageNet counts omit approximately 0.36G fused-attention products for each method, and SlimPajama counts omit approximately 0.60G queryโkey products. SuperNI parameter counts include only trainable adapters, and its compute column is the add-on relative to the frozen base, not full-model MACs.
Align is the percentage of categories for which the branch causing the largest performance loss when zero-ablated is the assigned branch. Ties go to the lowest branch index. LoRA uses signed performance degradation, so this is not equivalent to the heatmap's largest absolute change. Only six SlimPajama domains have validation coverage, and only 15 SuperNI clusters have test tasks. Even 100% Align does not imply high accuracy, as HardCategory demonstrates.
Ablation Study¶
The information-matched logit-masking control is essential to interpreting the vision results: give every model the same category label and retain only the logits of that category's fine classes. The following values are three-seed means from Appendix E.3 Table 9, all in Top-1 percent.
| Dataset / Method | Without logit masking | With logit masking | Change |
|---|---|---|---|
| CIFAR / dense ResNet-110 | 74.33 | 85.23 | +10.90 |
| CIFAR / No-Routing | 63.07 | 78.47 | +15.40 |
| CIFAR / Soft SpecDrop | 79.23 | 79.23 | +0.00 |
| ImageNet / dense ViT-S/16 | 76.37 | 83.65 | +7.28 |
| ImageNet / No-Routing+SE | 73.36 | 81.44 | +8.08 |
| ImageNet / Soft SpecDrop | 79.89 | 80.95 | +1.06 |
Unmasked entries differ from the main tables by up to 0.15. The authors explain that two missing seed-42 dense checkpoints were retrained from stored configurations, with version-related drift. Each table's own values should be preserved rather than replacing 74.33 with 74.48 to manufacture consistency.
Another CIFAR analysis, Appendix A.3 Table 7, compares masking and denominator choices. It is not a fully controlled single-variable ablation: Soft uses 0.7/0.3, Bernoulli uses 0.9/0.1, and random dropout uses uniform 0.5 with no warmup.
| Config | Denominator | Top-1 (%) | Experimental boundary |
|---|---|---|---|
| Category-conditioned Soft | Fixed total | 79.23 ยฑ 0.17 | Final deployed configuration |
| Category-conditioned Bernoulli | Fixed total | 64.55 | Changes both soft weights and sampling |
| Category-conditioned Bernoulli | Sampled active count | 67.54 | Stochastic-denominator bias remains |
| Category-free random dropout | Fixed total | 58.97 | Probabilities and warmup also differ |
| Category-free random dropout | Sampled active count | 59.53 | Cannot attribute solely to removing category assignment |
Key Findings¶
- The matched-architecture vision gains are +16.15 and +6.53 percentage points. ImageNet's shared expert first contributes +2.06. Architecture matching fixes capacity and implemented compute, but not vision-label information.
- After matching label information, dense+mask is more accurate on both vision datasets. CIFAR logit masking adds zero to SpecDrop, suggesting that category restriction is largely internalized. Its added value is trained-in structure supporting branch ablations, not the best classification deployment.
- SlimPajama reaches 94.4% Align while remaining 0.10 PPL above the matched control; SuperNI routing adds only 0.0012 ROUGE-L. Neither yields a clear gain under the authors' seed-noise interpretation. Structural alignment cannot be used as evidence of improved performance.
- A 125M text-model replication remains 0.17 PPL above the matched control. The one-epoch check gives 61.91 versus 61.93, with batch also changed from 32 to 64, so it is not an epoch-only ablation. Although 56.1% of validation text chunks span at least two BGE clusters, purity has near-zero correlation with per-chunk loss differences. This supports a distribution-level interpretation, not per-example causal proof.
Highlights & Insights¶
- Fixed gating serves as a mechanism probe: category-asymmetric training signals can produce division of labor without a learned router. It separates router quality from the meaningfulness of the categories supplied to routing.
- Nonzero leakage and shared experts both provide cross-category knowledge pathways. Hard CIFAR routing without a shared expert reaches only 57.67%, while a different configuration with a shared expert reaches 76.95%. Hard-routing failures therefore cannot all be attributed to one probability choice.
- Specialization, performance, and label-information value require separate reporting. The transferable experiment design combines an equal-weight, same-architecture control for mechanism gains, an information-matched control for label benefits, and branch ablations for dependency structure.
Limitations & Future Work¶
- Inference requires trusted category labels. On CIFAR, tags from coarsened dense predictions are 83.8% accurate and yield 69.33 Top-1; ImageNet's 88.5% tag accuracy yields 74.25. Neither exceeds its dense reference, and label prediction approximately doubles deployment compute.
- The fuzzy-partition interpretation is supported by four settings and embedding diagnostics, not established as a universal causal law. Vision DINOv2 silhouette is only 0.069, also below a substantial-cluster-structure threshold. Modality, model, label definition, and training protocol change together; controlled changes to label granularity remain necessary.
- Scale is limited, the main text experiment traverses 500M tokens ten times, LoRA adapters are approximately 18% of the base, and several baselines deviate from original configurations. SMoE-Dropout's expert-count growth schedule did not execute, and Mod-Squad is an FFN-only adaptation. Rankings cannot be generalized to all original-paper configurations.
- CIFAR selects the best checkpoint on the test split, while SuperNI selects epochs using ROUGE-L on 119 held-out tasks and generates only 10 sampled instances per task for that metric. Independent validation splits, larger generation samples, and explicit label-prediction cost accounting would improve external validity.
- A possible extension is to train structure under fixed labels and then hand over to a label-free learned router. Dense prediction requires testing whether region/token-level tags recover gains. These are proposed directions, not completed detection, segmentation, or large-scale MoE evaluations.
Related Work & Insights¶
- vs Soft MoE / Switch: These methods learn input-dependent expert assignments; SpecDrop fixes category preferences from the start. It removes routers and auxiliary losses but leaves category acquisition to deployment, and its main Soft variant still computes all branches.
- vs DEMix / Branch-Train: Both use domain tags to organize experts. SpecDrop's small leakage weights preserve cross-domain learning, unlike hard domain isolation. Performance gains still depend on how domain tags align with training-unit granularity.
- vs Stochastic Depth / dropout: These methods primarily regularize through random deactivation. SpecDrop's strongest results come from deterministic category-conditioned soft weights. โDropโ in the name does not make it a better stochastic regularizer.
- vs StableMoE / auxiliary-loss-free balancing: StableMoE learns before freezing, and auxiliary-loss-free balancing adjusts loads through online biases. This paper fixes assignments at step zero and statically adjusts weights using category frequencies. Combining category priors with learned routing is a research direction, not an implemented result here.
Rating¶
- Novelty: 3/5 โ Fixed category routing is not entirely new, but separating granularity, structure, and performance has clear research value.
- Experimental Thoroughness: 4/5 โ Four settings, information-matched controls, and scaling checks provide broad evidence; privileged labels, baseline adaptations, and checkpoint selection limit generalization.
- Writing Quality: 4/5 โ v2 clearly qualifies Soft gating, label benefits, and theoretical assumptions; checkpoint-related numerical differences between main and control tables are disclosed.
- Value: 4/5 โ More useful as a diagnostic and category-signal reference for modular specialization than as a label-free compression or sparse-compute deployment method.