Beyond Missing Rates: Rethinking Incomplete Multi-View Clustering with Protocol Divergence¶
Conference: NeurIPS 2026 (Poster)
arXiv: 2606.04857
Code: https://github.com/dk23lhl/CRAFT
Area: Self-Supervised Learning
Keywords: incomplete multi-view clustering, protocol divergence, observation support, masked attention, train-once deployment
TL;DR¶
The paper exposes evaluation-protocol divergence in incomplete multi-view clustering through effective missing rate and complete-sample proportion, and introduces CRAFT, which fuses only each sample's observed views: it wins 12 of 13 information-matched incomplete-training conditions, while a separate evaluation tests cross-protocol deployment of checkpoints trained with all views available.
Background & Motivation¶
Incomplete Multi-View Clustering (IMVC) seeks meaningful clusters of unlabeled samples when some sensors, modalities, or feature sources are unavailable. Methods such as COMPLETER, DCP, and DCG learn cross-view prediction or reconstruction from views co-observed for the same sample, but a “missing rate of 0.5” does not specify how much supervision remains. If half the samples each lose one view, only approximately 8.3% of entries are missing in six-view data; deleting half of all entries substantially reduces both observed information and complete-sample support. Comparing methods at the same nominal missing rate can therefore mistake differences in data difficulty for algorithmic differences.
This is not merely an evaluation-label problem: it concerns which observations activate a reconstruction branch. A loss requiring every view depends on complete samples; a loss requiring a particular view pair depends on that pair's co-observation probability. These supports are not interchangeable. Separately, retraining for each missing configuration asks whether a model can adapt to a known missing distribution, whereas training once with all views available and making no deployment updates asks whether a fixed model handles changing observed subsets. The paper studies both questions but reports them separately: the information advantage of complete-view pretraining must not be presented as an information-matched incomplete-training advantage.
CRAFT consequently does not reject reconstruction outright. Instead, it separates supervision for representation learning from the architecture accepting missing inputs at inference. Reconstruction remains the main training signal, while missing views are strictly excluded from attention at inference, without requiring imputation before clustering. Core Idea: characterize the task through observation structure rather than nominal missing rate alone, and use a per-sample, mask-aware shared fusion model to evaluate incomplete training separately from fixed-checkpoint deployment across missing configurations.
Method¶
Overall Architecture¶
The paper first audits missing protocols and their hidden observation distributions, then introduces CRAFT (Co-occurrence-free Robust Attention-masked Fusion Transformer). CRAFT receives a sample's observed views and missingness mask, encodes features with view-specific MLPs, fuses them through a Transformer with a learnable CLS token, and outputs a cluster distribution through a cosine softmax head.
The network comprises “Per-sample masked fusion,” “Two-stage representation and cluster training,” and “Masked fine-tuning,” in that order. The latter two belong only to training: default deployment training starts with all views available, learns representations in Stage 1, aligns cluster structure through self-labeling in Stage 2, and optionally adds Masked Fine-Tuning (MFT). Once the final checkpoint is frozen, inference retains only encoding, masked fusion, and the shared cluster head; it performs no reconstruction decoding, pseudo-label refresh, or optimization.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Sample views + mask"] --> B["Per-sample masked fusion<br/>View encoding → CLS → cosine head"]
B -->|Training: all views by default| C["Two-stage representation<br/>and cluster training<br/>Reconstruction and consistency → self-labeling"]
C -->|Training: optional continuation| D["Masked fine-tuning<br/>View subsets and KL consistency"]
C -->|MFT disabled| E["Freeze final checkpoint"]
D --> E
E -.->|Fixed parameters| B
B -->|Inference: observed views only| F["Cluster distribution / label"]
The diagram describes the network rather than turning the four protocols into serial modules: protocols determine input-mask distributions, not four processing stages of CRAFT. Information-matched experiments instead train a separate model using each condition's genuinely available inputs and supervision; they do not reuse the default complete-view supervision shown above.
Key Designs¶
1. Per-sample masked fusion: exclude absent views from the current sample's computation
CRAFT uses two architectural conditions. C1 states that, with shared parameters fixed, a sample's fused representation depends only on its own observed views, not other samples' observation patterns. C2 requires the fusion function to be defined for every nonempty observed subset and to aggregate only observed views. C1 does not prohibit cross-sample training statistics: the batch-averaged entropy regularizer in Stage 2 aggregates predictions across samples. It restricts the data dependence of forward representations and does not require that training never observes complete samples.
View-specific encoders map heterogeneous feature dimensions into a common embedding space, preceded by an always-valid CLS token. The fixed-slot implementation retains every view position, but a per-sample key-padding mask sets absent positions' attention logits to negative infinity. These positions consequently leave the effective softmax denominator and receive exactly zero weight. Zero values or placeholder contents in absent positions no longer affect CLS aggregation. This differs from merely filling zeros, using mean features, or learning a missing token: without masking, these substitutes can still absorb attention mass.
The CLS hidden state becomes the fused representation. A cosine softmax head computes scaled cosine similarities against unit-norm cluster prototypes and produces cluster probabilities. Every observed subset shares the fusion network and clustering rule, without a separate head for each missing pattern. Attention can reweight observed views according to their content, but accepting any nonempty subset does not imply unchanged accuracy when an informative view disappears. Fixed-slot matrices also do not automatically shrink when fewer views are observed; variable observed subsets must not be equated with inference complexity necessarily decreasing with the retained-view count.
2. Two-stage representation and cluster training: preserve multi-view structure before shaping cluster structure
Stage 1 decodes the fused representation back into each view's features, using squared reconstruction error to preserve input structure. It also adds SimSiam-style consistency between different view embeddings of the same sample. An embedding from one view passes through a predictor and aligns with another view's stop-gradient embedding; the objective averages negative cosine similarity over ordered view pairs. This is neither a contrastive loss treating other samples as negatives nor consistency applied only to final cluster probabilities.
All views are available in default deployment training, so every view supplies a reconstruction target and every view pair supplies consistency supervision. Information-matched incomplete training instead encodes observed inputs only, reconstructs observed entries only, and averages consistency over eligible observed pairs; singleton samples have zero consistency loss. “Co-occurrence-free” in the method name must not be interpreted as every training signal being independent of co-observation: default training uses complete data, and pairwise consistency requires eligible view pairs. What is removed is the inference-time input dependence on other samples' completeness and on reconstructing absent views.
Stage 2 moves beyond explaining inputs to producing useful clusters. At the start of every epoch, pseudo-labels are refreshed by selecting each sample's most probable cluster. Cross-entropy on cosine softmax outputs then jointly fine-tunes the representation and cluster head. Negative entropy of batch-averaged predictions acts as an anti-collapse regularizer: minimizing it encourages a more uniform average cluster distribution rather than assigning every sample to one cluster. Ordinary Stage 2 disables the subset KL term. Self-labeling thus begins from a representation that already captures data structure rather than relying on pseudo-label reinforcement from random features.
3. Masked fine-tuning: continue subset-robustness training with the deployment masking mechanism
MFT continues optimization from the Stage 2 checkpoint rather than retraining separately for every test protocol. For each sample, it randomly draws how many views to delete and then selects the deleted views, always retaining at least two. Two sampled subsets of the same sample pass through masked fusion. Their cluster predictions receive a KL consistency term in addition to self-labeling and entropy regularization. This encourages a shared clustering rule to remain coherent across observed subsets without fitting a dedicated model to one deployment missing rate.
The deletion count is uniform from zero through the total view count minus two. Subset sizes therefore receive equal probability, but individual subsets do not. With six views, 57 subsets retaining at least two views participate; with two views, the only possible deletion count is zero, so MFT deletes no views. Any performance change in the latter case concerns additional fine-tuning or separate feature noise, not learning single-view masking. Inference still supports singleton inputs, but the MFT sampling rule itself does not cover that boundary.
4. Protocol and supervision-support audit: distinguish lost information from losses that can still learn
The nominal missing rate is \(r\), the fraction of genuinely missing sample–view entries is \(\hat r\), and the proportion of samples retaining every view is \(p_c\). Mask entry \(M_v^{(i)}=1\) means view \(v\) of sample \(i\) is observed. The definitions below must be used rather than directly treating \(r\) as the realized entry-wise missing rate.
All four protocols retain at least one view per sample, but differ in deletion budgets and allocation rules. The formulas below hold for fixed view count as the sample count tends to infinity. At finite sample size, Protocol 4 uses fixed-budget deletion without replacement, not independent Bernoulli entries.
| Protocol | Deletion mechanism | Effective missing rate \(\hat r\) | Complete-sample proportion \(p_c\) |
|---|---|---|---|
| P1 | Select fraction \(r\) of samples; randomly delete one view from each | \(r/V\) | \(1-r\) |
| P2 | Protect one view per sample; independently delete other views with probability \(r/(V-1)\) | \(r/V\) | \((1-r/(V-1))^{V-1}\) |
| P3 | Independently delete all entries with probability \(r\); restore one random view in empty samples | \(r-r^V/V\) | \((1-r)^V\) |
| P4 | Protect one view per sample; delete a fixed budget without replacement from other entries | \(\min(r,(V-1)/V)\) | \([\max(0,1-Vr/(V-1))]^{V-1}\) |
The “hidden distribution” is therefore the observation-mask distribution behind a nominal missing rate, not merely an unreported average. With six views and nominal rate 0.5, P1 retains 50% complete samples and P4 approximately 1.024%, a roughly 49-fold difference; their effective missing rates are also approximately 8.3% and 50%, respectively. This reveals protocol divergence but does not control total observed information, so accuracy differences cannot be attributed causally to complete-sample proportion alone. Even equal \(\hat r\) and \(p_c\) can hide different retained-view identities, view-pair support, and information content.
The theory has two layers. The capability bound applies to any method: the observed view set contains at least as much cluster-relevant mutual information as its most informative individual view. The upper bound by the sum of individual-view mutual information requires conditional independence of views given the cluster label and finite information quantities. A fixed encoder also obeys data processing, but the observed-set lower bound does not transfer to a learned representation: a constant encoder can discard all information. Deterministic imputation from observed views adds no inference-time information, without negating the useful gradients reconstruction can provide during training.
The trainability bound concerns only a support-gated reconstruction branch. Define \(q_2(P)=\Pr_{M\sim P}(|O|\geq2)\). When the loss is normalized over all sampled examples, examples are sampled uniformly, and each eligible sample's gradient is uniformly bounded by \(C_g\) along the trajectory, the branch's expected gradient norm is at most \(C_gq_2(P)\). If only this loss is optimized using constant-step SGD with step size \(\eta\) and fresh independent samples, expected parameter displacement after \(T\) steps is at most \(\eta TC_gq_2(P)\). Only a strictly complete-sample branch replaces \(q_2\) with \(p_c\); a particular view pair requires its own co-observation probability.
These assumptions are part of the conclusion rather than technical footnotes. Normalizing only over eligible samples makes the corresponding strict-complete gradient bound depend on the probability that a batch contains an eligible sample, rather than linearly multiplying by \(p_c\). AdamW also does not directly inherit the SGD displacement result. Other losses can update shared parameters; small displacement need not imply random clustering, and positive support does not guarantee accuracy. For six-view P1 at \(r=1\), every sample still retains five views: \(p_c=0\) but \(q_2=1\). Thus “no complete samples means all cross-view reconstruction fails” is an invalid generalization.
A Worked Example¶
Consider a three-view sample. During training, all three feature sources are available: Stage 1 reconstructs them from the CLS representation and uses each view embedding as a stop-gradient consistency target for the others. Stage 2 refreshes the sample's pseudo-label from its currently most likely cluster, aligning the representation with shared cluster prototypes.
In MFT, zero or one view can be deleted on each draw. Two sampled subsets might retain views 1 and 2, and views 1 and 3, respectively. They use identical network parameters but different key-padding masks; KL consistency constrains their cluster predictions. If deployment leaves only view 2, the network still performs a valid forward pass: views 1 and 3 receive zero attention weight, and CLS aggregates view 2 before passing to the existing cluster head. This is an architectural input boundary, not proof that MFT trained on singleton samples.
The example explains why complete-data training followed by frozen deployment is not equivalent to retraining with incomplete data alone. In the former, three-view supervision has already shaped representations and cluster prototypes. In the latter, missing targets are unavailable from the outset, and losses must be constructed from actual observed information. The paper treats them as separate experiments rather than substitutes.
Loss & Training¶
The two-stage objectives clarify their respective roles: Stage 1 combines reconstruction and embedding consistency; Stage 2 combines self-labeled cross-entropy, batch-averaged negative entropy, and MFT subset KL. Ordinary Stage 2 sets \(\gamma=0\); subset KL is used only during MFT continuation.
The default configuration uses a single-layer, four-head Transformer, AdamW, batch size 256, and weight decay \(10^{-4}\). Embedding dimension is 128 for CUB and 256 for default HandWritten. MultiFashion uses shallow encoders and 200 epochs for each ordinary stage, whereas CUB uses 100 epochs per stage. Deployment configuration is selected once and the final checkpoint reused without protocol-specific test adjustments; the paper acknowledges that historical records do not fully establish the checkpoint-selection criterion for every run.
MFT adds training budget: 100 post-Stage-2 epochs for HandWritten and Out-Scene, and 300 for MultiFashion, with learning rate \(10^{-5}\). These epochs are not included in ordinary Stage 2 budgets. Default CUB deployment results disable MFT. CRAFT-Core retains Stage 1 reconstruction, the cosine head, and attention masking while jointly removing several training components; its gap to canonical CRAFT is not an isolated MFT ablation.
Key Experimental Results¶
Main Results¶
The following selection from original Table 2 reports ACC (%). CRAFT per-condition trains independently for each condition; CRAFT train-once starts with complete-view data and reuses a checkpoint. Reported per-condition CUB and MultiFashion entries constitute the information-matched comparison. HandWritten is a separate missing-rate-matched control and must not be included in the “12/13” count.
| Dataset | Protocol / \(r\) | CRAFT per-condition | CRAFT train-once | Highest baseline in this condition |
|---|---|---|---|---|
| CUB | P1 / 0.1 | 87.17 | 82.09 | Energy-DIMC 75.20 |
| CUB | P1 / 0.7 | 69.00 | 73.61 | DCG 63.50† |
| CUB | P4 / 0.3 | 72.00 | 74.15 | DCG 65.39† |
| CUB | P4 / 0.5 | N/A | 68.99 | Energy-DIMC 42.38 |
| HandWritten | P4 / 0.5 | 91.05 | 93.69 | Energy-DIMC 92.75 |
| HandWritten | P4 / 0.7 | 82.56 | 85.21 | Energy-DIMC 83.26† |
| MultiFashion | P1 / 0.1 | 98.07 | 93.21 | HSACC 97.46 |
| MultiFashion | P1 / 0.7 | 89.61 | 91.97 | COMPLETER 90.09 |
| MultiFashion | P4 / 0.5 | 86.59 | 88.12 | DVIMC 80.43 |
| MultiFashion | P4 / 0.7 | N/A | 85.41 | Energy-DIMC 29.14 |
† denotes one effective independent seed. Unmarked CRAFT and HSACC entries are five-seed means; other baselines follow the independent-run records in Appendix A.1 and must not all be described as five-seed results. N/A means an unreported incomplete-training result, not failure or zero accuracy. DVIMC and COMPLETER fail at MultiFashion P4 / 0.7 and are excluded from selecting the highest baseline. In Table 2, HandWritten enables MFT only for train-once, not per-condition; Appendix D.6 disables MFT in both variants using a separate recipe.
Across the full information-matched comparison, CRAFT leads all six CUB conditions and six of seven MultiFashion conditions. The exception is MultiFashion P1 / 0.7: COMPLETER scores 90.09, Energy-DIMC 89.90, and CRAFT per-condition 89.61. Fixed checkpoints cover seven datasets with a four-protocol by four-rate deployment grid per dataset, but supplementary UCI-Digit, Out-Scene, and Caltech do not rerun baselines, and YTF-31 uses one seed. This is not equally strong SOTA verification on all seven datasets.
Ablation Study¶
The following selection from Appendix Table 20 reports three-seed ACC means ± standard deviations under P4. These historical ablations fix the learnable-placeholder recipe rather than the canonical attention-masking recipe of Table 2. Comparisons must therefore use their within-table Full reference, not subtract these numbers from main results to infer component effects.
| Dataset / \(r\) | Full | No Stage 2 | No Entropy | Recon Only | Repr Only | No Pretrain |
|---|---|---|---|---|---|---|
| HandWritten / 0.5 | 74.42 ± 1.3 | 73.62 ± 1.5 | 73.47 ± 1.6 | 74.22 ± 5.8 | 11.33 ± 0.1 | 43.15 ± 6.4 |
| HandWritten / 0.7 | 54.27 ± 1.9 | 53.67 ± 1.9 | 53.20 ± 1.7 | 54.83 ± 6.8 | 10.46 ± 0.1 | 31.40 ± 4.4 |
| CUB / 0.3 | 68.62 ± 4.4 | 65.63 ± 1.9 | 66.03 ± 2.1 | 73.93 ± 4.8 | 31.60 ± 1.8 | 49.34 ± 1.3 |
| CUB / 0.5 | 61.54 ± 5.7 | 57.99 ± 2.5 | 58.42 ± 2.7 | 69.28 ± 5.5 | 15.34 ± 1.1 | 47.49 ± 3.3 |
Recon Only removes Stage 1 consistency while retaining subsequent training; Repr Only removes Stage 1 reconstruction; No Pretrain skips Stage 1 entirely. At HandWritten / 0.5, Repr Only is 63.09 percentage points below Full, identifying reconstruction as a crucial anchor in this recipe. At CUB / 0.5, Recon Only is instead 7.74 points higher, so consistency is not universally beneficial. The table does not establish statistical significance.
Key Findings¶
- Table 4 isolates inference-time missing-view handling: with the same complete-view-trained checkpoint at CUB P4 / 0.5, attention masking scores 63.84 ± 7.9, learnable placeholders 61.54 ± 5.7, zero filling 58.13 ± 4.4, and mean filling 55.89 ± 0.9, all over three seeds. Table 2's 68.99 must not replace the masking number because the recipes differ.
- Table 3 compares fusion blocks with shared encoders and two-stage training, all without MFT. At HandWritten P4 / 0.7, Transformer scores 79.06 ± 5.7, SetMLP-Light 62.09 ± 2.3, SetMLP-Match 55.15 ± 1.9, and concat 56.55 ± 0.7, all over five seeds. Pooling architectures satisfying C1+C2 also run successfully but perform worse here; this does not prove Transformer is the only viable architecture.
- Two-view P4 caps effective missing rate at 0.5, so nominal rates 0.5 and 0.7 share the same saturated mask distribution. CUB train-once scoring 68.99 in both columns is not independent evidence of robustness to a stronger missing condition. Three-view P4 instead caps the rate at \(2/3\).
- MultiFashion timing records give 128 minutes for sixteen DVIMC runs versus 14.5 minutes for one CRAFT run. The approximately 8.8-fold ratio concerns cumulative grid cost; one DVIMC run takes only 8 minutes and is shorter. Records do not itemize model selection, search, or extra MFT, so they do not show an 8.8-fold speedup of the complete development pipeline.
Highlights & Insights¶
- Separating nominal missing rate into coverage and supervision support exposes hidden difficulty differences in older protocols. Checking complete-sample or view-pair support for each actual loss is more diagnostic than reporting only overall missingness.
- Strict masking is more explicit than learning a “missingness meaning”: zero attention to absent positions is structurally guaranteed, whereas higher accuracy remains an empirical claim. Correct missing-input handling and successful representation learning are distinct conclusions.
- Reconstruction is retained as training supervision shaping representations rather than as mandatory deployment imputation. Transferring this separation to sensor fusion should likewise evaluate information-matched training separately from fixed-model deployment to avoid conflating information budgets.
Limitations & Future Work¶
- Missingness mainly follows four controlled synthetic protocols; naturally occurring, quality-dependent, and correlated sensor failures remain untested. The two summary statistics also do not fully describe retained-view identity at equal coverage.
- Default deployment training requires all views to be available initially. A modality never observed during training and strictly zero-complete-sample training beyond reported conditions remain open. MFT retains at least two views, leaving adaptation to singleton deployment as another boundary.
- C1+C2 are forward-computation conditions, not a successful-learning theorem. Support-gated bounds depend on normalization, bounded gradients, sampling, and optimizer assumptions; they cannot assign a common failure cause to all reconstruction methods, distributional methods, or low-accuracy runs.
- Historical tables mix independent-run counts, training recipes, and completeness of checkpoint records, limiting significance and fairness interpretations. Further evaluation should standardize seeds, information budgets, model selection, and full cost accounting, and add real missingness and imbalanced-cluster tests.
Related Work & Insights¶
- vs COMPLETER / DCP / DCG: These methods train recovery or prediction branches from co-observed information, but their support requirements differ. The paper's multi-view COMPLETER extension uses pairwise-complete training and still scores 68.76 ± 5.34 at HandWritten P1 / 1.0, demonstrating that zero complete samples does not invalidate every pairwise reconstruction objective.
- vs Energy-DIMC: Its cross-sample or distributional supervision lies outside the support-gated reconstruction gradient bound. Strong results in some stringent conditions show why performance cannot be explained solely by whether imputation is used; its per-configuration training also differs informationally from CRAFT's default deployment setting.
- vs DVIMC / FreeCSL / I2MVC: Per-sample fusion without explicit recovery is not introduced by CRAFT. Its concrete contribution combines masked attention, a shared cluster head, and separated training/deployment evaluation. DVIMC initialization failures concern the evaluated implementation, not the entire architectural family.
- vs Set Transformer / missing-modality Transformers: Set attention and observed-modality fusion have established foundations. This paper applies them to unsupervised IMVC with protocol auditing; the reusable value of imvc-audit is checking actual mask structure rather than simply adding another network.
Rating¶
- Novelty: 4/5. Protocol auditing and training/deployment separation are valuable; attention fusion itself is not entirely new.
- Experimental Thoroughness: 4/5. Information-matched training, fixed checkpoints, and multiple diagnostics are included, but historical seeds and recipes are not uniform.
- Writing Quality: 4/5. Theoretical assumptions and support types are explicit, although experimental recipes require careful distinction.
- Value: 4/5. Provides an actionable framework for fair IMVC evaluation and dynamic missing-view deployment.