๐ Self-Supervised Learning¶
๐ง NeurIPS2026 ยท 6 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (21) ยท ๐ท CVPR2026 (92) ยท ๐ฌ ICLR2026 (81) ยท ๐ฌ ACL2026 (1) ยท ๐งช ICML2026 (28) ยท ๐ค AAAI2026 (16)
๐ฅ Top topics: Self-Supervised Learning ร2
- Beyond Missing Rates: Rethinking Incomplete Multi-View Clustering with Protocol Divergence
-
The paper exposes evaluation-protocol divergence in incomplete multi-view clustering through effective missing rate and complete-sample proportion, and introduces CRAFT, which fuses only each sample's observed views: it wins 12 of 13 information-matched incomplete-training conditions, while a separate evaluation tests cross-protocol deployment of checkpoints trained with all views available.
- Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching
-
SFM approximately interprets linear gradient matching on frozen pretrained vision backbones as matching relative vectors between class centers, then supervises synthetic images with global statistics computed once, improving classification with one image per class; its approximately 10-fold memory reduction and fourfold speedup compare single-augmentation SFM against ten-augmentation LGM, rather than equal-budget configurations.
- Handwritten Text Recognition Lives in the High-Pixel Variance Subspace
-
Using pixel-PCA projections and frozen-encoder probes, this paper studies the distribution of handwritten text recognition signals: high-variance pixel directions support transcription better on six Latin-script benchmarks, while real-data-pretrained MAE achieves 5.5% mean probe CER and 4.5% mean CER after adding a language decoder and full fine-tuning.
- I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
-
StreamMAE preserves MAE's pixel reconstruction objective while adapting temporally ordered video training through stronger regularization, example-level DataDrop, and motion-biased two-stage cropping, improving classification, segmentation, and depth transfer on WT++12h and approaching same-data offline i.i.d. MAE on most metrics.
- Latent Video Prediction for World Modeling: An Evaluation Uncovering Intriguing Favorable Evidence
-
Under a unified readout protocol, this study evaluates four frozen video encoders, finds that V-JEPA variants better preserve usable action semantics under degraded inputs and support more reliable planning in limited simulated manipulation experiments, but neither evaluates generation quality nor causally isolates the pretraining objective.
- Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding
-
SGMA combines content-adaptive quadtree tokenization with cross-scale structure-conditioned masking for MAE pretraining, preserving scientific microstructure within a fixed sequence budget for standard ViTs; SGMA-SAM reaches 95.68% Dice on SpringXCT and SGMA-SAM 2 reaches 83.21% on PAIP, while the reported maximum 24.8ร inference speedup comes from a separate, shorter-sequence configuration.