Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training¶
Conference: ECCV 2026
Paper: CVF Open Access
Code: https://github.com/MingliangLiang3/DynamiCS
Area: Multimodal VLM
Keywords: Vision-Language Pretraining, Dynamic Data Sampling, Cluster Scaling, Long-Tail Concepts, Computational Efficiency
TL;DR¶
Addressing the excessive computational cost and severe underrepresentation of rare concepts in vision-language pre-training on noisy web data, this paper proposes DynamiCS, a framework combining power-law cluster scaling with cross-epoch dynamic resampling that slashes GPU pre-training hours to ~3% while dramatically boosting long-tail recognition.
Background & Motivation¶
Large-scale vision-language models (VLMs) such as CLIP learn transferable multimodal representations via contrastive learning, serving as core backbones for zero-shot classification, cross-modal retrieval, and multimodal LLMs. However, full-scale pre-training typically requires billions of web-crawled image-text pairs and tens of thousands of GPU hours, creating a massive computational bottleneck for reproducing and iterating on these architectures. To alleviate these costs, existing research has primarily explored two avenues: token-reduction methods that compress per-sample computation via lower image resolutions, patch masking, or syntax-based text masking (e.g., RECLIP, FLIP, CLIPA); and dataset pruning methods that filter out redundant or poorly aligned pairs (e.g., DataComp, DFN, DBP, MetaCLIP).
Although dataset pruning reduces overall training volume, raw web-crawled data naturally exhibits a severely skewed long-tail distribution: a handful of common topics contain massive redundant instances (the fat head), whereas vast numbers of fine-grained, niche concepts contain very few samples (the long tail). Existing semantic balancing approaches, such as MetaCLIP's hard truncation to 20k per metadata class or DBP's cluster complexity pruning, adhere to an "aim for even" philosophy. By aggressively flattening density across categories, these methods disproportionately discard or outright eliminate rare long-tail concepts and disrupt natural semantic hierarchies. Furthermore, dual-purpose data curation methods like DFN and HQ-CLIP rely on another pre-trained teacher VLM to filter pairs or synthesize descriptive captions, introducing substantial additional dependencies and overhead.
This paper shifts the paradigm from blindly pursuing uniform distributions to an "aim for utility" philosophy. The key insight is that pre-training does not require flattening semantic distributions; instead, by preserving the natural relative ordering of semantic clusters while using a smooth power-law scaling factor to moderately downsample the head and upsample sparse clusters with replacement, combined with dynamic resampling across epochs, models can achieve superior efficiency and representation coverage. Core idea: DynamiCS introduces dynamic cluster-based data sampling that downsamples large semantic clusters via power-law scaling while upsampling rare clusters with replacement, dynamically redrawing distinct subsets at each epoch to achieve competitive accuracy with full-scale pre-training at only ~3% of the computational cost.
Method¶
Overall Architecture¶
The DynamiCS pipeline consists of four coordinated stages: offline semantic clustering, power-law cluster quota calculation, cross-epoch dynamic resampling, and two-stage efficient pre-training. First, image embeddings are extracted from unlabeled datasets using a pre-trained visual encoder (DINOv2-ViT-B/16), grouped into tens of thousands of semantic clusters via spherical K-Means, and merged if centroid cosine similarity exceeds a threshold. Second, target sampling capacities are computed for all clusters via power-law scaling, balancing head compression with long-tail oversampling. Third, during pre-training, a fresh random subset is drawn dynamically from each cluster at every epoch, preventing sample ossification. Finally, the model trains predominantly at a reduced resolution (\(112 \times 112\)) and truncated text length, followed by a brief high-resolution (\(224 \times 224\)) adaptation phase to bridge the distribution gap.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Raw Web-Scale Multimodal Data<br/>LAION-400M / DataComp"] --> B1["Offline Clustering & Centroid Merging<br/>DINOv2 Features + K-Means + Cosine Deduplication"]
B1 --> B["Cluster Scaling<br/>Power-law downsampling of head & upsampling of tail"]
B --> C["Dynamic Sampling<br/>Per-epoch random subset draw for cross-epoch diversity"]
C --> D["Two-Stage Efficient Pre-training<br/>112ร112 large-batch pre-training + 224ร224 adaptation"]
D --> E["Long-Tail-Aware Multimodal Representations<br/>ImageNet / Let It Wag! / Retrieval / Robustness"]
Key Designs¶
1. Cluster Scaling: Preserving Semantic Order while Upsampling Long-Tail Concepts
Targeting the limitation of traditional pruning methods that either randomly discard rare samples or completely eliminate long-tail clusters via hard thresholds, this design establishes a cluster scaling formula that preserves relative semantic order while actively emphasizing the tail. Once embeddings are partitioned into \(N\) clusters with original sizes \(c_i\), setting a target total sample budget \(T\) (e.g., 50% of the dataset) determines the resampled size \(S_i\) of cluster \(i\) via scaling exponent \(\alpha \ge 0\):
The parameter \(\alpha\) smoothly controls the distribution profile: \(\alpha = 0\) corresponds to extreme uniform allocation where every cluster receives an identical quota regardless of size; \(\alpha = 1\) reduces to standard unbiased random sampling; and \(\alpha > 1\) severely amplifies the fat head. Setting \(\alpha = 0.2\) achieves the optimal utility trade-off: large redundant clusters are downsampled, while preserving the natural precedence of dominant concepts. Crucially, whenever \(S_i > c_i\) for sparse tail clusters, samples are drawn with replacement, ensuring rare concepts receive sufficient gradient updates during pre-training.
2. Dynamic Sampling: Expanding Cross-Epoch Diversity without Additional Cost
To counter the severe risk of overfitting caused by repeatedly training on an identical static subset of upsampled tail instances, and to avoid permanently discarding half of the head cluster samples, this design applies dynamic resampling at every epoch. In each training epoch, samples are drawn from cluster \(i\) according to the sampling ratio:
For downsampled large clusters (\(P_i < 1\)), the data loader draws a different random subset in each epoch, effectively allowing the model to see a much larger fraction of the overall dataset across multiple epochs compared to static pruning. For upsampled small clusters (\(P_i > 1\)), drawing with replacement per epoch injects dynamic stochasticity and stabilizes gradient variance. This dynamic selection substantially expands cross-epoch data diversity while keeping the exact computational budget per epoch strictly capped at \(T\).
Loss & Training¶
The network is optimized using standard symmetric InfoNCE contrastive loss over paired visual and textual embeddings. Training employs an aggressive efficiency curriculum: during the initial epochs (scaling across 0.64B to 2.56B total samples seen), images are processed at \(112 \times 112\) resolution with a large batch size of 28k and a maximum text sequence length of 32 tokens. After low-resolution pre-training, the model undergoes a short fine-tuning stage for 128M samples seen at the standard \(224 \times 224\) resolution to eliminate resolution discrepancy before downstream evaluation.
Key Experimental Results¶
Main Results¶
The models were evaluated using the ViT-B/16 architecture pre-trained on LAION-400M and DataComp-DFN. Zero-shot top-1 classification was tested on ImageNet-1K and the long-tail benchmark Let It Wag! (containing 290 tail concepts with 130k test images).
| Models | Dataset (Data Size) | Samples Seen @ Resolution | ImageNet-1K Top-1 (%) | Let It Wag! Top-1 (%) | GPU Hours |
|---|---|---|---|---|---|
| OpenCLIP | LAION-400M (400M) | 12.8B @ 224 | 67.1 | 39.1 | 10736 |
| MetaCLIP-400M | LAION-400M (400M) | 12.8B @ 224 | 70.8 | 46.5 | โ10700 |
| RECLIP* | LAION-400M (298M) | 2.56B@112 + 128M@224 | 62.9 | 36.0 | 280 |
| CLIPA* | LAION-400M (298M) | 2.56B@112 + 128M@224 | 63.2 | 36.4 | 269 |
| DynamiCS (Ours) | LAION-400M (298M) | 1.28B@112 + 128M@224 | 65.0 | 42.1 | 163 |
| DynamiCS (Ours) | LAION-400M (298M) | 2.56B@112 + 128M@224 | 67.5 | 45.5 | 299 |
| DataComp (ImageโฉCLIP) | DataComp (1.28B) | 1.28B @ 224 | 63.1 | 33.7 | โ1070 |
| DFN* | DataComp-DFN (130M) | 1.28B@112 + 128M@224 | 68.7 | 42.4 | 151 |
| HQ-CLIP | DataComp-DFN | 3.20B @ 224 | 70.6 | 38.2 | โ2675 |
| DynamiCS (Ours) | DataComp-DFN (130M) | 0.64B@112 + 128M@224 | 69.2 | 46.5 | 95 |
| DynamiCS (Ours) | DataComp-DFN (130M) | 1.28B@112 + 128M@224 | 71.3 | 50.2 | 163 |
| DynamiCS (Ours) | DataComp-DFN (130M) | 2.56B@112 + 128M@224 | 72.6 | 52.0 | 299 |
Ablation Study¶
The ablation experiments isolate the contributions of cluster scaling versus dynamic sampling, alongside the sensitivity of the scaling factor \(\alpha\).
Table 1: Sampling Strategies and Dynamic Mechanism Ablation (DataComp 0.64B@112 + 128M@224)
| Config | ImageNet-1K Top-1 (%) | Let It Wag! Top-1 (%) | Note |
|---|---|---|---|
| Random Pruning | 64.5 | 35.5 | Fixed static 50% subset |
| Random-Dynamic | 66.2 | 36.2 | Independent 50% dynamic sampling (+1.7% / +0.7%) |
| Cluster-Scaling (Static) | 68.0 | 43.7 | Cluster scaling with fixed subset (+3.5% / +8.2%) |
| DynamiCS (Full Model) | 69.2 | 46.5 | Cluster scaling + dynamic sampling (+4.7% / +11.0% over random) |
Table 2: Scaling Exponent \(\alpha\) Sensitivity Analysis (DataComp 106M samples seen @ 112ร112)
| Scaling Factor \(\alpha\) | ImageNet-1K Top-1 (%) | Let It Wag! Top-1 (%) | Behavior Note |
|---|---|---|---|
| \(\alpha = 0.0\) | 38.5 | 19.5 | Strictly uniform distribution across all clusters |
| \(\alpha = 0.2\) | 39.2 | 20.2 | Optimal utility trade-off (best overall and tail performance) |
| \(\alpha = 0.4\) | 38.2 | 19.6 | Moderate scaling |
| \(\alpha = 0.6\) | 36.9 | 17.5 | Diminishing long-tail upsampling |
| \(\alpha = 0.8\) | 36.4 | 15.6 | Approaching natural distribution proportions |
| \(\alpha = 1.0\) | 33.8 | 13.4 | Standard random sampling baseline |
| \(\alpha = 2.0\) | 19.4 | 5.1 | Severe over-concentration on head clusters |
Key Findings¶
- Long-tail upsampling is the primary driver of performance gains: Switching from random pruning to cluster scaling produces an immediate +8.2% jump on the long-tail benchmark Let It Wag! (35.5% to 43.7%), verifying that discarding rare concepts is the dominant source of error in previous efficient methods.
- Dynamic sampling mitigates overfitting while expanding coverage: Dynamic sampling adds +1.2% on ImageNet-1K and +2.8% on Let It Wag! over static cluster scaling. It broadens sample exposure for head clusters across epochs and prevents over-memorization of upsampled tail instances.
- Robustness across hyperparameter selections: Any \(\alpha\) between 0.0 and 0.8 clearly outperforms the unbiased baseline (\(\alpha = 1.0\)), demonstrating that DynamiCS does not require delicate hyperparameter tuning.
- Outperforming full-scale models on fine-grained and robustness tasks: DynamiCS achieves substantial gains on fine-grained benchmarks like CUB-200 (70.8% vs. 58.1% for DFN) and Flowers102 (83.5% vs. 73.2% for DFN), where differentiating long-tail classes is essential.
Highlights & Insights¶
- From "Aim for Even" to "Aim for Utility": Rather than forcing visual concept distributions into an artificial uniform flatland, preserving natural semantic ranking with a dampened power-law exponent (\(\alpha=0.2\)) yields superior representation quality for contrastive learning.
- Drastic 97% training cost reduction: DynamiCS achieves 72.6% on ImageNet-1K and 52.0% on Let It Wag! with only 299 GPU hours, outperforming OpenAI CLIP, OpenCLIP, and MetaCLIP which required upwards of 10,700 GPU hours.
- Zero reliance on external VLM or LLM teachers: Unlike DFN, DataComp, or synthetic captioning pipelines that require pre-existing models for filtering or recaptioning, DynamiCS relies solely on offline vision self-supervised clustering (DINOv2) and simple sampling rules.
Limitations & Future Work¶
- Visual-only clustering assumptions: Clustering relies on DINOv2 vision embeddings; polysemous images or scenes dominated by subtle textual relations rather than visual appearance might be misclustered into coarse semantic groups. Incorporating joint multimodal text-image embeddings could refine cluster boundaries.
- Initial clustering cost on web-scale data: While K-Means clustering and centroid deduplication are one-time offline operations, performing them over billions of web images requires notable preprocessing compute. Exploring streaming or approximate online clustering could streamline deployment.
Related Work & Insights¶
- vs MetaCLIP / DBP: MetaCLIP enforces hard truncation via metadata dictionaries and DBP flattens cluster densities; both focus solely on downsampling. DynamiCS proves that combining power-law downsampling with with-replacement upsampling on tail clusters provides superior coverage of rare concepts.
- vs DFN / HQ-CLIP / WhatIf: These methods depend heavily on pre-trained CLIP models for filtering or LLMs for synthetic recaptioning. DynamiCS operates independently of teacher VLMs and preserves original web data diversity.
- vs RECLIP / FLIP / CLIPA: These methods focus on reducing tokens per sample (lower resolution, patch masking, text masking). DynamiCS tackles semantic data distribution and selection, serving as an orthogonal enhancement that can be stacked with token-reduction techniques for higher efficiency.
Rating¶
- Novelty: โญโญโญโญโ Challenges the conventional wisdom of uniform balancing by introducing utility-driven relative-order scaling and dynamic upsampling.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive comparisons across LAION-400M and DataComp-DFN against full-scale, cost-reducing, and dual-purpose baselines.
- Writing Quality: โญโญโญโญโญ Rigorous motivation, coherent mathematical formulation, and exceptionally thorough empirical validation.
- Value: โญโญโญโญโญ Enables academic labs to train state-of-the-art vision-language models at a fraction of standard computational budgets while excelling on long-tail concepts.