Skip to content

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

Conference: NeurIPS2026 (task-list assignment; local full text is arXiv v3, 2026-09-30)
arXiv: 2602.05391
Code: https://github.com/einsteinxia/SFM
Area: Self-Supervised Learning; Dataset Distillation
Keywords: dataset distillation, statistical flow matching, frozen pretrained backbone, linear probing, classifier inheritance

TL;DR

SFM approximately interprets linear gradient matching on frozen pretrained vision backbones as matching relative vectors between class centers, then supervises synthetic images with global statistics computed once, improving classification with one image per class; its approximately 10-fold memory reduction and fourfold speedup compare single-augmentation SFM against ten-augmentation LGM, rather than equal-budget configurations.

Background & Motivation

Traditional dataset distillation aims to replace an entire training set with very few synthetic images, but one image per class often cannot support high accuracy when the downstream model is trained from scratch. This paper follows a setting better suited to pretrained models: CLIP, DINO-v2, EVA-02, or MoCo-v3 already supplies transferable visual features, and downstream training only fits a linear classifier. Distillation therefore need not teach a model visual recognition anew; it packages the training signals for a particular classification task into a small image set. The compressed object is the dataset, not the backbone, and the procedure is not another round of unlabeled pretraining.

The direct baseline, Linear Gradient Matching (LGM), randomly initializes a linear head at each step and encourages real and synthetic images to induce similar classifier gradients. However, its real-data target comes from changing local batches, while synthetic images require repeated differentiable augmentations and retained back-propagation graphs through a large backbone. The authors ask how much information still requires inner-loop differentiation when the backbone is frozen and the random head has small variance. If gradients mainly reflect geometry between class centers, repeatedly sampling real data and random heads may be an expensive way to estimate a target that could instead remain fixed.

The paper starts from the linear head's cross-entropy gradient, connects it to differences between class centers, and replaces a local, dynamic target with a global, fixed statistical target. Core Idea: first summarize each class's “flow” relative to all other classes in a frozen feature space, then directly optimize synthetic images to align their flows with these real-data statistics, without recomputing real-data classifier gradients at every step.

Method

Overall Architecture

The inputs are a labeled real dataset and a frozen pretrained vision backbone; the output is a very small labeled synthetic dataset, with the central experiments using images per class IPC=1. Global Statistical Flow is constructed first, followed by Synthetic Flow Alignment: real data is traversed during statistics computation, while optimization repeatedly augments synthetic images, extracts features, and updates their image representations through back-propagation. Backbone parameters remain unchanged throughout.

The standard downstream procedure trains a new linear head on synthetic images using the evaluation backbone; Classifier Inheritance (CI) is an optional alternative evaluation route. CI additionally retains a classifier trained on the complete real dataset and uses synthetic images to learn a linear projection between two backbones. Consequently, SFM alone and SFM+CI differ in delivered artifacts, training information, and computational costs.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Labeled real images<br/>Frozen distillation backbone"] --> B["Global Statistical Flow"]
    S["Trainable synthetic images<br/>Augmentation and frozen backbone"] --> C["Synthetic Flow Alignment"]
    B -.->|Fixed training supervision| C
    C --> D["Labeled synthetic dataset"]
    D -->|Standard evaluation| E["Train a new linear head<br/>Classify test images"]
    D -->|Optional evaluation training| F["Classifier Inheritance"]
    G["Frozen classifier trained<br/>on the full real dataset"] -.->|Additional artifact| F
    F --> H["Test images: evaluation backbone<br/>Projector and inherited classifier"]

Dashed edges denote fixed supervision or an additional artifact, while solid edges describe construction and use. The synthetic-flow loss applies only during distillation; test images do not participate in statistical flow matching or CI feature-alignment training.

Key Designs

1. Global Statistical Flow: extract precomputable class geometry from random linear-head gradients

LGM compares cross-entropy gradients with respect to linear weights, not gradients of every backbone parameter. Let the frozen backbone be \(\phi\), batch size be \(B\), and number of classes be \(C\); each sample has predicted probabilities \(p_i\) and a one-hot label \(y_i\). The exact gradient is:

\[ \frac{\partial \ell}{\partial W}=\frac{1}{B}\sum_{i=1}^{B}(p_i-y_i)\phi(x_i)^\top. \]

The key observation is not that all random classifiers are identical. Rather, when class weights are independently initialized from the same zero-mean Gaussian distribution, exchange symmetry makes each class probability have expectation \(1/C\). When logit variance is sufficiently small and the class count is large, the authors additionally approximate individual probabilities as close to \(1/C\). In a balanced batch with equal numbers of samples or augmented views per class, row \(c\) of the expected gradient becomes the mean non-target feature minus the mean target feature, multiplied by the common class-count-dependent factor \((C-1)/C^2\).

This gives “flow” a geometric meaning: it points from the target-class center toward the average center of all other classes. There is no continuous time, velocity-field network, or ODE sampling; the term only resembles flow matching in generative modeling. The common factor does not affect cosine matching, so center differences can directly define the matching object. The real-data statistical flow is:

\[ \mathcal{F}_c^*=\phi^*(x^o\mid y_c^o=0)^\top-\phi^*(x^o\mid y_c^o=1)^\top. \]

Here \(\phi^*\) denotes conditional means computed over the complete real dataset, not another learned encoder. The first term averages all non-target samples, and the second averages target-class samples. These vectors are stored and held fixed after precomputation, rather than drifting with each optimization batch. Class centers can be computed by accumulating sums and counts over small batches, without loading an entire class into GPU memory; the retained objects are statistics, not all real images or their computation graphs.

This interpretation has explicit conditions. Theorem 1 is stated using class-dependent variances \(\sigma_c^2\), but Appendix A.2 actually invokes identically distributed weights. Independence and zero means with arbitrarily different class variances do not establish general multiclass exchangeability. Appendix Equation (23) also writes the expectation of a ratio as a ratio of expectations, which is not a general identity; under identical distributions, symmetry and probabilities summing to one instead justify the uniform expectation. The initialization actually used in this paper has equal variances and satisfies that narrower condition.

2. Synthetic Flow Alignment: back-propagate only through synthetic images toward a fixed relative statistical target

Each iteration applies differentiable augmentation to synthetic images, then computes their class centers and synthetic flow \(\mathcal{F}^s\) through the frozen backbone. Flows for all classes are concatenated and flattened, and a single global cosine distance aligns real and synthetic statistics:

\[ \mathcal{L}_{\mathit{sfm}}=1-\cos[\operatorname{flat}(\mathcal{F}^*),\operatorname{flat}(\mathcal{F}^s)]. \]

This neither minimizes a separate pixel error for each class nor requires synthetic images to sit at absolute real-data class centers. Relative center differences preserve relationships between classes that matter for classification; global cosine matching also permits a common scale mismatch, allowing synthetic images to lie near more discriminative boundaries. Once the matching target is fixed, a single augmentation can provide useful optimization signals without many views compensating for changing local real-data targets.

Compared with LGM, SFM removes per-step real-data loading, the random linear head, and inner gradient computation, but it still propagates the loss through a large frozen backbone to update synthetic images. Frozen weights do not permit disabling automatic differentiation on the synthetic-image branch: more augmented views still require more intermediate activations. This explains why memory usage is very similar at equal APB, whereas low-APB SFM versus high-APB LGM produces the order-of-magnitude difference.

Statistically, global means remove the batch-sampling fluctuations considered here, but do not fully represent within-class multimodality, covariance, or every downstream decision boundary. A global statistical target does not prove that image optimization reaches a global optimum. The small-variance derivation also does not make the nonlinear random LGM cosine objective strictly equivalent to SFM. It is better treated as a mechanism interpretation and simplification rationale whose usefulness is tested empirically.

3. Classifier Inheritance: align representations using synthetic images and reuse a real-data decision boundary

Across backbones, distilled images inherit the feature geometry of the distillation backbone \(\phi_d\), which may not suit another evaluation backbone \(\phi_e\). Rather than estimating an entire decision boundary again from one labeled image per class, CI first trains a classifier \(f\) on complete real-data features from the distillation backbone and freezes it. It then trains a single-layer linear projector \(\mathcal{P}\) only on synthetic images, mapping evaluation features into the distillation feature space:

\[ \mathcal{L}_{eval}=\|\phi_d(x^s)-\mathcal{P}(\phi_e(x^s))\|_2^2. \]

This stage performs feature regression, needs no synthetic labels to calculate its loss, and updates neither backbone. New test images follow “evaluation backbone → trained projector → inherited classifier.” The projector addresses both dimensionality differences and representation alignment, while the classifier supplies a boundary already learned from the complete real dataset. This is not evidence that one image alone recovers the same boundary information from scratch.

Appendix D trains this additional classifier for only 10 epochs and reports its real-data accuracy. Relative to soft-label methods, CI avoids tuning the balance between label cross-entropy and KL divergence, but requires retaining the classifier artifact and accessing the distillation backbone when training the projector. Main experiments do not use CI by default; only explicitly labeled SFM+CI results include this route, so it should not be folded into fair comparisons of SFM alone.

Loss & Training

Synthetic images use LGM's pyramid representation rather than a learned generator. Appendix B uses Adam with learning rate 0.002 for 5000 distillation iterations, adding a pyramid level every 200 iterations. Augmentations include random horizontal flipping, random resized cropping to 224×224, and Gaussian noise with standard deviation 0.2. APB denotes augmentations per batch and defaults to 1 for SFM. Multi-augmentation SFM uses 5, 8, and 10 views on CUB-200, Stanford Dogs, and ImageNet-100, respectively; LGM uses 3 on ImageNet-1k and 10 on the other stated datasets. “Multi-augmentation” therefore does not always imply equal view budgets between methods.

Standard evaluation trains a linear head for at most 1000 epochs with batch size 100, Adam initial learning rate 0.001/256, and cosine decay. The appendix explicitly stops training when test accuracy fails to improve for 50 epochs. This uses the test set for model selection and introduces selection bias; stopping should instead be selected on an independent validation set before test results are reported again. Existing numbers should not be read as fully isolated final-test estimates.

The authors report using an earlier augmentation implementation with Gaussian noise mean 0.5 rather than 0. Appendix Table 7 shows that switching to zero mean raises EVA-02-distilled accuracy on MoCo-v3 from 70.5 to 76.6, but changes same-backbone EVA-02 accuracy from 88.9 to 88.6. Thus the stated improvement from correcting the noise is not monotonic across all models. Reproduction should document this implementation-version difference.

Key Experimental Results

Main Results

Images have resolution 224×224, the main experiments use four ViT-B backbones, and results are means over 3 trials. The following selection from original Table 2 reports average accuracy (%) over four backbones at IPC=1 when distillation and evaluation use the same backbone. The “±” values preserve the paper's reported aggregates and are not reinterpreted as confidence intervals. All numbers exclude CI.

Table 1: Same-backbone classification results.

Dataset LGM, single augmentation SFM, single augmentation LGM, multiple augmentations SFM, multiple augmentations Full real dataset
ImageNet-100 83.5±0.1 87.7±0.1 87.2±0.1 88.5±0.1 92.8±0.1
ImageNet-1k 64.6±0.1 69.3±0.1 67.9±0.0 Not reported 80.0±0.0
Stanford Dogs 55.8±0.2 69.2±0.3 69.9±0.2 71.9±0.1 80.1±0.2
CUB-200 43.6±0.1 66.8±0.2 66.2±0.2 69.4±0.1 73.8±0.5

Single-augmentation SFM substantially improves over single-augmentation LGM, but does not universally exceed multi-augmentation LGM: the Stanford Dogs averages remain 69.2 versus 69.9, and individual DINO-v2 results on ImageNet-100 are 90.6 versus 91.5. The complete dataset also retains a clear advantage. One image under strong pretrained priors should not be described as an equivalent replacement for all real training data.

Ablation Study

Table 2 is selected from original Table 4, uses CLIP for distillation, and reports average accuracy across four evaluation backbones (%, including the distillation backbone itself). TCDD matches only target-class centers, NCDD matches only non-target-class centers, and their combination is SFM's relative flow.

Table 2: Contributions of target and non-target statistics.

Config ImageFruits ImageNet-100 ImageNet-1k
NCDD alone 15.8±1.2 0.77±0.0 0.10±0.1
TCDD alone 77.6±0.7 79.4±0.1 57.9±0.0
TCDD+NCDD 76.5±1.1 81.5±0.2 60.7±0.1

Non-target centers alone cannot supply target-class identification signals. Adding them increases average accuracy by 2.1 and 2.8 percentage points on ImageNet-100 and ImageNet-1k, but reduces it by 1.1 points on ten-class ImageFruits. These are differences between reported means, not claims of statistical significance. Benefits from relative geometry depend on task scale rather than unconditionally exceeding absolute-center matching.

Table 3 is selected from Appendix Table 13 with EVA-02 as the distillation backbone; together with Figure 1, this efficiency comparison concerns ImageNet-100 at IPC=1. Memory and time are reported distillation-stage measurements, not total end-to-end deployment costs.

Table 3: Efficiency at equal and different augmentation budgets.

Method APB GPU memory (GB) Time (minutes)
LGM 1 16.6 21
SFM 1 16.3 18
LGM 10 165.2 81
SFM 10 164.5 62

The approximately tenfold memory difference and fourfold time difference compare SFM/APB=1 with LGM/APB=10. At fixed APB=1, the comparison is 16.3 versus 16.6 GB and 18 versus 21 minutes; at fixed APB=10, it is 164.5 versus 165.2 GB and 62 versus 81 minutes. Equal-budget results support runtime savings from removing data loading and gradient computation, not an order-of-magnitude memory reduction at the same view count.

Key Findings

  • Cross-model generalization usually improves, but does not universally dominate real-image baselines. In original Table 3, MoCo-v3-distilled SFM reaches 66.3 on CLIP versus 75.5 for Centroids; CLIP-distilled SFM reaches 68.7 on MoCo-v3 versus 74.7 for Centroids.
  • Increasing IPC helps. In the DINO-v2 → EVA-02 setting of original Table 5, SFM reaches 86.2, 86.5, 87.5, and 88.4 at IPC=1, 2, 5, and 10; LGM reaches 82.4, 84.9, 85.5, and 85.5. This trend is specific to the reported model pair.
  • CI gains depend on the evaluation backbone. With DINO-v2 distillation in original Table 6, SFM → SFM+CI changes from 78.6 → 82.1 on CLIP, 90.6 → 95.1 on DINO-v2, 86.2 → 88.8 on EVA-02, and 80.5 → 80.6 on MoCo-v3. Improvements carrying additional classifier information should not be attributed to flow matching alone.
  • Appendix Table 12 shows that model scale does not produce monotonic gains: with DINO-v2/ViT-B distillation and ViT-S/B/L evaluation, SFM scores 84.9/90.6/90.4 and SFM+CI scores 86.0/95.1/92.4. Appendix ArtBench style-transfer examples are qualitative only, not quantitative evidence of advantages in detection, segmentation, or style transfer.

Highlights & Insights

  • Frozen models invite reassessment of gradient objectives. The simplification comes from identifying class geometry under a small random head, not engineering a faster general-purpose bi-level optimizer. For strong-prior/lightweight-head tasks, it is useful to ask what unique information inner differentiation still carries.
  • Stable real-data supervision matters more than simply adding augmentations. Fixed full-data statistics make low-view-budget optimization viable. Augmentation remains useful, but no longer primarily compensates for batch fluctuations in the real-data target.
  • The boundary of the compressed artifact must remain explicit. Synthetic data plus labels and synthetic data plus a classifier trained on complete real data are different information packages. CI's practical value is boundary reuse, not proof that images alone encode all task knowledge.

Limitations & Future Work

  • Small variance, identical initialization laws, and balanced batches limit the gradient interpretation. Mean flows do not fully model within-class diversity, and cosine matching cannot guarantee a global optimum for nonconvex image optimization.
  • Appendix B explicitly uses test accuracy for early stopping, creating model-selection leakage concerns. Train, validation, and test sets should be separated, with consistent evaluation and augmentation budgets used to verify whether gains persist.
  • Multiple augmentations or simultaneous optimization of 1000 classes still require substantial GPU memory. Class blocking and inner/outer-loop processing are proposed future work, not an implemented constant-memory algorithm.
  • Frozen backbones and their original pretraining costs are external priors; CI additionally requires a complete-real-data classifier and access to the distillation backbone. Deployment accounting should include statistical preprocessing, required models, and extra artifacts, rather than only synthetic image counts.
  • Section 4.3 gives 86.6 for DINO-v2 evaluation after MoCo-v3 distillation, but the SFM row in Table 3 gives 86.7±0.2; 86.6±0.1 belongs to LGM*. This note distinguishes the table entries and preserves the conflict rather than silently reconciling them. Some table references are also displaced: the few-shot paragraph cites Table 6, while its results are in Table 5.
  • Synthetic images differing from originals do not automatically guarantee privacy. Appendix C's “privacy-friendly” application proposal lacks differential-privacy guarantees or membership-inference evaluation; detection and segmentation extensions also remain unverified.
  • vs LGM: Both exploit frozen pretrained features. LGM matches linear-head gradients on local real-data batches, whereas SFM directly matches fixed global relative means. The central contribution is interpreting and replacing supervision, not compressing the backbone or training a generative flow model.
  • vs distribution matching and TCDD: Absolute class centers already provide a strong distillation target, while SFM introduces relations to other classes. Small-dataset ablations motivate studying when relative terms are needed rather than assuming class coupling always helps.
  • vs MGD3 and soft labels: MGD3 synthesizes data using diffusion priors; this paper optimizes image representations without diffusion sampling. CI regresses features and inherits a classifier, differing from temperature- and loss-weight-sensitive soft labels, but retains extra information learned from complete real data.
  • Research direction: Test a validation-selected mixture of absolute centers and relative flows, and add multiple centers for multimodal classes. Equal APB, independent validation sets, and aligned delivered-artifact budgets are needed to distinguish benefits from statistical targets versus stronger supervision.

Rating

  • Novelty: 4/5. Interprets specific random-linear-head gradients as relative class statistics and builds fixed supervision around that interpretation.
  • Experimental Thoroughness: 3/5. Covers multiple datasets, backbones, IPC settings, CI, and efficiency, but test-set stopping, budget differences, and numerical conflicts weaken the evidence.
  • Writing Quality: 3/5. The main argument is accessible, but optimality claims are too strong and exchangeability conditions and appendix reasoning need greater rigor.
  • Value: 4/5. Useful when strong vision backbones already exist and classification training signals need lightweight transfer; not established for from-scratch training or arbitrary downstream tasks.