Steering Diffusion Models via Class-Contrastive Influence for Few-Shot Classification¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Medical Imaging / Few-Shot Learning / Diffusion Models
Keywords: Diffusion Models, Data Augmentation, Few-Shot Medical Classification, Class-Contrastive Influence (C2I), Reinforcement Learning Fine-Tuning
TL;DR¶
Addressing the core issue that synthetic data augmentation in few-shot classification prioritizes image realism over downstream utility, this paper proposes Class-Contrastive Influence (C2I) along with theoretical grounding, demonstrating that high-C2I samples are boundary-proximal hard examples, and fine-tunes diffusion models via reinforcement learning using C2I rewards to generate task-informative samples that substantially boost downstream accuracy and noise robustness.
Background & Motivation¶
In data-constrained regimes—particularly medical image classification where expert annotation is extraordinarily costly—augmenting limited training sets with off-the-shelf text-to-image diffusion models has become a prevailing paradigm. However, existing generative augmentation techniques largely focus on enhancing visual realism, intra-class diversity, or adapting models to in-domain distributions through small-scale fine-tuning. This practice rests on an untested implicit premise: that as long as generated samples look visually convincing and diverse, downstream classifiers will naturally benefit. Yet empirical observations reveal a perplexing discrepancy: different subsets of synthetic images with nearly identical visual fidelity yield wildly divergent downstream classification performance, where some subsets produce clear gains while others contribute virtually nothing or even degrade classification accuracy.
The fundamental tension behind this failure is the lack of a principled, task-oriented metric for sample usefulness. Directly borrowing first-order gradient influence functions from large language model instruction tuning also fails. In classification tasks, the validation set comprises conflicting gradient signals from distinct classes; a synthetic sample that reduces validation loss for one class often penalizes another. Unstratified global gradient dot products wash out these conflicting dynamics, leaving raw influence scores completely uncorrelated with classification efficacy. Consequently, generative models have no guidance on which regions of the data manifold provide genuine utility, blindly producing prototypical or safe centroid samples that do not help refine the decision boundary.
This paper tackles the problem from the perspective of classifier gradient dynamics, discovering that a synthetic sample's true value lies in establishing a sharp contrastive gap between validation gradients of its own class and opposing classes. The core idea is to introduce Class-Contrastive Influence (C2I) to quantify a sample's downstream classification utility, prove theoretically that high-C2I samples correspond to boundary-proximal hard examples that converge toward the global dataset mean, and use C2I as an RL reward to steer diffusion models toward generating decision-boundary-refining samples.
Method¶
Overall Architecture¶
The framework establishes a closed-loop system transforming a general pre-trained diffusion model into a task-directed data generator tailored for downstream classification. The end-to-end pipeline proceeds through four sequential stages: first, both the classification backbone (ViT-B/16) and the text-to-image generator (Stable Diffusion 2.1) are adapted to the target few-shot training split using parameter-efficient LoRA fine-tuning; second, with the classifier checkpoint frozen, low-dimensional random projected gradients are precomputed for validation samples of each class; third, during RL fine-tuning, the diffusion model generates batches of class-conditioned images, the classifier computes projected gradients on-the-fly to evaluate alignment against same-class versus other-class validation gradients, producing a scalar C2I reward; finally, Denoising Diffusion Policy Optimization (DDPO) updates the generator's LoRA parameters, enabling it to synthesize high-utility hard examples that train downstream classifiers across multiple unseen architectures.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Few-Shot Medical Dataset<br/>(16–32 samples per class)"] --> B["Stage 1: Preliminary Adaptation<br/>LoRA tuning of ViT classifier & SD generator"]
B --> C["Stage 2: Class-Contrastive Influence (C2I)<br/>Positive same-class alignment & negative other-class repulsion"]
C --> D["Stage 3: Decision Boundary Hard Example Mechanism<br/>Convex combination converging to global dataset mean"]
D --> E["Stage 4: Reinforcement Learning Diffusion Tuning<br/>DDPO policy optimization driven by C2I reward"]
E --> F["Targeted Synthetic Hard Samples<br/>(500 high-utility samples per class)"]
F --> G["Downstream Classifier Evaluation<br/>ViT-B/16 & ResNet-18 accuracy & noise robustness gains"]
Key Designs¶
1. Class-Contrastive Influence (C2I): Disentangling Same-Class Affinity from Opposite-Class Repulsion
Standard first-order influence functions estimate the utility of a training sample by approximating its reduction on validation loss through gradient inner products \(\langle \nabla \ell(x), \nabla \ell(v) \rangle\). In classification architectures, loss gradients with respect to classifier weights and representations exhibit opposite signs across class labels: gradients of same-class pairs naturally align positively, whereas opposite-class pairs oppose each other with negative inner products. Averaging influence naively over a multi-class validation split causes these opposing signals to cancel out, completely decoupling raw influence from downstream performance. To resolve this, C2I formalizes influence on a per-class basis. For a generated batch of class-conditioned samples \(\mathbf{x} \subset \mathcal{D}_c\) and validation set \(\mathcal{V}_c\) from class \(c \in \{0, 1\}\), the method computes the class-specific mean \(\mu_c(\mathbf{x}, \mathcal{V}_c)\) and variance \(\sigma_c^2(\mathbf{x}, \mathcal{V}_c)\) of projected gradient cosine similarities:
where \(\Pi\) represents a random projection matrix compressing high-dimensional LoRA gradients, and \(\tilde{\Gamma}\) denotes the Adam optimizer gradient vector. In binary classification, the C2I score penalizes intra-class variance while maximizing mean separation:
For \(K\)-class classification, C2I normalizes class-conditional mean influences via a Softmax distribution \(\text{C2I}(x_a) = \frac{\exp(\mu_{aa})}{\sum_{b=1}^K \exp(\mu_{ab})}\), which smoothly reduces to a Sigmoid of the influence difference \(\sigma(\mu_{aa} - \mu_{ab})\) for \(K=2\), rewarding synthetic samples that widen the class-contrastive separation gap.
2. Decision Boundary Proximity and Hard Example Dynamics: The Theoretical Grounding of C2I
Why does widening the C2I separation gap consistently translate into superior downstream generalization? The authors establish a formal theoretical foundation (Theorem 1 and Lemma 2). In logistic regression, the sample \(x^\star\) that maximizes the absolute gradient alignment gap \(|\mu_0(x) - \mu_1(x)|\) over balanced validation splits is analytically characterized by a convex combination of validation features:
When data vectors have comparable norms, the weights become approximately uniform (\(\alpha_i \approx 1/N\)), meaning \(x^\star \approx \bar{v}\) converges directly to the global dataset mean across all classes. Under Lemma 2, for any classifier with reasonable class predictions, the predicted probability at the global mixture mean \(\bar{v}\) exhibits strictly lower confidence and strictly higher cross-entropy loss than at either individual class centroid. Consequently, maximizing C2I steers the generator away from safe, easily classifiable prototype centroids toward hard, high-uncertainty examples proximal to the decision boundary. Empirical analysis on ViT features confirms this theoretical mechanism: as C2I reward increases during fine-tuning, synthetic image representations migrate toward validation centroids and inter-class feature distance contracts, populating the critical decision boundary margin.
3. Reinforcement Learning Fine-Tuning: Steering Diffusion via C2I Policy Optimization
To guide the generative trajectory toward high-C2I manifold regions, the approach formulates diffusion fine-tuning under the Denoising Diffusion Policy Optimization (DDPO) framework. The multi-step sampling chain is treated as a Markov decision process, where the scalar C2I score evaluated on the final synthetic images acts as the reward \(r(\mathbf{x}, \mathcal{V}; \phi_a)\). The frozen ViT-B/16 classifier provides online gradient feedback, while LoRA parameters on the diffusion model's attention projections are updated via importance-sampled policy gradients:
To maintain optimization stability, set-level mini-batches are formed where all samples in a generated set share a common group C2I reward. Because influence evaluation relies solely on a small validation split with low-dimensional projections, the RL fine-tuning process is computationally lightweight (completing in ~5 hours on a single A100 GPU). Crucially, this optimization incurs zero test-time overhead during generation, retaining standard diffusion sampling efficiency.
Loss & Training¶
- Classifier Adaptation & Checkpoint Freezing: ViT-B/16 initialized with ImageNet pre-training is fine-tuned on the few-shot set \(\mathcal{D}\) using LoRA (rank 16, learning rate \(5 \times 10^{-4}\)) for 20 epochs. Its weights \(\phi_a\) are frozen to serve as the gradient reference model.
- Diffusion RL Fine-Tuning: Pre-trained Stable Diffusion 2.1 is adapted with LoRA (rank 16, \(\alpha=16\)) under DDPO for 30 epochs with learning rate \(3 \times 10^{-4}\), guidance scale 0.2, and clip range \(1 \times 10^{-4}\). The checkpoint with the highest validation C2I reward is selected for downstream augmentation.
- Downstream Classifier Training: 500 synthetic images per class are generated to augment the few-shot training pool. Classifiers (ViT-B/16 with LoRA and ResNet-18 trained from scratch) are optimized for 100 epochs, with the best model selected based on validation AUC.
Key Experimental Results¶
Main Results¶
On MedMNIST benchmarks (BreastMNIST 32-shot, DermaMNIST-binary 16-shot, PneumoniaMNIST 16-shot), the method is compared against standard transformations (RandAugment, RandomErasing, Mixup) and diffusion-based baselines (DataDream, Dataset Expansion, DistDiff). Downstream test AUC results are summarized below:
| Backbone | Augmentation Method | BreastMNIST | DermaMNIST-binary | PneumoniaMNIST | Average AUC |
|---|---|---|---|---|---|
| ViT-B/16 | Original only | 0.828 | 0.846 | 0.941 | 0.873 |
| ViT-B/16 | + RandAugment | 0.858 | 0.824 | 0.954 | 0.879 |
| ViT-B/16 | + RandomErasing | 0.873 | 0.839 | 0.945 | 0.885 |
| ViT-B/16 | + Mixup | 0.823 | 0.845 | 0.890 | 0.867 |
| ViT-B/16 | + DataDream | 0.822 | 0.819 | 0.958 | 0.866 |
| ViT-B/16 | + Dataset Expansion | 0.844 | 0.852 | 0.943 | 0.880 |
| ViT-B/16 | + DistDiff | 0.764 | 0.805 | 0.938 | 0.784 |
| ViT-B/16 | + Ours | 0.885 | 0.853 | 0.945 | 0.894 |
| ResNet-18 | Original only | 0.815 | 0.777 | 0.935 | 0.842 |
| ResNet-18 | + RandAugment | 0.764 | 0.787 | 0.936 | 0.829 |
| ResNet-18 | + RandomErasing | 0.758 | 0.747 | 0.900 | 0.802 |
| ResNet-18 | + DataDream | 0.844 | 0.804 | 0.947 | 0.865 |
| ResNet-18 | + Dataset Expansion | 0.804 | 0.831 | 0.956 | 0.864 |
| ResNet-18 | + Ours | 0.854 | 0.836 | 0.956 | 0.882 |
Ablation Study¶
1. Decision Boundary Robustness Under Input Perturbations (Test AUC) To verify whether boundary-proximal samples genuinely stabilize the decision boundary, classifiers are evaluated under test perturbations including Salt & Pepper noise (amount 0.01), JPEG compression (quality 25%), and Gaussian blur (radius 2):
| Dataset | Noise Type | Original Only | Dataset Exp. | DataDream | Ours |
|---|---|---|---|---|---|
| DermaMNIST-binary | Salt & Pepper | 0.766 | 0.813 | 0.810 | 0.830 |
| DermaMNIST-binary | JPEG Compression | 0.806 | 0.800 | 0.821 | 0.831 |
| DermaMNIST-binary | Gaussian Blur | 0.827 | 0.848 | 0.828 | 0.841 |
| DermaMNIST-binary | Average | 0.800 | 0.820 | 0.820 | 0.834 |
| BreastMNIST | Salt & Pepper | 0.764 | 0.772 | 0.817 | 0.832 |
| BreastMNIST | JPEG Compression | 0.760 | 0.804 | 0.814 | 0.816 |
| BreastMNIST | Gaussian Blur | 0.758 | 0.727 | 0.765 | 0.810 |
| BreastMNIST | Average | 0.761 | 0.768 | 0.799 | 0.819 |
| PneumoniaMNIST | Salt & Pepper | 0.868 | 0.861 | 0.823 | 0.792 |
| PneumoniaMNIST | JPEG Compression | 0.922 | 0.907 | 0.950 | 0.940 |
| PneumoniaMNIST | Gaussian Blur | 0.930 | 0.881 | 0.956 | 0.946 |
| PneumoniaMNIST | Average | 0.907 | 0.883 | 0.910 | 0.893 |
2. Comparison Against Naive Misclassified-Sample Hard-Example Fine-Tuning To test whether generating hard examples can be achieved by simply fine-tuning Stable Diffusion on misclassified validation samples, the authors evaluate this naive heuristic on BreastMNIST:
| Dataset | Method / Configuration | Test AUC | Note |
|---|---|---|---|
| BreastMNIST | Original Only | 0.828 | Baseline without augmentation |
| BreastMNIST | Naive Hard Baseline (fine-tuning SD on 15 misclassified images) | 0.778 | Severe drop (-0.050) due to noise drift and overfitting |
| BreastMNIST | C2I RL Fine-Tuning (Ours) | 0.885 | Principled boundary anchoring (+0.057 gain over baseline) |
Key Findings¶
- Cross-Architecture Generalizability: Even though C2I rewards are computed exclusively using a ViT-B/16 model during RL, the generated samples yield an average AUC of 0.882 on ResNet-18 trained from scratch (+4.0% over original data). This demonstrates that C2I captures intrinsic data manifold properties that transfer seamlessly across model families.
- Scalability with Synthetic Data Volume: While existing diffusion methods like DataDream and Dataset Expansion plateau or experience performance drops as synthetic sample counts increase (due to off-manifold drift), C2I-guided augmentation exhibits monotonic performance gains as synthetic samples scale from 50 to 500 per class.
- Robustness to Validation Set Size: On PneumoniaMNIST, reducing the validation set size from 135 samples per class down to a scarce 16 samples per class (matching the few-shot training size) results in virtually identical performance (0.943 vs. 0.944 average AUC), highlighting its viability in extremely low-data clinical settings.
Highlights & Insights¶
- Shifting Augmentation from Realism to Task Utility: The work demonstrates that visual fidelity and statistical diversity are insufficient proxies for downstream classification value; aligning generator updates directly with classifier gradient influence provides a principled metric of sample utility.
- Rigorous Theoretical Connection to Decision Boundaries: Proving that the maximum C2I point is the feature-weighted dataset mean elegantly clarifies why contrastive gradient alignment forces the generator to produce hard, boundary-refining examples rather than trivial prototypes.
- Zero Inference Overhead: Unlike test-time optimization methods (such as Dataset Expansion, which requires ~25 seconds of optimization per image), C2I steers the diffusion weights offline, generating images at standard inference speeds.
Limitations & Future Work¶
- Training Budget Requirements: RL fine-tuning of Stable Diffusion requires approximately 5 GPU hours on an NVIDIA A100 per dataset, introducing upfront computational overhead compared to direct training on few-shot data.
- Sensitivity to Initial Classifier Competence: When label scarcity is extreme and initial classification accuracy is too low (e.g., BreastMNIST at 16-shot achieving only 66% accuracy), gradient feedback becomes excessively noisy, necessitating an expansion to 32-shot to stabilize RL rewards.
- Expansion to Dense and 3D Tasks: The current evaluation focuses on 2D medical classification benchmarks. Extending C2I guidance to 3D volumetric imaging (e.g., CT/MRI) and dense prediction tasks (segmentation and detection) represents a promising future avenue.
Related Work & Insights¶
- vs. DataDream & Dataset Expansion: DataDream and Dataset Expansion optimize for distribution alignment or test-time image perturbation without modeling downstream classifier gradients, often generating redundant centroid samples. C2I directly embeds classification utility into the generative reward, producing boundary-proximal samples that strengthen decision margins.
- vs. LESS & Data Selection Influences: Gradient influence frameworks like LESS operate as post-hoc data filtering or reweighting tools on existing datasets. This work shows that raw influence fails in classification due to cross-class sign cancellation, introduces class-contrastive formulation, and employs it constructively as an RL reward to generate novel high-value data from scratch.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneers Class-Contrastive Influence to quantify task usefulness and uses it as an RL reward to steer diffusion models toward hard boundary samples]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Comprehensive evaluation across MedMNIST benchmarks, multi-class extensions, cross-architecture transfer, noise robustness, and sample-size scaling]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical derivations, intuitive geometric motivation, and thorough empirical validation]
- Value: ⭐⭐⭐⭐⭐ [Provides a paradigm shift for generative data augmentation in low-resource and clinical imaging domains]