Self-Improving Diffusion Classifiers with Minority Preference Optimization¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: Conference PDF
Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5692.txt
Area: Model Compression
Keywords: Diffusion Classifier, Minority Sampling, Reinforcement Learning, Group Relative Policy Optimization, Parameter-Efficient Fine-Tuning
TL;DR¶
MiPO introduces a self-improving framework that broadens diffusion classifiers' perceptual coverage over low-density manifold regions using only arbitrary text prompts and Tweedie-based minority rewards, optimized via selective Group Relative Policy Optimization and LoRA without requiring downstream images or external reward models.
Background & Motivation¶
In recent years, diffusion models have emerged not only as dominant engines for photorealistic visual generation but also as compelling foundational representations for perceptual understanding tasks. Among these paradigms, the Diffusion Classifier (DC) reformulates zero-shot visual recognition by leveraging class-conditional denoising error as an effective surrogate for log-likelihood estimation. This generative classification pipeline boasts distinctive advantages over conventional discriminative models, notably superior out-of-distribution (OOD) robustness, intrinsic resilience against shortcut features, and a nuanced understanding of semantic compositionality. Nevertheless, the perceptual capability of standard diffusion classifiers is fundamentally handcuffed by the density bias of their pretraining data distribution: optimized under mean-squared noise-matching objectives, diffusion models naturally prioritize high-density majority modes while leaving low-density long-tail minority concepts poorly represented and plagued with severe reconstruction uncertainty.
To tackle this minority sampling failure, prior investigations have predominantly focused on inference-time guidance, employing auxiliary classifier gradients or iterative per-prompt token optimizations (such as MinorityPrompt). While these techniques enhance the diversity of generated outputs, they suffer from two debilitating bottlenecks. First, performing continuous test-time backpropagation per prompt incurs massive computational overhead (often exceeding one minute per image), rendering them impractical for scalable inference. Second, and more critically, inference-time guidance leaves the underlying generative backbone entirely frozen; the internal representation remains unmodified, preventing any generation-side diversity gains from translating into improved perception capabilities for diffusion classifiers. Inspired by Richard Feynman's celebrated dictum, "What I cannot create, I do not understand," the perceptual blindness of diffusion classifiers on long-tail concepts is directly rooted in their generative omission of minority manifold regions.
This work breaks the disconnect between minority generation and zero-shot perception by enabling diffusion models to self-improve their long-tail manifold coverage autonomously without downstream images or human-labeled supervision. The core idea is to convert Tweedie-based DDIM reconstruction discrepancies into an intrinsic minority preference reward, optimizing early denoising steps via Group Relative Policy Optimization (GRPO) with LoRA and KL regularization to broaden low-density generative coverage and self-improve zero-shot diffusion classification.
Method¶
Overall Architecture¶
MiPO formulates minority-aware representation enhancement as a self-improving reinforcement learning framework that operates solely on arbitrary caption collections (e.g., HPSv2 prompts), requiring neither downstream target images nor external foundation reward models (e.g., CLIPScore or ImageReward). For each text prompt, the model deploys stochastic differential equation (SDE) sampling to generate a cohort of parallel candidate trajectories conditioned on identical initial Gaussian noise. Each generated sample is evaluated via an intrinsic minority preference reward derived from Tweedie-approximated single-step DDIM reconstruction error. Advantage estimates are subsequently normalized across prompt groups via Group Relative Policy Optimization (GRPO), balanced by a Kullback–Leibler (KL) regularization term against the frozen pretraining policy to prevent manifold divergence. Finally, policy updates are strictly confined to the initial 40% of diffusion timesteps and parameterized by a lightweight LoRA adapter.
The end-to-end operational pipeline and modular data flow are illustrated below:
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Arbitrary Prompt Set<br/>Unlabeled text captions"] --> B["Multi-trajectory SDE Sampling<br/>Candidate generation from shared noise"]
B --> C["DDIM Inverse Minority Preference Reward<br/>Tweedie error measuring manifold sparsity"]
C --> D["Group Relative Policy Optimization<br/>Group-wise advantage normalization & LoRA updates"]
D --> E["Prior-Preserving KL Regularization<br/>Anchor pretrained prior to avert reward hacking"]
E --> F["Selective Policy Optimization<br/>Restricted updates on early 40% coarse timesteps"]
F --> G["Self-Improving Diffusion Classifier<br/>Minority-enhanced zero-shot recognition"]
Key Designs¶
1. DDIM Inverse Minority Preference Reward: Intrinsic Low-Density Manifold Metric Conventional reinforcement learning from human feedback (RLHF) in vision models relies heavily on massive external multimodal scoring networks, which are incapable of directly probing the density landscape of the generative manifold. MiPO circumvents this limitation by designing a fully self-supervised minority reward grounded in Tweedie's formula and single-step DDIM inversion. Given a candidate image \(\mathbf{x}_0\) synthesized from prompt \(c\), Gaussian noise \(\bm{\epsilon} \sim \mathcal{N}(0, \mathbf{I})\) is injected at an intermediate diffusion step \(t = 0.9T\), producing a perturbed latent state \(\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\bm{\epsilon}\). The un-tuned pretrained diffusion denoiser then reconstructs the clean estimate \(\hat{\mathbf{x}}_0\) via the single-step DDIM equation: $\(\hat{\mathbf{x}}_0 = \frac{\mathbf{x}_t - \sqrt{1 - \bar{\alpha}_t}\bm{\epsilon}_{\text{pre}}(\mathbf{x}_t, c, t)}{\sqrt{\bar{\alpha}_t}}\)$ The minority preference reward is directly quantified as the Euclidean distance between the generated sample and its posterior mean approximation, \(\mathcal{M}(\mathbf{x}_0) = \|\hat{\mathbf{x}}_0 - \mathbf{x}_0\|_2\). When a sample lands in a well-represented majority mode, the denoiser accurately cancels the perturbation, yielding negligible reconstruction error. Conversely, when a sample populates an underrepresented minority mode, high epistemic uncertainty produces substantial reconstruction residuals, naturally translating to a high reward. This formulation operates entirely endogenously, detecting low-density manifold boundaries without external annotations.
2. Group Relative Policy Optimization: Image-Free Self-Improving Policy Gradients To eliminate the requirement for an explicit critic network while dramatically suppressing variance in policy gradient estimation, MiPO adapts Group Relative Policy Optimization (GRPO) to lightweight LoRA adapter parameters \(\theta\). For each prompt \(c\), a group of \(G\) stochastic SDE trajectories produce outputs evaluated by the minority reward \(\{r^{(g)}\}_{g=1}^G\). The policy computes group-normalized scalar advantages: $\(A^{(g)} = \frac{r^{(g)} - \text{mean}\big(\{r^{(j)}\}_{j=1}^G\big)}{\text{std}\big(\{r^{(j)}\}_{j=1}^G\big)}\)$ This group normalization ensures invariance to arbitrary reward scaling and prompt heterogeneity: candidate trajectories exhibiting stronger minority features receive positive advantages, while dominant majority trajectories are down-weighted. Parameter updates follow a clipped PPO-style surrogate objective over importance sampling ratios \(\rho_{t,i}(\theta) = \frac{\pi_\theta(a_{t,i}|x_{t,i})}{\pi_{\theta_{\text{old}}}(a_{t,i}|x_{t,i})}\), allowing LoRA weights to gently expand generative coverage into low-density regions without compromising the underlying backbone.
3. Prior-Preserving KL Regularization: Mitigating Reward Hacking and Manifold Collapse Aggressively pursuing minority generation without constraints introduces severe stability risks; reinforcement learning policies are notoriously prone to "reward hacking," synthesizing pathological high-frequency artifacts or degenerated patterns simply to maximize DDIM reconstruction errors. To preserve generative fidelity, MiPO balances the minority objective \(\mathcal{J}_{\text{Minority}}(\theta)\) with an explicit Kullback–Leibler (KL) divergence penalty against the frozen pretraining policy \(\pi_{\text{pre}}\): $\(\mathcal{J}(\theta) = \alpha \mathcal{J}_{\text{Minority}}(\theta) - \beta D_{\text{KL}}(\pi_\theta \| \pi_{\text{pre}})\)$ Configuring \(\alpha = 0.7\) and \(\beta = 0.15\) establishes an elastic constraint anchored to the base generative manifold. This mechanism prevents the adapter from straying into degenerate distributional modes, ensuring that while the model broadens its reach into underrepresented categories, the perceptual realism and high-density majority fidelity remain solidly preserved.
4. Selective Policy Optimization: Focusing on Coarse Semantic Formation Timesteps Backpropagating policy gradients across all \(T\) timesteps of a diffusion trajectory is computationally prohibitive and empirically suboptimal. Recent findings in diffusion dynamics indicate that early denoising steps dictate coarse semantic layout and high-level categorical structure, whereas later steps govern fine-grained pixel smoothing and local texture synthesis. MiPO restricts policy updates strictly to the initial 40% of diffusion steps (e.g., steps \(0 \le t \le 20\) in a 50-step schedule). By confining LoRA fine-tuning to this foundational semantic phase, the optimization avoids overfitting to local high-frequency pixel noise, substantially reduces training memory, and focuses capacity on rectifying semantic-level classification blind spots.
Loss & Training¶
Training uses 100k prompts from HPSv2 without any paired training images. Low-rank LoRA adapters are injected into the cross-attention and self-attention projections of the Stable Diffusion (SD 1.5 and SD 2.0) UNet. Optimization is conducted using AdamW with \(G = 4\) trajectories per prompt group under stochastic SDE sampling. The policy objective incorporates clipping parameter \(\epsilon = 0.2\) and KL penalty weight \(\beta = 0.15\), applied exclusively over \(t \in [0, 0.4T]\). For zero-shot diffusion classification, test images are perturbed across time steps, conditioned on respective candidate class textual labels, and classified according to the label minimizing the L2 error of the noise estimation output.
Key Experimental Results¶
Main Results¶
Zero-shot classification performance across standard diffusion classifier benchmarks and out-of-distribution (OOD) testbeds is detailed in the table below. Across both SD 1.5 and SD 2.0 backbones, MiPO delivers consistent improvements without observing a single downstream training image.
| Backbone & Method | CIFAR-10 | CIFAR-10-C (Robustness) | ImageNet-Tiny | Caltech-101 | SUN09 |
|---|---|---|---|---|---|
| SD 1.5 (Baseline) | 85.38% | 57.72% | 46.50% | 97.81% | 64.69% |
| SD 1.5 + MiPO (Ours) | 87.89% | 59.76% | 47.30% | 99.51% | 68.25% |
| SD 1.5 Gain | +2.51% | +2.04% | +0.80% | +1.70% | +3.56% |
| SD 2.0 (Baseline) | 88.58% | 64.61% | 51.45% | 99.65% | 64.72% |
| SD 2.0 + MiPO (Ours) | 86.98% | 65.33% | 55.25% | 99.93% | 69.26% |
| SD 2.0 Gain | -1.60% | +0.72% | +3.80% | +0.28% | +4.54% |
Ablation Study¶
To ensure rigorous evaluation independent of the training reward metric, CIFAR-10 was partitioned into distinct Minority and Majority groups using an external k-Nearest Neighbors (KNN) density estimator.
Table 1: Effect of KL Regularization on Classification Accuracy (SD 1.5)
| Configuration | Minority Group Acc. | Majority Group Acc. | Mechanism Analysis |
|---|---|---|---|
| SD 1.5 (Vanilla Baseline) | 73.07% | 79.18% | Significant performance gap on underrepresented modes |
| MiPO w/o KL Penalty (\(\beta = 0\)) | 78.58% | 82.08% | Reward hacking degrades distribution stability |
| MiPO Full Model (w/ KL, Ours) | 80.58% | 82.68% | Prevents drift: +7.51% on minority, +3.50% on majority |
Table 2: Ablation on Denoising Timestep Selection Strategies (SD 1.5)
| Selection Strategy | Timestep Interval | Minority Group Acc. | Majority Group Acc. | Impact Analysis |
|---|---|---|---|---|
| SD 1.5 (Baseline) | Unmodified | 73.07% | 79.18% | Pretrained baseline reference |
| Full Trajectory (Full) | 0–50 steps (100%) | 80.58% | 82.68% | Strong performance but highest memory & compute cost |
| Late Phase (Late) | 30–50 steps (last 40%) | 73.77% | 78.18% | Adjusts high-frequency pixels without altering semantics |
| Middle Phase (Middle) | 15–35 steps (mid 40%) | 79.90% | 78.28% | Incomplete semantic shift; slight majority drop |
| Early Phase (Early, Ours) | 0–20 steps (first 40%) | 81.68% | 82.48% | Optimal structural adaptation; highest minority accuracy |
Table 3: Minority Image Generation and Inference Efficiency (SD 1.5, MS-COCO Val)
| Method | Minority Gen. | DC Compatible | Prompt-Adaptive | CLIPScore ↑ | HPSv3 ↑ | Minority Score ↑ | Time/Image ↓ |
|---|---|---|---|---|---|---|---|
| Vanilla DDIM | No | Yes | Yes | 31.7571 | 5.7581 | 0.2132 | 1s |
| MinorityPrompt | Yes | No | No | 30.4351 | 1.5324 | 0.2311 | 63s |
| MiPO (Ours) | Yes | Yes | Yes | 30.8026 | 4.4090 | 0.2289 | 1s |
Key Findings¶
- Generative Manifold Expansion Directly Empowers Perception: Targeted evaluation on low-density CIFAR-10 partitions reveals that MiPO propels minority classification accuracy from 73.07% to 81.68% (an +8.61% absolute gain), empirically corroborating that generative coverage over low-density modes is a prerequisite for robust zero-shot classification.
- Early Timesteps Govern Semantic Perceptual Alignment: Constraining policy updates to the first 40% of denoising steps achieves superior minority recognition (81.68%) compared to updating the entire trajectory (80.58%), confirming that later diffusion stages risk overfitting to local pixel artifacts rather than enhancing conceptual understanding.
- Dramatically Outperforms Inference-Time Guidance in Speed and Quality: While prior state-of-the-art MinorityPrompt requires 63 seconds of gradient backpropagation per prompt and degrades human preference scores (ImageReward -0.026), MiPO amortizes minority exploration into static LoRA weights, executing at an instantaneous 1-second inference latency while delivering top human preference scores (ImageReward 0.228).
Highlights & Insights¶
- Self-Improving Generative-Perceptual Synergy: Establishes a closed-loop self-training paradigm wherein a diffusion model probes its own epistemic uncertainty to generate synthetic long-tail rewards, bridging the long-standing divide between generative diversity and discriminative recognition.
- Inference-Free Plug-and-Play Architecture: Encapsulates minority preference adaptation within modular LoRA weights, completely discarding the cumbersome test-time gradient optimization of prior works and enabling zero-latency switching across classification and synthesis tasks.
- Extensible Density-Aware Policy Alignment: Utilizing inverse Tweedie residuals as intrinsic manifold density estimators presents an adaptable template readily transferable to Flow Matching architectures, long-tail synthesis, and self-supervised anomaly detection.
Limitations & Future Work¶
- Performance Degradation on Cluttered Multi-Object Scenes: In uncurated datasets featuring heavy occlusion and multiple co-occurring instances (e.g., LabelME and VOC2007), MiPO experiences noticeable accuracy declines (e.g., SD 1.5 dropping from 62.90% to 55.28% on LabelME). Global minority rewards can inadvertently incentivize background clutter or irrelevant rare context rather than the primary target object.
- Fidelity-Diversity Equilibrium: Hyper-optimizing minority rewards introduces a delicate trade-off between conceptual distinctiveness and perceptual aesthetics. Dynamically scheduled temperature mechanisms are required to calibrate exploration against visual fidelity.
- Future Directions: Developing spatially-grounded, object-aware minority rewards and investigating the applicability of self-improving policy alignment to video diffusion and 3D representations.
Related Work & Insights¶
- vs. MinorityPrompt / Don't Play Favorites: Existing minority generators rely on slow test-time guidance (~63s/image) that leaves underlying model weights untouched; MiPO fine-tunes model representations offline, simultaneously boosting generation speed by 60x and transferring manifold improvements to zero-shot classification.
- vs. Standard Diffusion Classifiers: Conventional diffusion classifiers inherit pretraining mode collapse towards majority concepts; MiPO is the first self-improving approach to actively correct long-tail perceptual deficiencies without extra images.
- vs. DanceGRPO / FlowGRPO: While contemporary diffusion RL frameworks target subjective human aesthetics or text alignment, MiPO breaks new ground by leveraging intrinsic manifold geometric sparsity as a policy reward.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates the first direct link between minority generative sampling and diffusion classifier perception, resolved via an elegant, reward-model-free RL framework.
- Experimental Thoroughness: ⭐⭐⭐⭐☆ Extensive evaluation across multiple benchmarks, OOD sets, generation metrics, and independent KNN ablations, though multi-object complex scene degradation warrants further resolution.
- Writing Quality: ⭐⭐⭐⭐⭐ Exceptionally clear narrative progression from intuitive motivation to rigorous mathematical formulation and empirical validation.
- Value: ⭐⭐⭐⭐⭐ Delivers critical insights for resolving long-tail representation failures and broadening the utility of generative perception models.