DisRM: Reward Modeling as Discriminative Prediction¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: Image Generation
Keywords: reward modeling / discriminative prediction / preference alignment / diffusion models / sample selection
TL;DR¶
DisRM replaces expensive pairwise human preference annotations and complex multi-metric engineering by training a lightweight discriminator on just 500 unlabeled representative samples to distinguish generator outputs from preferred targets, enabling iterative generator-reward model co-refinement across Best-of-N, SFT, and DPO.
Background & Motivation¶
Post-training alignment techniques—such as Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO)—as well as test-time Best-of-N sampling depend heavily on accurate reward models to steer visual generative models toward human expectations. However, building reliable reward models remains prohibitively expensive. Prevailing methodologies (such as PickScore, ImageReward, and HPSv2) rely on massive pairwise preference datasets, frequently requiring hundreds of thousands to over one million crowdsourced pairwise comparisons. Collecting these annotations at scale is not only economically daunting, but also inflexible when target preferences shift across domains or application goals. An alternative paradigm attempts to bypass pairwise labels by manually defining and calibrating dozens of explicit content-quality metrics (e.g., aesthetics, text alignment, anatomical fidelity), but this dimension-engineering strategy incurs immense development overhead and often fails to reflect nuanced, holistic human perceptions.
The core tension stems from an artificial constraint imposed by existing pipelines: they require explicit, comparative human judgments or exhaustively hand-crafted scoring criteria, even though humans intuitively recognize a small group of "exemplar outputs" representing their target preference with minimal effort. Furthermore, standard alignment frameworks operating on static preference data (such as DiffusionDPO) are intrinsically restricted to single-round optimization; once the generator advances, the fixed dataset fails to supply harder, informative training signals, causing post-training gains to plateau or overfit.
This paper tackles this bottleneck by reformulating the objective of reward modeling: rather than learning to compare arbitrary pairs or scoring multidimensional quality attributes, one can simply train a discriminator to distinguish a compact set of target exemplars from the generator's current outputs. Core idea: reformulate reward modeling as a discriminative prediction task using a small set of unlabeled Preference Proxy Data (PPD), implicitly learning robust preference signals from target versus generated distributions while leveraging generator updates to supply dynamic hard negatives for iterative co-refinement.
Method¶
Overall Architecture¶
The DisRM pipeline establishes an iterative closed loop between a visual generator \(G_t\) and a discriminative reward model \(R_t\). The framework requires only a tiny collection of positive-only Preference Proxy Data \(D_p\) (e.g., 500 unlabeled images from a high-quality source) without comparative labels. In each round, visual features extracted by a frozen foundation encoder pass into a lightweight Reward Projection Layer (RPL) trained to discriminate \(D_p\) from raw generator outputs. A rank-based bootstrapping mechanism subsequently expands the training distribution with pseudo-labeled generator samples. The resulting DisRM model can either directly evaluate and rank candidate samples during inference (Best-of-N selection) or annotate online preference pairs to drive SFT or DPO updates for the generator.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Preference Proxy Data (PPD) & Generator Outputs"] --> B["Preference Proxy Data & Discriminative Reward Modeling<br/>CLIP features + RPL discriminator distinguish target from generated distributions"]
B --> C["Rank-based Bootstrapping<br/>Score confident generator samples with rank-decayed pseudo-labels"]
C --> D["Iterative Co-Refinement & Post-Training<br/>Score candidate samples (Best-of-N) or construct preference pairs for SFT / DPO"]
D -->|Updated generator generates harder negative samples| B
Key Designs¶
1. Preference Proxy Data & Discriminative Reward Modeling: replacing pairwise labels with exemplar discrimination To circumvent the prohibitive costs and subjectivity of pairwise human labeling, DisRM introduces the concept of Preference Proxy Data (PPD). For image quality enhancement, PPD consists of a compact set of \(N=500\) unlabeled exemplar images \(D_p = \{x_i^+\}_{i=1}^N\) randomly drawn from a high-quality synthetic gallery (JourneyDB); for safety alignment, it employs synthetic safe images without comparative ranking. To robustly capture high-level visual representations under extreme data scarcity, DisRM builds upon a pre-trained foundation vision encoder (CLIP-Vision for images or ViCLIP for video) coupled with a lightweight Reward Projection Layer (RPL), implemented as an MLP with normalization. The model optimizes a binary cross-entropy loss:
where \(y=1\) denotes PPD exemplars and \(y=0\) denotes raw generator outputs. The normalized probability of the positive class serves directly as the reward score: \(r(x) = \text{softmax}(\text{RPL}(\text{CLIP-Vision}(x)))_1\). This discriminative formulation implicitly captures multi-faceted aesthetic and structural attributes without manual dimension weighting.
2. Rank-based Bootstrapping: maximizing data efficiency on limited exemplars Because initial PPD contains only a few hundred samples, training a discriminator solely on raw outputs risks inadequate coverage of the generator's vast latent support. DisRM incorporates a Rank-based Bootstrapping strategy to expand its training space. After initial training, DisRM evaluates a larger pool of images generated across diverse prompts. The top-\(M\) highest-scoring samples form a pseudo-positive set \(D_f^+\), while the \(M\) lowest-scoring samples form a pseudo-negative set \(D_f^-\). Rather than assigning binary hard labels to pseudo-positives, DisRM assigns soft labels that decay exponentially according to their score rank \(r\):
where \(\alpha > 0\) regulates the confidence decay rate across ranks. Pseudo-negatives are assigned hard label 0. Retraining DisRM on the augmented dataset \(D = D_r \cup D_f^+ \cup D_f^-\) anchors the discriminator on high-confidence model outputs, significantly boosting generalization across the generator's output distribution.
3. Iterative Co-Refinement & Post-Training: dynamic hard negatives for multi-round alignment Unlike standard offline DPO (such as DiffusionDPO), which is constrained to a static preference set, or RAFT-style iterative loops that risk reward hacking under a static reward model, DisRM establishes an evolving co-refinement loop. In round \(t\), DisRM \(R_t\) scores batches of candidate samples from generator \(G_t\). For each prompt, the highest-scoring candidate \(x_h\) and lowest-scoring candidate \(x_l\) can either supply the optimal sample at inference (Best-of-N), feed into SFT on \(x_h\), or form online preference pairs \((x_h, x_l)\) for DPO. When the generator improves to \(G_{t+1}\), its outputs more closely resemble the static PPD \(D_p\), naturally acting as "harder negatives" for the next discriminator round \(R_{t+1}\). Consequently, both components continually co-evolve across multiple rounds without additional human intervention.
Key Experimental Results¶
Main Results¶
DisRM was thoroughly evaluated across Stable Diffusion 1.5 and SDXL under Best-of-N sample selection (Ours-RM@10), SFT, and DPO. Evaluations measure distributional fidelity (FID), human preference alignment (ImageReward, PickScore, HPS), and text-image alignment (CLIPScore).
| Model | Optimization | Preference Data | Data Volume | FID↓ | ImageReward↑ | PickScore↑ | HPS↑ | CLIPScore↑ |
|---|---|---|---|---|---|---|---|---|
| SD1.5 | Base-model | - | - | 72.06 | -0.040 | 19.460 | 0.277 | 0.698 |
| SD1.5 | SPO | PickScore | 1M pairs | 70.19 | 0.310 | 20.248 | 0.262 | 0.666 |
| SD1.5 | DiffusionDPO | Pickapic | 1M pairs | 68.15 | 0.180 | 19.869 | 0.281 | 0.709 |
| SD1.5 | DiffusionNPO | PickScore | 1M pairs | 71.60 | -0.017 | 19.520 | 0.272 | 0.684 |
| SD1.5 | Ours-RM@10 | DisRM | 0.5k unpaired | 68.51 | 0.072 | 19.650 | 0.282 | 0.703 |
| SD1.5 | Ours-SFT | DisRM | 0.5k unpaired | 64.98 | 0.217 | 19.980 | 0.284 | 0.720 |
| SD1.5 | Ours-DPO | DisRM | 0.5k unpaired | 63.61 | 0.240 | 20.032 | 0.281 | 0.710 |
| SDXL | Base-model | - | - | 62.83 | 0.790 | 21.235 | 0.293 | 0.744 |
| SDXL | DiffusionDPO | Pickapic | 1M pairs | 63.24 | 1.033 | 21.628 | 0.301 | 0.765 |
| SDXL | Ours-RM@10 | DisRM | 0.5k unpaired | 62.05 | 0.890 | 21.311 | 0.297 | 0.753 |
| SDXL | Ours-SFT | DisRM | 0.5k unpaired | 61.74 | 0.915 | 21.275 | 0.297 | 0.756 |
| SDXL | Ours-DPO | DisRM | 0.5k unpaired | 61.95 | 0.893 | 21.305 | 0.296 | 0.753 |
When compared directly as an automated annotation engine for 10,000 DPO pairs on SD1.5, single-metric baselines overfit heavily: ImageReward annotations inflate ImageReward to 0.186 at the cost of degrading FID to 72.59 and CLIPScore to 0.686; PickScore annotations push PickScore to 19.849 while leaving FID at 70.97. In contrast, DisRM-annotated DPO yields the best FID of 67.42 along with balanced gains across all preference metrics (IR 0.131, PS 19.715, CLIPScore 0.702), avoiding reward hacking.
Ablation Study¶
The ablations investigate the impact of different reward model training strategies on sample selection, as well as the progressive benefits of multi-round iterative co-refinement under DPO.
| Configuration | Strategy / Round | FID↓ | ImageReward↑ | PickScore↑ | HPS↑ | Note |
|---|---|---|---|---|---|---|
| Single Checkpoint | Naiive | 14.48 | 0.048 | 19.612 | 0.280 | Single checkpoint after fixed training steps |
| Checkpoint Averaging | Average | 14.56 | 0.067 | 19.624 | 0.280 | Weight averaging across regular intervals |
| Score Voting | Voting | 14.61 | 0.063 | 19.618 | 0.281 | Score voting ensemble across checkpoints |
| Rank Bootstrapping (DisRM) | Bootstrap | 14.18 | 0.071 | 19.651 | 0.282 | Soft-label rank bootstrapping achieves best performance |
| Alignment Baseline | Base model | 72.06 | -0.037 | 19.467 | 0.277 | Initial unaligned SD1.5 generator |
| Round 1 Alignment | Round 1 | 66.20 | 0.099 | 19.631 | 0.279 | FID drops by 5.86 points; immediate gains |
| Round 2 Alignment | Round 2 | 64.98 | 0.223 | 19.960 | 0.281 | Harder negatives shift boundary; IR doubles |
| Round 3 Alignment | Round 3 | 63.61 | 0.240 | 20.032 | 0.282 | Cumulative 8.45 FID reduction across three rounds |
Key Findings¶
- Unprecedented data efficiency: Using only 500 unlabeled PPD images, DisRM-driven DPO achieves an FID of 63.61 on SD1.5, substantially outperforming DiffusionDPO (68.15) and DiffusionNPO (71.60) trained on 1,000,000 pairwise human labels. In blind human evaluations, Ours-DPO achieves a 74.4% win rate over base SD1.5, matching human preferences with 70.5% agreement.
- Prevention of metric overfitting: Unlike pairwise reward models that overfit to their own reward function and sacrifice image fidelity or text alignment, DisRM's discriminative signal provides a balanced density-ratio gradient that improves aesthetic, distributional, and semantic scores simultaneously.
- Cross-task and modal generalization: In safety alignment, using synthetic safe images as PPD drops the Inappropriate Probability (IP) from 42/51 down to 18/17 on SD1.5/SDXL. In video generation with VideoCrafter2, DisRM trained on 500 Artgrid clips reduces FVD from 1021.77 to 983.28 and increases VBench to 81.38, matching heavy multi-expert pipelines without extra manual annotation.
Highlights & Insights¶
- Paradigm shift from pairwise rankings to discriminative exemplars: DisRM proves that distinguishing positive exemplars from current generative outputs is sufficient to recover smooth, human-aligned reward gradients, bypassing millions of pairwise comparisons.
- Rank-based exponential soft labels: The rank-decayed pseudo-label formulation \(y=e^{-\alpha r}\) mitigates confirmation bias in bootstrapping, effectively smoothing the estimated density ratio across the generator's manifold.
- Self-sustaining co-refinement dynamics: The fixed PPD acts as a high-dimensional anchor while the advancing generator continuously generates harder negative samples, unlocking multi-round post-training gains without refreshing human feedback.
Limitations & Future Work¶
- Vulnerability to exemplar distribution bias: Although only 500 samples are required, any stylistic or compositional bias in PPD (e.g., specific artistic rendering styles) could inadvertently narrow generation diversity.
- Backbone representation bottlenecks: DisRM relies on pre-trained vision representations (such as CLIP or ViCLIP); fine-grained attributes not captured well in CLIP feature space (e.g., complex text spelling or precise spatial counts) remain difficult to optimize.
- Stability across extended iterations: The relative balance of updates between discriminator and generator requires careful scheduling across rounds to prevent reward collapse or mode dropping over prolonged training.
Related Work & Insights¶
- vs DiffusionDPO [50] / D3PO [58]: Traditional DPO relies on static offline pairwise datasets limited to single-round training; DisRM acts as an online annotation engine that generates dynamic preference pairs for continuous multi-round optimization.
- vs ImageReward [57] / PickScore [21]: Prior reward models require millions of human pairwise annotations and frequently suffer from metric overfitting; DisRM requires only hundreds of unpaired target samples and exhibits significantly more balanced cross-metric performance.
- vs VideoDPO [28]: VideoDPO relies on dozens of visual expert models and tedious dimension engineering; DisRM trains a compact discriminator over ViCLIP using 500 raw clips, achieving competitive alignment with dramatically lower engineering complexity.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates visual reward modeling as discriminative prediction over minimal positive exemplars with rank bootstrapping.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation spanning SD1.5, SDXL, VideoCrafter2 across aesthetics, safety alignment, and video generation with rigorous user studies.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear theoretical motivation, coherent structure, and self-consistent empirical analysis.
- Value: ⭐⭐⭐⭐⭐ Significantly lowers the data and compute barrier for visual generative alignment, establishing a practical blueprint for iterative post-training.