Reference-Free Quality Assessment for Virtual Try-On via Human Feedback¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/litelightlite/VTON-IQA
Area: Human Understanding
Keywords: Virtual Try-On, Image Quality Assessment, Human Feedback, Cross-Attention, Benchmark
TL;DR¶
To tackle the lack of ground-truth reference images in real-world virtual try-on deployment, this paper constructs VTON-QBench—the largest curated human feedback benchmark with 62,688 try-on images and 431,800 annotations—and introduces VTON-IQA, a reference-free evaluator with Interleaved Cross-Attention that achieves near-human quality prediction and model ranking.
Background & Motivation¶
In fashion e-commerce, image-based virtual try-on (VTON) synthesizes a realistic image of a target person wearing a specified garment, bridging the divide between online shopping and in-store fitting experiences. Despite notable generative advances spanning GANs, latent diffusion models, and diffusion transformers, evaluating try-on synthesis remains fundamentally bottlenecked by evaluation criteria. In real-world industrial settings, ground-truth images of the exact same customer wearing the unpurchased target garment are inherently unavailable. Consequently, traditional full-reference metrics such as SSIM and LPIPS are inapplicable in practical workflows, while distribution-level statistics like FID and KID evaluate global dataset domain similarity rather than the perceptual quality and localized fidelity of individual try-on images. While human subjective user studies remain the gold standard, their exorbitant cost, slow turnaround, and lack of reproducibility make them prohibitive for rapid model iteration.
Existing reference-free attempts in virtual try-on either leverage multimodal LLMs to output qualitative textual critiques or borrow off-the-shelf general aesthetic predictors, lacking rigorous alignment with human perceptual criteria under a large-scale standardized benchmark. Crucially, virtual try-on image quality assessment fundamentally departs from conventional single-image quality assessment (IQA): it is inherently a multi-image relation evaluation task. A valid try-on outcome must simultaneously ensure high-fidelity transfer of fine-grained garment attributes (collar structure, textures, typography) and strict preservation of non-target visual elements (facial identity, body pose, limb structure, and background). Existing reference-based metrics over-penalize natural posture adjustments or camera zoom variations, severely deviating from human perceptual judgments.
This work addresses the challenge by training a task-specific evaluation network directly from curated large-scale human feedback. Core idea: construct VTON-QBench, a rigorously curated human subjective benchmark comprising 62,688 try-on images and 431,800 reliable annotations, and design VTON-IQA, a reference-free framework equipped with Interleaved Cross-Attention (ICA) that models asymmetric garment fidelity and human identity preservation to output human-aligned continuous quality scores without ground-truth images.
Method¶
Overall Architecture¶
VTON-IQA accepts three input images: the flat target garment \(I_G\), the source person \(I_P\), and the synthesized try-on image \(I_V\). The model builds upon a three-branch Vision Transformer backbone (DINOv3 ViT-L/16) for hierarchical feature extraction. The initial \(L/2\) transformer layers perform independent self-attention feature extraction for each branch. The subsequent \(L/2\) layers introduce the proposed Interleaved Cross-Attention (ICA) module, which explicitly facilitates bidirectional cross-modal interaction between the try-on representation and the garment and person branches. Finally, global [CLS] tokens from all three branches are extracted and combined via an adaptive cosine similarity formulation followed by an affine transformation and a hyperbolic tangent function, producing a bounded continuous quality score \(\hat{s} \in [-1, 1]\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Three Input Images<br/>Target Garment IG / Source Person IP / Synthesized Try-On IV"] --> S1["Shallow Feature Extraction<br/>DINOv3 First L/2 Layers Independent Self-Attention"]
S1 --> S2["Interleaved Cross-Attention ICA<br/>Deep L/2 Layers Asymmetric Modeling V↔G and V↔P"]
S2 --> S3["Dual Relational Score Aggregation<br/>Garment Fidelity and Identity Preservation Cosine Weighting"]
S3 --> S4["Joint Alignment Optimization<br/>Pairwise Preference Cross-Entropy + Score Regression"]
S4 --> Out["Output Human-Aligned Continuous Score s ∈ [-1, 1]"]
Key Designs¶
1. VTON-QBench Benchmark Construction and Multi-Stage Curation To provide reliable subjective supervisory signals, the authors construct a large-scale benchmark featuring 62,688 try-on images generated across 14 representative VTON models (spanning GANs, U-Net diffusion, DiT-based models, and proprietary commercial engines). Using FLUX.1-dev and a specialized LoRA, synthetic garment-person pairs covering 24 fashion styles are incorporated to dramatically enhance demographic, pose, and aesthetic diversity. To counter crowdsourcing noise, inattention, and subjective rating bias, a strict two-stage curation pipeline is established: five deterministic sanity-check tasks are embedded in each questionnaire of 50 tasks, filtering out workers who fail checks, choose identical options in \(\ge 80\%\) of tasks, or deviate from majority votes in \(> 60\%\) of cases. At the questionnaire level, Krippendorff's \(\alpha\) is calculated across the three-level ordinal ratings, discarding questionnaires with \(\alpha \le 0.4\). This filtering raises mean inter-rater agreement \(\alpha\) from 0.286 to 0.550, yielding 431,800 high-confidence annotations from 13,838 qualified workers.
2. Interleaved Cross-Attention (ICA) Conventional multi-branch architectures often deploy symmetric all-to-all cross-attention, introducing heavy computational redundancy and irrelevant \(G \leftrightarrow P\) feature entanglement. In contrast, ICA restricts cross-attention layers in the latter \(L/2\) transformer blocks exclusively to the \((V, G)\) and \((V, P)\) modality pairs. In layer \(\ell\), the generated try-on representation integrates updates from both garment and person features: $\(\widehat{X}_V^{(\ell)} = \widetilde{X}_V^{(\ell)} + C_{V\leftarrow G}^{(\ell)} + C_{V\leftarrow P}^{(\ell)}\)$ Meanwhile, the garment and person branches only receive updates from the try-on stream: \(\widehat{X}_G^{(\ell)} = \widetilde{X}_G^{(\ell)} + C_{G\leftarrow V}^{(\ell)}\) and \(\widehat{X}_P^{(\ell)} = \widetilde{X}_P^{(\ell)} + C_{P\leftarrow V}^{(\ell)}\). This asymmetric topology enforces that quality assessment remains centered on the synthesized try-on image, guiding the attention maps to pinpoint local garment fidelity defects (such as distorted neckline patterns or blurred text logos) and identity corruptions (such as distorted hands, facial deformation, or background leakage).
3. Dual Relational Score Aggregation and Bounded Mapping At the final layer, global tokens \(c_G, c_P, c_V \in \mathbb{R}^d\) are extracted from each branch. The model explicitly decomposes try-on quality into two orthogonal perceptual dimensions—garment transfer consistency and non-target region preservation—balanced dynamically via a learnable parameter \(\alpha \in [0, 1]\): $\(\tilde{s} = \alpha \frac{c_G^\top c_V}{\|c_G\|\|c_V\|} + (1-\alpha) \frac{c_P^\top c_V}{\|c_P\|\|c_V\|}\)$ The intermediate score is then mapped into a normalized bounded range via learnable affine scalars \(a, b\) and a tanh activation: $\(\hat{s} = \tanh(a\tilde{s} + b)\)$ This design prevents extreme outlier predictions, improves optimization stability, and offers transparent interpretability by quantifying how human perception trades off garment preservation against subject consistency.
Loss & Training¶
Given that absolute human ratings exhibit inter-subject variance while relative preference between two candidate outputs for the same subject-garment pair remains highly robust, the objective function combines pairwise preference optimization with score regression. Following the Bradley-Terry formulation, for two generated outputs \(I_{V_i}\) and \(I_{V_j}\) given \((I_G, I_P)\), the predicted preference probability \(p_\theta\) and empirical human probability \(q_{ij}\) are defined as: $\(p_\theta(I_{V_i} \succ I_{V_j} \mid I_G, I_P) = \sigma\left(\frac{\Psi_\theta(I_{V_i}) - \Psi_\theta(I_{V_j})}{\tau}\right), \quad q_{ij} = \sigma\left(\frac{S_i - S_j}{\tau}\right)\)$ The overall loss combines soft-label cross-entropy with a score regression term: $\(\mathcal{L}_\theta = - \sum_{(i, j)} \left[ q_{ij} \log p_\theta + (1 - q_{ij}) \log(1 - p_\theta) \right] + \lambda \sum_{k} \|\Psi_\theta(I_{V_k}) - S_k\|^2\)$ The backbone freezes the initial 12 layers while fine-tuning the remaining 12 layers along with the inserted ICA modules on a single NVIDIA A100 GPU using AdamW and bfloat16 mixed precision.
Key Experimental Results¶
Main Results¶
Evaluated on the held-out test split of VTON-QBench (13,038 samples with disjoint garment and person identities), VTON-IQA is compared against reference-based metrics and a zero-shot DINOv3 baseline. Because raw SSIM, LPIPS, and zero-shot DINOv3 scores do not match the human scale linearly, PLCC and \(R^2\) are not applicable for them; evaluation focuses on Spearman's Rank Correlation Coefficient (SRCC) and macro/micro pairwise ranking accuracy (\(A_{\text{macro}}, A_{\text{micro}}\)).
| Scorer | Reference Required | Fine-Tuned | \(\rho_{\text{SRCC}} \uparrow\) | \(\rho_{\text{PLCC}} \uparrow\) | \(R^2 \uparrow\) | \(A_{\text{macro}} \uparrow\) | \(A_{\text{micro}} \uparrow\) |
|---|---|---|---|---|---|---|---|
| SSIM | Yes | No | 0.150 | — | — | 0.596 | 0.593 |
| LPIPS (negated) | Yes | No | 0.406 | — | — | 0.701 | 0.695 |
| DINOv3 (zero-shot) | No | No | 0.244 | — | — | 0.637 | 0.641 |
| VTON-IQA w/o ICA | No | Yes | 0.617 | 0.615 | 0.372 | 0.722 | 0.747 |
| VTON-IQA (Full) | No | Yes | 0.750 | 0.751 | 0.553 | 0.781 | 0.790 |
In a human consistency study where test annotators are randomly split into two groups to establish the human ceiling, human inter-annotator agreement measures \(\rho_{\text{SRCC}} = 0.760 \pm 0.004\), \(\rho_{\text{PLCC}} = 0.762 \pm 0.004\), \(A_{\text{macro}} = 0.782 \pm 0.002\), and \(A_{\text{micro}} = 0.791 \pm 0.002\). VTON-IQA achieves \(A_{\text{macro}} = 0.771 \pm 0.003\) and \(A_{\text{micro}} = 0.783 \pm 0.002\), closely matching the human upper bound on pairwise ranking consistency.
Ablation Study¶
To verify the necessity of Interleaved Cross-Attention (ICA), the model was ablated against a variant without cross-attention layers:
| Config | \(\rho_{\text{SRCC}} \uparrow\) | \(\rho_{\text{PLCC}} \uparrow\) | \(R^2 \uparrow\) | \(A_{\text{micro}} \uparrow\) | Note |
|---|---|---|---|---|---|
| Full VTON-IQA | 0.750 | 0.751 | 0.553 | 0.790 | Full asymmetric relational modeling |
| w/o ICA module | 0.617 | 0.615 | 0.372 | 0.747 | Independent branches; \(R^2\) drops by 32.7% |
| Zero-shot DINOv3 | 0.244 | — | — | 0.641 | Unadapted baseline without task-specific training |
Benchmark evaluation of 14 representative VTON models on the VITON-HD dataset highlights the divergence between VTON-IQA and traditional metrics:
| Model Category | Representative Model | VITON-HD Paired Ours \(\uparrow\) | VITON-HD Unpaired Ours \(\uparrow\) | VITON-HD FID \(\downarrow\) | VITON-HD SSIM \(\uparrow\) |
|---|---|---|---|---|---|
| Proprietary Model | Nano Banana Pro | 0.303 | 0.315 | 3.318 | 0.904 |
| Proprietary Model | GPT-Image-1.5 | 0.255 | 0.234 | 9.613 | 0.768 |
| DiT-based Diffusion | FitDit | 0.230 | 0.189 | 8.048 | 0.843 |
| DiT-based Diffusion | CatVTON-FLUX | 0.207 | -0.025 | 5.470 | 0.872 |
| U-Net-based Diffusion | IDM-VTON | 0.151 | 0.039 | 6.007 | 0.865 |
| U-Net-based Diffusion | CatVTON | -0.163 | -0.266 | 9.082 | 0.833 |
| GAN-based Model | SD-VITON | -0.368 | -0.479 | 7.746 | 0.886 |
| GAN-based Model | VITON-HD | -0.512 | -0.559 | 9.450 | 0.877 |
Key Findings¶
- ICA interaction is essential for multi-image relational assessment: Removing the ICA module causes \(R^2\) to drop sharply from 0.553 to 0.372, and \(\rho_{\text{SRCC}}\) falls by 0.133. Independent pooling fails to capture subtle garment deformations, boundary bleed, and texture mismatches, which require deep patch-level cross-attention.
- Reference-based metrics suffer from severe structural misalignment penalties: When models slightly modify posture or framing to enhance naturalness (e.g., GPT-Image-1.5), SSIM drops to 0.768—ranking it worse than early GANs. In contrast, VTON-IQA is robust to harmless global transformations and evaluates true visual and attribute fidelity.
- Proprietary models lead in visual quality, while open-source DiTs advance rapidly: Nano Banana Pro and GPT-Image-1.5 achieve the highest human-aligned scores across benchmarks. Among open-source approaches, DiT models (FitDit, CatVTON-FLUX) consistently outperform traditional U-Net diffusion and GAN baselines.
Highlights & Insights¶
- Rigorous data curation unlocks reliable perceptual modeling: By eliminating more than 60% of noisy crowdsourced responses via sanity checks and agreement thresholds, the average Krippendorff's \(\alpha\) doubled. This demonstrates that in subjective quality estimation, annotation consistency and data curation are more critical than raw sample count.
- Asymmetric cross-attention matches the physical priors of try-on: By restricting cross-attention to \(V \leftrightarrow G\) and \(V \leftrightarrow P\) and bypassing \(G \leftrightarrow P\), the architecture prevents redundant parameter learning while directly embedding the dual objectives of garment transfer and non-target identity preservation.
- Standardized plug-and-play evaluation replacement: VTON-IQA provides an open-source, reproducible, reference-free evaluation standard that can replace expensive and non-reproducible ad hoc user studies in future virtual try-on literature.
Limitations & Future Work¶
- Domain coverage constraints: The benchmark primarily focuses on single-person studio or clean catalog imagery with standing or slightly turned poses, leaving extreme occlusion, multi-person interactions, and complex dynamic poses unaddressed.
- Global pooling limitations: Aggregating features via global
[CLS]tokens and cosine similarity may under-detect tiny localized artifacts (e.g., small button misalignments or subtle lettering distortions) if dominated by large background regions. - Future applications: The differentiable nature of VTON-IQA opens promising avenues to serve as an online reward model for RLHF or direct test-time guidance in diffusion sampling pipelines.
Related Work & Insights¶
- vs VTONQA (Wei et al., 2026): VTONQA was trained on roughly 8,000 images with 40 annotators and is not fully open-source. VTON-QBench scales to 62,688 images, 431,800 annotations from 13,838 qualified workers, and provides fully open-sourced code and weights.
- vs VTON-VLLM (Wan et al., 2025): VTON-VLLM utilizes MLLMs to generate textual critiques, whereas VTON-IQA directly outputs calibrated, low-latency, differentiable scalar scores suited for benchmarking and optimization.
- vs SSIM / LPIPS: Full-reference metrics require ground-truth paired images and over-penalize zoom or pose shifts. VTON-IQA operates entirely reference-free and aligns closely with human subjective preference.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Introduces asymmetric Interleaved Cross-Attention and a reference-free evaluation formulation designed for try-on]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [62k images, 431k curated annotations, 14 evaluated models, and comprehensive human-level consistency validation]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear structural organization, thorough mathematical definitions, and detailed empirical reporting]
- Value: ⭐⭐⭐⭐⭐ [Resolves a foundational evaluation bottleneck in virtual try-on with fully open-sourced benchmarks and models]