Iterative Perceptual Alignment for VLMs via Deterministic Reconstruction Feedback¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/XC0053/iterative-perceptual-alignment
Area: Multimodal VLM
Keywords: Vision-Language Alignment, Deterministic Reconstruction Feedback, Direct Preference Optimization (DPO), Reward Model, Self-Supervised Data Generation
TL;DR¶
This paper introduces a self-supervised vision-language alignment framework based on deterministic visual reconstruction feedback, utilizing a few-step diffusion distillation engine and perceptual distance evaluation to construct a "reconstruct-compare-refine" loop, which trains a highly discriminative reward model using only 20% of typical data and substantially improves VLM fine-grained attribute grounding via DPO.
Background & Motivation¶
The advancement of vision-language models (VLMs) and text-to-image (T2I) synthesis increasingly relies on large-scale, fine-grained preference alignment data. Traditional multi-modal preference dataset construction primarily depends on costly human crowdsourcing or heuristic negative sample synthesis. However, human annotations are hindered by prohibitive expenses, subjective bias, and label noise, whereas heuristic methods—such as phrase truncation or random attribute swapping—create negative samples with excessive semantic gaps. These discrete, coarse negatives lack the smooth gradient discriminability required to guide models toward subtle visual nuances. Concurrently, using existing VLMs as evaluators (VLM-as-a-judge) frequently introduces hallucinations and inference instability, while standard automated metrics like CLIP Score operate at a coarse global matching level and remain blind to high-density detailed attributes and relational semantics.
The core tension is that evaluating fine-grained multimodal alignment demands an objective perceptual ground truth devoid of text-only heuristics or evaluator hallucinations, yet scoring semantic fidelity purely in abstract linguistic space is inherently ambiguous. If a generated caption faithfully captures every subtle visual attribute and spatial configuration of a reference image, mapping those semantics back to the visual space through a highly deterministic synthesis engine should yield a reconstruction that is perceptually indistinguishable from the original image. Conversely, if the visual generator introduces substantial stochastic disturbances, reconstruction discrepancies cannot be deterministically attributed to caption defects.
This paper's angle of attack is to transform elusive semantic scoring into a concrete, observable perceptual reconstruction closed loop, suppressing generator stochasticity to a negligible level via low-randomness distilled diffusion transformers. Core idea: leverage a few-step deterministic diffusion distillation engine and visual perceptual distance to construct a self-supervised "reconstruct-compare-refine" loop, generating continuous quality-gradient preference trajectories to train a lightweight consistency reward model and drive VLM Direct Preference Optimization (DPO).
Method¶
Overall Architecture¶
The proposed approach models vision-language alignment as a reconstruction-driven optimization problem in visual perceptual space: given a reference image \(I_{\text{ref}}\), the goal is to find an optimal description \(C^*\) that minimizes the perceptual distance \(\mathcal{D}(I_{\text{ref}}, g(C))\) of the reconstructed image \(g(C)\). The framework operates across three distinct phases: multi-candidate initialization, closed-loop semantic refinement, and preference trajectory extraction. In initialization, candidate captions from multiple diverse VLMs are generated, reconstructed deterministically via Z-Image-Turbo, and evaluated by DreamSim to identify the optimal starting caption \(C_0\). Next, a strong Visual Critic drives an iterative "reconstruct-compare-refine" loop by performing cross-image comparative analysis to eliminate attribute omissions and spatial discrepancies until perceptual convergence. Finally, complete optimization trajectories are mined to generate dense training pairs for reward modeling as well as adjacent hard pairs for VLM preference fine-tuning.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Reference Image I_ref"] --> B["Candidate Initialization<br/>Multi-VLM generation + DreamSim C0 filtering"]
B --> C["Closed-Loop Semantic Refinement<br/>Critic cross-image discrepancy + Z-Image-Turbo"]
C -->|sk > τ and k < K| C
C -->|sk ≤ τ or k = K| D["Preference Trajectory Extraction<br/>T = (C0, s0), ..., (CK, sK)"]
D --> E["Decomposed Consistency Scoring Model<br/>Cross-attention interaction decoupled from text quality"]
D --> F["Direct Preference Optimization (DPO)<br/>Adjacent high-quality chosen-rejected pairs"]
Key Designs¶
1. Candidate Initialization: Eliminating Cold-Start Bias and Sampling Noise Sampling initial candidates from a single VLM easily traps the generation process in that specific architecture's stylistic and semantic blind spots. To establish a robust initialization pool, the framework queries four distinct open-source VLMs—Qwen3-VL-8B, InternVL3.5-8B, Gemma-3n-E4B-it, and LLaVA-1.5-7B—yielding candidate set \(\{C_1, \dots, C_n\}\) for each reference image \(I_{\text{ref}}\). Crucially, to prevent visual evaluation from being corrupted by generation randomness, reconstruction is performed by Z-Image-Turbo \(g(\cdot)\), a single-stream diffusion transformer utilizing few-step deterministic sampling. Perceptual alignment is measured via DreamSim \(s_i = h(I_{\text{ref}}, I_i)\), which mirrors human perceptual granularity. The candidate with the minimum perceptual distance is selected as \(C_0\). Empirical verification demonstrates that generator variance is negligible compared to inter-caption score differences, ensuring visual disparities faithfully reflect caption fidelity.
2. Closed-Loop Semantic Refinement: Directing Optimization via Visual Discrepancy When the perceptual score \(s_{k-1}\) of the previous iteration fails to satisfy the target convergence threshold \(\tau\), the reference image \(I_{\text{ref}}\) and the current reconstructed image \(I_{k-1}\) are provided side-by-side to a Visual Critic (e.g., GPT-5.1). Rather than producing an arbitrary numerical rating, the critic conducts explicit cross-image visual comparison, isolating precise semantic gaps—such as missing background elements, incorrect material textures, or erroneous spatial relations—in caption \(C_{k-1}\) and performing targeted semantic editing to produce an evolved caption \(C_k\). The refined caption is deterministically reconstructed into image \(I_k = g(C_k)\) and re-evaluated by DreamSim. If \(s_k \le s_{k-1}\), the refinement is validated, and the loop repeats until \(s_k \le \tau\) or the maximum iteration limit \(K=4\) is reached. This design grounds every text edit in verifiable image fidelity improvements, producing an evolutionary trajectory \(\mathcal{T} = \{(C_0, s_0), (C_1, s_1), \dots, (C_K, s_K)\}\) with monotonically decreasing perceptual distance.
3. Preference Trajectory Extraction: Dual Strategies for Pre-training and DPO Directly leveraging uncurated combinatorial pairs from trajectories creates severe noise and trivial contrasts. The framework applies two task-specific filtering strategies. For reward model pre-training, ordered pairs \((C_i, C_j)\) are drawn from the same reference image's trajectory and filtered by a score margin threshold \(\delta > 0\) (set to 0.03) to eliminate ambiguous samples: $\(\mathcal{D}_{\text{pair}} = \left\{ (C_i, C_j) \mid s_j - s_i \ge \delta, \, (C_i, s_i), (C_j, s_j) \in \mathcal{T} \right\}\)$ For DPO training, where policy optimization requires high-fidelity contrastive supervision, the framework selects the global best caption \(C_{k^*}\) (lowest score) and the suboptimal runner-up \(C_{m^*}\) (second-lowest score) to form pair \((C^{\text{chosen}}, C^{\text{rejected}})\), enforcing \(s_{m^*} - s_{k^*} \ge \tau_{\text{dpo}}\). This "adjacent-pair strategy" guarantees that chosen and rejected captions share overall structural and semantic validity, differing strictly in nuanced visual details. This avoids uninformative trivial gradients while maintaining clear causal discriminability.
4. Decomposed Consistency Scoring Model: Mitigating Text-Length and Fluency Shortcuts Invoking full diffusion reconstruction at inference time is computationally prohibitive, necessitating a standalone regression scoring model \(f_\theta(I_{\text{ref}}, C)\). Standard multimodal reward models often bypass visual grounding by exploiting text-only shortcuts such as length and lexical fluency. The proposed scoring architecture addresses this by pairing a frozen InternVL3.5-1B encoder (\(D=1024\)) with a decoupled dual-branch design. The cross-modal branch feeds text tokens as queries and image tokens as keys/values across \(L=4\) cross-attention blocks, using a learnable CLS token and a 5-layer MLP to predict consistency score \(s_{\text{cross}}\). Simultaneously, an independent text-only branch scores the unconditioned text tokens via an isolated MLP, predicting text quality score \(s_{\text{text}}\). The final scalar reward subtracts text quality: $\(r = f_\theta(I_{\text{ref}}, C) = s_{\text{cross}}(I_{\text{ref}}, C) - s_{\text{text}}(C)\)$ This subtraction penalizes the model for relying solely on linguistic fluency. The composite training objective optimizes \(\mathcal{L} = \mathcal{L}_{\text{main}} + \lambda_t \mathcal{L}_{\text{text}} + \lambda_c \mathcal{L}_{\text{contrast}} + \lambda_x \mathcal{L}_{\text{cross}}\), applying Bradley-Terry ranking loss with \(s_{\text{text}}\) detached alongside in-batch negative image permutations.
Loss & Training¶
Direct Preference Optimization (DPO) aligns the VLM policy on adjacent pairs \((C_w, C_l)\) using the standard Bradley-Terry parameterization without explicit reward modeling: $\(\mathcal{L}_{\text{DPO}}(\pi_\theta; \pi_{\text{ref}}) = -\mathbb{E}_{(I_{\text{ref}}, C_w, C_l) \sim \mathcal{D}_{\text{dpo}}} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(C_w \mid I_{\text{ref}})}{\pi_{\text{ref}}(C_w \mid I_{\text{ref}})} - \beta \log \frac{\pi_\theta(C_l \mid I_{\text{ref}})}{\pi_{\text{ref}}(C_l \mid I_{\text{ref}})} \right) \right]\)$ where the KL penalty coefficient is set to \(\beta = 0.1\). Base models undergo full-parameter fine-tuning for 1 epoch to prevent overfitting, with a learning rate of \(5 \times 10^{-6}\) governed by a cosine decay schedule with 10% warmup, optimized with per-device batch size 4, gradient accumulation steps 8, and gradient checkpointing.
Key Experimental Results¶
Main Results¶
The standalone T2I-Consistency-Score model was evaluated on in-domain held-out splits (1k unseen reference images) across three difficulty tiers and zero-shot tested on public preference benchmarks.
In-domain preference classification accuracy (%):
| Model | Best-Worst | Best-Second | Best-Second-Thr |
|---|---|---|---|
| T2I-Consistency-Score (Ours) | 94.44 | 60.91 | 76.80 |
| Cyclereward-Combo | 90.00 | 55.35 | 62.00 |
| Cyclereward-T2I | 66.97 | 52.53 | 51.00 |
| Cyclereward-I2T | 90.00 | 55.35 | 61.33 |
| HPSv2 | 66.87 | 55.86 | 62.43 |
| PickScore | 51.01 | 54.04 | 59.67 |
| LongCLIP-L | 87.11 | 52.63 | 54.70 |
Zero-shot generalization performance on public preference benchmarks (Accuracy %):
| Model | RLHF-V | POVID | VLFeedback | RLAIF-V | Training Triplets |
|---|---|---|---|---|---|
| T2I-Consistency-Score (Ours) | 74.12 | 67.43 | 75.10 | 53.47 | ~223k (20%) |
| Cyclereward-Combo | 65.25 | 77.41 | 59.38 | 54.09 | ~866k (100%) |
| Cyclereward-T2I | 57.30 | 86.49 | 39.21 | 55.46 | - |
| Cyclereward-I2T | 63.52 | 78.90 | 54.39 | 54.22 | - |
| HPSv2 | 63.52 | 82.69 | 53.83 | 53.62 | - |
| PickScore | 59.73 | 76.88 | 48.49 | 51.94 | - |
| LongCLIP-L | 58.55 | 84.15 | 58.03 | 54.33 | - |
Ablation Study¶
Evaluation of DPO alignment across varying data thresholds (\(\tau_{\text{dpo}} > 0.05\) yielding ~3k pairs, \(\tau_{\text{dpo}} > 0.03\) yielding ~5k pairs) against base models and the human-annotated RLHF-V (3.5k) baseline across comprehensive VLM benchmarks:
| Model Configuration | POPE (%) | Hallusion (%) | MME-P | MME-C | RealWorldQA | MM-Bench | MMStar | Note |
|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-2B-Instruct | ||||||||
| Base (No DPO) | 88.85 | 46.27 | 1510.46 | 125.71 | 0.66 | 0.66 | 0.42 | Unaligned base policy |
| + Ours (> 0.05, 3k) | 89.26 | 56.99 | 1486.91 | 121.43 | 0.63 | 0.70 | 0.45 | High threshold favors detailed attribute reasoning |
| + Ours (> 0.03, 5k) | 89.32 | 57.83 | 1493.92 | 121.43 | 0.63 | 0.67 | 0.45 | Broader dataset yields superior anti-hallucination |
| + RLHF-V (3.5k) | 89.55 | 56.15 | 1478.92 | 127.86 | 0.64 | 0.69 | 0.46 | Human-annotated baseline |
| InternVL3.5-1B-HF | ||||||||
| Base (No DPO) | 85.95 | 54.47 | 1369.87 | 112.86 | 0.45 | 0.32 | 0.25 | Base 1B model capabilities |
| + Ours (> 0.05, 3k) | 85.12 | 51.10 | 1357.73 | 117.14 | 0.52 | 0.49 | 0.34 | Capacity trade-off: shifts toward dense descriptions |
| + Ours (> 0.03, 5k) | 85.00 | 51.74 | 1359.83 | 117.14 | 0.52 | 0.49 | 0.33 | Massive comprehension gains; slight binary regression |
| + RLHF-V (3.5k) | 85.03 | 51.31 | 1351.56 | 115.00 | 0.51 | 0.50 | 0.33 | Exhibits identical capacity bottleneck pattern |
| InternVL3.5-2B-HF | ||||||||
| Base (No DPO) | 86.67 | 59.62 | 1498.41 | 129.29 | 0.39 | 0.35 | 0.33 | Base 2B model |
| + Ours (> 0.05, 3k) | 88.66 | 60.15 | 1532.34 | 135.71 | 0.35 | 0.42 | 0.36 | Consistent across-the-board gains |
| + Ours (> 0.03, 5k) | 88.54 | 60.99 | 1518.25 | 136.43 | 0.36 | 0.42 | 0.36 | Leads on Hallusion anti-illusion |
| + RLHF-V (3.5k) | 88.56 | 60.46 | 1518.54 | 138.57 | 0.35 | 0.43 | 0.38 | Human-annotated baseline comparison |
Key Findings¶
- Fine-Grained Discriminability Outperforms Baselines: On the challenging Best-Second task where candidate captions differ only in subtle attributes, existing reward models hover around random guessing (52%–55%), while the proposed model reaches 60.91% and advances to 76.80% with score-margin thresholding.
- Superior Sample Efficiency: Using only ~223k preference triplets (roughly 20% of CycleReward-Combo's 866k training volume), the scoring model outperforms competitors by 8–15% on zero-shot benchmarks RLHF-V (74.12%) and VLFeedback (75.10%).
- Model Capacity Bifurcation Between 1B and 2B: Fine-tuning the 1B backbone trades binary calibration (mild drops on POPE and Hallusion) for substantial leaps in multimodal perception (MM-Bench jumping from 0.32 to 0.49 and MMStar from 0.25 to 0.34). At 2B parameters, the model absorbs fine-grained descriptive alignment without sacrificing discrimination, registering widespread gains across benchmarks.
- Precision vs. Coverage Threshold Tuning: The stringent \(\tau_{\text{dpo}} > 0.05\) threshold isolates sharp attribute contrasts beneficial for structured reasoning (MM-Bench), whereas the broader \(\tau_{\text{dpo}} > 0.03\) pool provides richer variation that enhances generalized anti-hallucination metrics.
Highlights & Insights¶
- Grounding Abstract Alignment in Observable Visual Synthesis: The framework sidesteps text-domain evaluation ambiguity by using deterministic diffusion reconstruction, turning subjective consistency scoring into an objective perceptual distance minimization problem.
- Score Decomposition Eliminates Length and Fluency Shortcuts: Introducing an explicit text-only scoring branch that is subtracted from the cross-attention reward prevents the model from falling into trivial heuristics such as preferring verbose or ornate language.
- Adjacent Optimization Trajectory Extraction: Generating hard negative samples with minimal perceptual distance from the optimal caption provides dense, non-trivial learning signals, bypassing the flat gradients typical of heuristic negative generation.
Limitations & Future Work¶
- Residual Stochasticity in Base Generators: Although few-step distillation significantly stabilizes synthesis, diffusion models still present subtle inference variance when rendering intricate multi-object compositions or complex text typography.
- Synthetic-to-Real Domain Gap: Training on synthetic images (NIGHTS dataset) leads to weaker out-of-distribution transfer when evaluated on real-world photographic spatial reasoning (e.g., RealWorldQA) and photo-specific distortion benchmarks (e.g., POVID).
- Future Directions: Incorporating explicit spatial condition controls (such as Depth or Edge conditioning) into the reconstruction loop and synthesizing preference trajectories across heterogeneous, real-world photographic corpora.
Related Work & Insights¶
- vs CycleReward: CycleReward relies on bidirectional cycle consistency using unconstrained random sampling; in contrast, this paper develops a directed "reconstruct-compare-refine" trajectory with explicit score decomposition to suppress length bias, surpassing CycleReward with one-fifth of the training data.
- vs RLHF-V / RLAIF-V: RLHF-V depends heavily on labor-intensive human correction, while RLAIF-V inherits the perceptual blind spots and hallucinations of high-level VLMs; this work replaces subjective evaluation with objective perceptual reconstruction distance, achieving comparable or superior alignment quality fully self-supervised.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Clean conceptualization of deterministic visual reconstruction feedback coupled with evolutionary preference trajectory extraction.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across in-domain difficulty tiers, zero-shot benchmarks, and multiple VLM backbones under varying DPO thresholds.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulation, disciplined methodology description, and thorough analysis of capacity-dependent model behaviors.
- Value: ⭐⭐⭐⭐⭐ Offers a cost-effective, self-supervised blueprint to bypass the data curation bottlenecks currently limiting fine-grained multimodal alignment.