Contrastive-Guided Self-Supervised Latent Visual Reasoning for Hallucination Mitigation¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Multimodal VLM / Hallucination / Latent Reasoning
Keywords: Hallucination Mitigation / Latent Visual Reasoning / Self-Supervised Learning / Reinforcement Learning / Contrastive Probing
TL;DR¶
CoLVR introduces an annotation-free self-supervised latent visual reasoning framework that identifies query-relevant salient regions via visual contrastive probing and aligns internal continuous latent tokens with critical visual evidence using GRPO, significantly curbing LVLM hallucinations without any manual annotations.
Background & Motivation¶
Large Vision-Language Models (LVLMs) have achieved substantial breakthroughs across diverse multimodal tasks, yet object hallucination remains a severe impediment that undermines their factual reliability in mission-critical applications. Recent investigations into attention dynamics reveal that during hallucinated token generation, the model's self-attention heavily over-relies on textual language priors while disregarding authentic visual grounding. To suppress hallucinations, prevailing paradigms predominantly focus on logit-level contrastive post-processing, attention steering, or explicit visual chain-of-thought ("thinking with images") leveraging external tools like croppers and detection agents. However, external-tool pipelines introduce prohibitive multi-turn interaction latency, tool-calling overhead, and fundamentally lack an intrinsic perceptual exploration mechanism within the model itself.
To eliminate external dependencies, latent visual reasoning has emerged to internalize intermediate visual thoughts within the model's continuous latent space, inserting continuous latent embeddings directly into the autoregressive decoding stream. Nonetheless, current latent reasoning approaches rely heavily on strong supervision: they require labor-intensive manual curation of auxiliary images or regional bounding-box chains, training latent features via supervised fine-tuning (SFT). This dependency introduces prohibitive annotation costs and establishes an intrinsic performance ceiling bounded by human-annotated granularity.
Core Idea: Leverage visual contrastive probing to filter out language priors and autonomously discover query-relevant salient regions, and guide continuous latent reasoning tokens to align with these critical visual features via GRPO reinforcement learning in an annotation-free, self-supervised manner.
Method¶
Overall Architecture¶
The overall architecture of CoLVR comprises an autoregressive latent inference interface and a two-stage self-supervised training pipeline (coarse SFT warmup followed by contrastive-guided RL). The model expands its vocabulary with three special tokens: โจlatentโฉ, โจ/latentโฉ, and a placeholder โจlatent_padโฉ. When decoding โจlatentโฉ, the model transitions into latent reasoning mode to generate a sequence of continuous latent embeddings \(Z\) of fixed length \(K\). Subsequently, โจ/latentโฉ terminates the latent mode and resumes standard linguistic text generation.
During training, Stage 1 (SFT) utilizes unannotated coarse pseudo-regions to equip the model with basic syntax for generating control tokens and mapping latent states into the visual representation subspace. Stage 2 (RL) deploys visual contrastive probing by contrasting attention maps between the original image and a perturbed image, thereby isolating query-relevant visual evidence without ground-truth labels. Finally, GRPO optimizes an alignment reward that reinforces the latent tokens to actively attend to these critical visual patches.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input: Image I and Query q"] --> Probing["1. Visual Contrastive Probing: Contrast attention between original and perturbed images to identify salient patch weights"]
In --> SFTStage["2. Coarse Latent SFT: Self-supervised pseudo-region alignment and format warmup with stop-gradient"]
Probing --> RLReward["3. Alignment Reward Design: Measure cosine similarity between latent embeddings and Top-M salient patches"]
SFTStage --> RLReward
RLReward --> GRPOTrain["4. GRPO Policy Optimization: Group-relative policy optimization without ground-truth answers"]
GRPOTrain --> Out["Output: De-hallucinated response guided by latent visual thoughts"]
Key Designs¶
1. Visual Contrastive Probing: Annotation-Free Saliency Identification Standard attention weights inherently conflate language priors, positional bias, and genuine visual grounding. To extract query-relevant visual evidence without human annotations, CoLVR constructs a visual contrastive probe. Given the original input pair \((I, q)\), the model also processes a content-irrelevant Gaussian-perturbed image \(\tilde{I}\) (empirical results confirm Gaussian noise effectively destroys semantic content while retaining low-level baseline structures better than blank images). The probe computes the cross-attention matrix from text query tokens to all visual patch tokens across all layers and attention heads, isolating the pure visual contribution via subtraction: $\(a_t = g_\theta(I, q)_t - g_\theta(\tilde{I}, q)_t\)$ where \(g_\theta(I, q)_t = \frac{1}{LH}\sum_{\ell=1}^L \sum_{h=1}^H \left(\frac{1}{|\mathcal{P}_{\mathrm{q}}|} \sum_{i \in \mathcal{P}_{\mathrm{q}}} \mathbf{A}_{i,t}^{(\ell,h)}\right)\). Rectifying negative activations and normalizing yields a valid non-negative saliency distribution \(w_t\) over image tokens. This enables the model to autonomously pinpoint critical visual regions without external detectors or ground-truth bounding boxes.
2. Coarse Latent SFT and Latent-Only Backpropagation: Robust Latent Space Formation
Before reinforcement learning, the model must acquire the ability to transition in and out of latent mode and represent spatial regions. CoLVR segments images into canonical pseudo-regions (e.g., quadrants and center) without manual curation. During autoregressive decoding at โจlatent_padโฉ slots, the last-layer hidden states are extracted as latent vectors \(z_k = h_{t_k}^{(L)}\). Their mean embedding \(\bar{z}\) is trained to maximize cosine similarity with the mean pooled representation \(r(b)\) of the corresponding pseudo-region. Crucially, to prevent shortcut degeneration where the model alters visual token features rather than learning latent representations, gradients from the alignment loss are strictly isolated via a stop-gradient operator:
$\(\widetilde{\mathcal{L}}_{\mathrm{align}} = \sum_{k=1}^K z_k^\top \mathrm{sg}(\nabla_{z_k} \mathcal{L}_{\mathrm{align}})\)$
Restricting gradient flow exclusively to latent tokens forces the latent representations to actively adapt to the visual feature space, preventing catastrophic feature degradation.
3. Contrastive Alignment Reward Driven GRPO: Label-Free Latent RL
Following the SFT warmup, the model undergoes reinforcement learning using Group Relative Policy Optimization (GRPO). For each instance, \(G\) completion candidates are sampled from the current policy. The reward function combines a format reward \(r_2\) (enforcing valid โจlatentโฉ...โจ/latentโฉ block syntax with a penalty of -1 if violated, 0 otherwise) and an alignment reward \(r_1\). In practice, hard alignment over the top-\(M\) salient patch set \(\mathcal{S}\) identified by contrastive probing achieves optimal empirical performance:
$\(r_1^{\mathrm{hard}}(I, q, y) = \frac{1}{|\mathcal{S}|} \sum_{t \in \mathcal{S}} \frac{1 + \cos(\bar{z}, h_t^{(L)})}{2}\)$
The total reward \(R = r_1 + \beta r_2\) provides group-relative advantage estimates \(A^{(i,g)}\) for policy updates. Without relying on any gold reference answers, the model self-supervises its latent tokens to anchor onto critical visual evidence, naturally mitigating object hallucination.
Loss & Training¶
In the SFT stage, the joint optimization objective balances standard cross-entropy and the latent alignment loss:
$\(\mathcal{L}_{\mathrm{SFT}} = \mathcal{L}_{\mathrm{CE}} + \lambda \widetilde{\mathcal{L}}_{\mathrm{align}}\)$
where โจlatentโฉ and โจ/latentโฉ tokens are included in \(\mathcal{L}_{\mathrm{CE}}\) to encourage autonomous mode transitions, while โจlatent_padโฉ positions are masked out; \(\lambda\) is set to 1.
In the RL stage, GRPO optimizes the policy using group advantage-weighted likelihood alongside a reference policy KL divergence constraint:
$\(\mathcal{L}_{\mathrm{GRPO}} = -\mathbb{E}_{i,g} \left[ A^{(i,g)} \sum_{t \in \mathcal{P}_{\mathrm{txt}}} \log p_\theta(y_t^{(i,g)} \mid I^{(i)}, q^{(i)}, y_{<t}^{(i,g)}) \right] + \beta \mathbb{E}_{i,g}[\mathrm{KL}(\pi_\theta \parallel \pi_{\mathrm{ref}})]\)$
LoRA fine-tuning is adopted with the vision encoder frozen. The sampling group size \(G\) is set to 8, KL coefficient \(\beta = 0.01\), training across 4 A800 GPUs in approximately 18 hours.
Key Experimental Results¶
Main Results¶
Evaluated on benchmark hallucination suites including CHAIR, POPE, Hallbench, and GPT-4-judged MMHal, CoLVR demonstrates broad superiority over train-free methods, tool-based "thinking with images" approaches, and fully supervised latent reasoning baselines across both Qwen2.5-VL-7B and 3B backbones.
| Backbone | Category | Method | Tr. | Lb. | CHAIR-Cs โ | CHAIR-Ci โ | CHAIR-F1 โ | POPE Acc. โ | Hallbench Acc. โ | MMHal Score โ | MMHal Hal. โ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-VL-7B | Baseline | Base | ร | ร | 30.8 | 8.1 | 78.3 | 84.9 | 52.5 | 4.86 | 32.3 |
| Qwen2.5-VL-7B | Train-free | VCD | ร | ร | 34.6 | 8.7 | 78.6 | 86.0 | 52.9 | 4.85 | 33.3 |
| Qwen2.5-VL-7B | Train-free | OPERA | ร | ร | 31.6 | 8.5 | 78.7 | 85.4 | 52.9 | 5.02 | 31.3 |
| Qwen2.5-VL-7B | Train-free | PAI | ร | ร | 34.4 | 8.8 | 77.8 | 88.2 | 53.2 | 5.00 | 37.5 |
| Qwen2.5-VL-7B | Tool Visual-CoT | Deepeyes-7B | โ | โ | 27.6 | 7.3 | 78.6 | 86.1 | 53.9 | 5.10 | 29.2 |
| Qwen2.5-VL-7B | Tool Visual-CoT | Pixel-Reasoner-7B | โ | โ | 23.2 | 6.7 | 77.1 | 84.8 | 42.6 | 4.77 | 33.3 |
| Qwen2.5-VL-7B | Supervised Latent | LVR-7B | โ | โ | 33.2 | 9.8 | 79.4 | 84.8 | 50.4 | 4.88 | 37.5 |
| Qwen2.5-VL-7B | Supervised Latent | Monet-7B | โ | โ | 34.1 | 9.9 | 67.8 | 86.2 | 52.1 | 5.02 | 31.3 |
| Qwen2.5-VL-7B | Ours | CoLVR-7B | โ | ร | 20.2 | 6.1 | 80.1 | 89.6 | 54.3 | 5.12 | 30.2 |
| Qwen2.5-VL-3B | Baseline | Base | ร | ร | 38.4 | 8.3 | 79.0 | 85.9 | 40.6 | 4.69 | 34.4 |
| Qwen2.5-VL-3B | Train-free | OPERA | ร | ร | 43.2 | 9.6 | 79.9 | 85.9 | 40.8 | 4.91 | 31.3 |
| Qwen2.5-VL-3B | Ours | CoLVR-3B | โ | ร | 30.8 | 8.5 | 78.6 | 88.2 | 42.6 | 4.78 | 30.2 |
Ablation Study¶
Systematic ablations on Qwen2.5-VL-3B examine the core training stages, gradient isolation, contrastive perturbation sources, and latent token lengths.
| Configuration | CHAIR-Cs โ | CHAIR-Ci โ | CHAIR-F1 โ | POPE Acc. โ | MMHal Score โ | Note |
|---|---|---|---|---|---|---|
| Base model | 38.4 | 8.3 | 79.0 | 85.9 | 4.69 | Unmodified baseline without latent mode |
| SFT-only | 38.6 | 8.5 | 78.3 | 86.0 | 4.46 | Coarse pseudo-region warmup without RL |
| Full CoLVR (SFT + RL) | 30.8 | 8.5 | 78.6 | 88.2 | 4.78 | Complete framework with contrastive RL |
| w/o stop-gradient | โ | โ | โ | 84.6 | 4.28 | Shortcut optimization damages visual features |
| Blank contrast | 36.9 | 8.9 | 76.5 | โ | โ | Under-informative baseline compared to noise |
| Original-image attn | 33.3 | 8.3 | 75.9 | โ | โ | Lacks contrastive removal of text prior |
| Reverse contrast | 46.9 | 12.8 | 67.6 | โ | โ | Inverts saliency, causing severe hallucinations |
Key Findings¶
- Contrastive RL alignment is the primary driver of hallucination mitigation: Coarse SFT alone only provides basic structural syntax without mitigating hallucination metrics. Incorporating contrastive-guided RL boosts POPE Accuracy from 86.0% to 88.2% and lowers MMHal hallucination rate from 34.4% to 30.2%.
- Latent-only backpropagation prevents representation collapse: Allowing alignment gradients to propagate into standard visual representations prompts the network to alter visual tokens rather than learning latent reasoning, degrading POPE Accuracy to 84.6% and MMHal score to 4.28.
- Gaussian perturbation provides the cleanest contrastive signal: Contrastive probing with Gaussian noise effectively wipes out semantics while preserving spatial scale, substantially outperforming blank images and plain original-image attention. Inverting the contrast direction sharply deteriorates CHAIR-Cs to 46.9%.
- Latent sequence length exhibits robust plateauing: Performance peaks at \(K=4\), balancing expressive capacity and RL optimization stability. Across \(K \in [2, 32]\), performance remains consistently superior, indicating minimal hyperparameter sensitivity.
Highlights & Insights¶
- Annotation-free visual grounding via contrastive probes: By computing cross-attention differentials against perturbed inputs, CoLVR cleanly filters out language and positional biases, turning manual bounding-box annotations into an automated, unsupervised attention signal.
- Self-supervised latent reasoning surpasses supervised baselines: Without any human-annotated Visual-CoT traces, CoLVR-7B achieves a CHAIR-Cs score of 20.2%, convincingly outperforming supervised counterparts like Monet-7B (34.1%) and LVR-7B (33.2%).
- Superior extensibility to general visual reasoning: Initialized with CoLVR's self-supervised alignment, subsequent fine-tuning on multimodal reasoning datasets (CoLVR) achieves 73.4% on MathVista and 83.8% on V, validating that contrastive self-supervised pre-alignment provides a robust foundation for general visual understanding.
Limitations & Future Work¶
- Training-time computational overhead: Generating dual-stream attention maps for original and perturbed images during contrastive probing incurs additional memory and inference overhead during the RL phase.
- Latent interpretability vs. discrete CoT: Although latent tokens can be visualized via cosine heatmaps on image patches, continuous latent vectors lack the explicit linguistic readability of natural language chain-of-thought traces.
- Future directions: Exploring dynamic latent token length allocation based on query difficulty, and extending contrastive probing to video temporal dynamics and 3D multimodal reasoning.
Related Work & Insights¶
- vs OPERA / PAI (Inference-time feature steering): Post-processing approaches adjust logits or attention during decoding without altering model representations; CoLVR internalizes perceptual attention into the model's weights via RL, delivering more consistent intrinsic de-hallucination without inference-time computational burdens.
- vs LVR / Monet (Supervised latent visual reasoning): Existing latent methods rely on dense bounding boxes and auxiliary image labels, which restricts scaling; CoLVR demonstrates that self-supervised contrastive probing and GRPO fully match and exceed supervised visual chains.
- vs Deepeyes / Pixel-Reasoner (Tool-based Visual-CoT): Multi-turn tool invocation introduces latency and error propagation from external detectors; CoLVR achieves autonomous single-turn visual reasoning entirely within the model's continuous latent space.
Rating¶
- Novelty: โญโญโญโญโญ Elegant contrastive probing to extract unsupervised saliency combined with GRPO-driven latent token alignment.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluation across multiple hallucination benchmarks, two model scales, in-depth ablations, and downstream reasoning benchmarks.
- Writing Quality: โญโญโญโญโญ Clear motivation, rigorous mathematical formulation, and well-structured empirical validation.
- Value: โญโญโญโญโญ Establishes a practical, annotation-free training paradigm for latent visual reasoning and hallucination mitigation in LVLMs.