Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/mlvlab/C2-DPO
Area: Multimodal VLM
Keywords: object hallucination, direct preference optimization, context calibration, multimodal large language models, preference gain
TL;DR¶
Reveals the "context blindness" phenomenon in existing direct preference optimization (DPO) for MLLM alignment, and proposes Context-Calibrated Direct Preference Optimization (C2-DPO), which explicitly maximizes Contextual Preference Gain (CPG) while anchoring preference ranking under degraded context, significantly cutting object hallucination without compromising general reasoning.
Background & Motivation¶
Multimodal large language models (MLLMs) integrate visual encoders with large language model backbones and large-scale multimodal instruction tuning, demonstrating notable progress across core vision-language tasks including detailed image captioning and visual question answering. Nevertheless, these models frequently suffer from the persistent object hallucination problem, generating fluent, semantically plausible descriptions of objects or attributes that contradict the actual visual input. To suppress object hallucination, preference alignment methodsโmost notably Direct Preference Optimization (DPO)โhave become the prevailing paradigm. Recent follow-up efforts (such as HA-DPO, POVID, and C-DPO) have focused heavily on enriching preference datasets, for instance by augmenting the input with non-hallucinated auxiliary captions, aiming to provide stronger contextual grounding so that non-hallucinated and hallucinated responses become more discriminative.
However, existing research leaves a fundamental question unanswered: does the standard DPO objective actually encourage an MLLM to leverage this helpful context? From an information-theoretic viewpoint, providing relevant contextual cues should reduce uncertainty and thereby strengthen the model's preference margin for the grounded response over the hallucinated one. Yet, standard DPO optimizes the preference margin on a single static input prompt; its gradient never explicitly rewards achieving a wider margin when richer context is supplied. To investigate this, the authors introduce Contextual Preference Gain (CPG), a diagnostic metric quantifying how much the model's preference margin expands when auxiliary grounding context is provided. Empirical inspection reveals two striking findings: first, CPG exhibits a strong negative correlation with benchmark hallucination rates, meaning models with higher CPG consistently hallucinate less; second, standard DPO and its variants (SimPO, RDPO) show CPG distributions tightly clustered around zero or negative values, demonstrating pervasive "context blindness."
This implies that modifying preference data alone cannot guarantee that the model actually internalizes contextual cues during alignment. This paper's angle of attack is that preference optimization should explicitly model and amplify the preference gain between full and degraded contexts, while anchoring the baseline preference ordering when auxiliary context is absent. Core idea: introduce Context-Calibrated Direct Preference Optimization (C2-DPO), which formulates a contrastive Contextual Preference Calibration loss to directly maximize Contextual Preference Gain (CPG) alongside a degraded-context DPO anchor, significantly suppressing object hallucination while preserving general multimodal reasoning capabilities.
Method¶
Overall Architecture¶
C2-DPO is designed to eliminate context blindness and enforce context-aware preference calibration during multimodal alignment. The input space consists of a full context \(x = (v, q, c)\) (comprising an image \(v\), a text query \(q\), and an auxiliary image description \(c\)) and a degraded counterpart \(x' = (v, q, \emptyset)\) where the auxiliary description is stripped. Training candidate pairs consist of a non-hallucinated response \(y_w\) and a hallucinated response \(y_l\).
The entire framework targets a dual-constraint preference hierarchy: on one hand, the preference margin under the full context \(x\) must strictly exceed the margin under the degraded context \(x'\), yielding a strictly positive CPG; on the other hand, under the degraded input \(x'\), the model must maintain a positive preference margin for \(y_w\) over \(y_l\). C2-DPO relaxes these constraints into differentiable surrogate objectives, forming a unified three-component loss: the full-context DPO loss, the Contextual Preference Calibration loss, and the degraded-context DPO anchoring loss.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Multimodal Input Construction<br/>Full context x=(v,q,c) & Degraded context x'=(v,q,โ
)"] --> B["Policy & Reference Forward Pass<br/>Compute implicit reward margins ฮr_ฮธ(x) & ฮr_ฮธ(x')"]
B --> C["Contextual Preference Gain Diagnosis<br/>Quantify CPG(x, x') = ฮr_ฮธ(x) - ฮr_ฮธ(x')"]
C --> D["Contextual Preference Calibration<br/>Contrastive NCE loss maximizing CPG"]
C --> E["Degraded Context Preference Anchoring<br/>DPO objective securing base visual ranking"]
D --> F["Joint Multi-Objective Optimization<br/>Weighted combination of primary & surrogate terms"]
E --> F
F --> G["Output: Aligned MLLM with low hallucination & strong reasoning"]
Key Designs¶
1. Contextual Preference Gain Diagnosis: Quantifying contextual sensitivity and context blindness
To quantify the extent to which an MLLM relies on input context during preference optimization, the authors formulate the preference score based on the implicit reward margin derived from the Bradley-Terry model: $\(\Delta \hat{r}_\theta(x, y_w, y_l) = \hat{r}_\theta(x, y_w) - \hat{r}_\theta(x, y_l) = \beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}\)$ This score measures how decisively the model prefers the grounded response \(y_w\) over the hallucinated response \(y_l\) under context \(x\). Building on this, the authors formalize Contextual Preference Gain (CPG): $\(\text{CPG}(x, x') = \Delta \hat{r}_\theta(x, y_w, y_l) - \Delta \hat{r}_\theta(x', y_w, y_l)\)$ A positive CPG indicates that supplementary contextual information effectively amplifies the preference for correct visual grounding. Conversely, when CPG hovers around zero, the model's preference margin remains unchanged regardless of context availability, diagnosing context blindness. Diagnostic evaluation shows that existing preference optimization objectives (DPO, SimPO, RDPO) cluster near zero, whereas CPG strongly correlates negatively with hallucination rates on Object HalBench and AMBER, proving that scaling CPG is pivotal for hallucination reduction.
2. Contextual Preference Calibration: Maximizing CPG via contrastive regularization
To ensure that the model assigns a stronger preference margin when presented with richer context, the target preference ordering requires \(\Delta \hat{r}_\theta(x, y_w, y_l) > \Delta \hat{r}_\theta(x', y_w, y_l)\). To optimize this inequality through backpropagation, C2-DPO establishes a smooth, binary NCE-style contrastive calibration loss, treating the full context \(x\) as the positive instance and the degraded context \(x'\) as the negative instance: $\(\mathcal{L}_c(x, x') = - \log \frac{\exp(\Delta \hat{r}_\theta(x, y_w, y_l))}{\exp(\Delta \hat{r}_\theta(x, y_w, y_l)) + \exp(\Delta \hat{r}_\theta(x', y_w, y_l))} = - \log \sigma \left( \Delta \hat{r}_\theta(x, y_w, y_l) - \Delta \hat{r}_\theta(x', y_w, y_l) \right)\)$ Minimizing \(\mathcal{L}_c\) monotonically increases \(\text{CPG}(x, x')\), compelling the policy model to actively utilize auxiliary context \(c\) rather than ignoring it, converting contextual grounding into an explicit training signal.
3. Degraded Context Preference Anchoring: Preventing degenerate margins and ranking collapse
Optimizing only the full-context DPO loss and the calibration loss \(\mathcal{L}_c\) can lead to a degenerate failure mode: the gradient might widen \(\text{CPG}(x, x')\) not by improving grounding on \(x\), but by driving \(\Delta \hat{r}_\theta(x', y_w, y_l)\) to arbitrarily large negative values, destroying the model's preference ranking when auxiliary captions are missing at test time. To prevent this pathology, C2-DPO imposes the baseline constraint \(\Delta \hat{r}_\theta(x', y_w, y_l) > 0\) via a standard DPO loss on the degraded context \(x'\): $\(\mathcal{L}_{\text{DPO}}(x') = - \log \sigma \left( \Delta \hat{r}_\theta(x', y_w, y_l) \right)\)$ This anchoring loss serves as a vital safeguard, guaranteeing that even when stripped of all auxiliary textual descriptions, the model continues to prefer grounded responses over hallucinated ones based purely on the original visual input and prompt query.
Loss & Training¶
Combining the primary full-context objective with both surrogate regularization terms yields the complete C2-DPO training objective: $\(\mathcal{L}_{\text{C}^2\text{-DPO}}(x, x') = \mathcal{L}_{\text{DPO}}(x) + \lambda_c \mathcal{L}_c(x, x') + \lambda_u \mathcal{L}_{\text{DPO}}(x')\)$ where hyperparameters \(\lambda_c > 0\) and \(\lambda_u > 0\) balance calibration strength against degraded-context anchoring. The models are fine-tuned for one epoch using LoRA and AdamW with a learning rate of \(2 \times 10^{-6}\), global batch size 64, and \(\beta = 0.1\). Sensitivity analysis demonstrates stable performance within \([0.3, 0.5]\); default weights are set to \((\lambda_c, \lambda_u) = (0.3, 0.5)\) for LLaVA-v1.5-7B and \((0.5, 0.3)\) for Qwen2-VL-Instruct-2B.
Key Experimental Results¶
Main Results¶
C2-DPO is evaluated across two representative architectures (LLaVA-v1.5-7B and Qwen2-VL-Instruct-2B) against leading contrastive decoding methods (VCD, OPERA, DoLa) and preference optimization methods (HA-DPO, POVID, CLIP-DPO, RLAIF-V, TPO, vanilla-DPO, C-DPO) across hallucination benchmarks and general multimodal reasoning benchmarks.
| Model | Method | Object HalBench Rsp. โ | Object HalBench Men. โ | AMBER CHAIR โ | AMBER Hal. โ | AMBER Cog. โ | ScienceQA Image Acc. โ | MM-Vet Overall โ | TextVQA Acc. โ |
|---|---|---|---|---|---|---|---|---|---|
| LLaVA-v1.5-7B | Base | 52.7 | 28.0 | 8.4 | 35.5 | 4.0 | 66.8 | 31.0 | 58.2 |
| LLaVA-v1.5-7B | VCD | 51.3 | 25.9 | 9.1 | 39.8 | 4.2 | 68.7 | 29.8 | 56.1 |
| LLaVA-v1.5-7B | OPERA | 45.3 | 22.9 | 6.5 | 28.5 | 3.1 | 68.2 | 30.3 | 58.2 |
| LLaVA-v1.5-7B | DoLa | 44.0 | 25.1 | 6.2 | 27.7 | 2.9 | 67.5 | 30.8 | 56.6 |
| LLaVA-v1.5-7B | HA-DPO | 37.0 | 20.9 | 6.7 | 30.9 | 3.3 | 69.7 | 30.6 | 56.7 |
| LLaVA-v1.5-7B | POVID | 33.4 | 16.6 | 5.3 | 28.7 | 3.0 | 68.8 | 31.8 | 56.6 |
| LLaVA-v1.5-7B | RLAIF-V | 7.8 | 4.2 | 2.8 | 15.7 | 0.9 | 68.2 | 29.9 | 55.1 |
| LLaVA-v1.5-7B | TPO | 5.6 | 3.2 | 3.6 | 20.5 | 1.6 | 67.1 | 25.7 | 55.3 |
| LLaVA-v1.5-7B | vanilla-DPO | 27.4 | 15.9 | 5.4 | 25.5 | 2.5 | 69.3 | 32.0 | 58.2 |
| LLaVA-v1.5-7B | C-DPO | 5.9 | 3.3 | 3.0 | 14.9 | 1.3 | 69.4 | 33.4 | 58.2 |
| LLaVA-v1.5-7B | C2-DPO | 4.8 | 2.7 | 2.8 | 13.8 | 1.2 | 69.5 | 33.4 | 58.2 |
| Qwen2-VL-Instruct-2B | Base | 16.1 | 8.3 | 6.2 | 35.5 | 2.7 | 76.9 | 49.9 | 78.2 |
| Qwen2-VL-Instruct-2B | vanilla-DPO | 6.0 | 3.6 | 4.2 | 39.9 | 3.2 | 77.0 | 49.8 | 78.4 |
| Qwen2-VL-Instruct-2B | C-DPO | 2.5 | 1.6 | 2.7 | 17.5 | 0.8 | 77.3 | 48.2 | 78.4 |
| Qwen2-VL-Instruct-2B | C2-DPO | 1.6 | 1.0 | 2.7 | 16.1 | 0.9 | 77.2 | 49.1 | 78.4 |
Ablation Study¶
The ablation isolates each loss component to verify their complementary synergy:
| Model | \(\mathcal{L}_{\text{DPO}}(x)\) | \(\mathcal{L}_c(x, x')\) | \(\mathcal{L}_{\text{DPO}}(x')\) | Object HalBench Rsp. โ | Object HalBench Men. โ | AMBER CHAIR โ | AMBER Hal. โ | AMBER Cog. โ |
|---|---|---|---|---|---|---|---|---|
| LLaVA-v1.5-7B | โ | โ | โ | 5.9 | 3.3 | 3.0 | 14.9 | 1.3 |
| LLaVA-v1.5-7B | โ | โ | โ | 7.1 | 4.1 | 3.1 | 15.9 | 1.2 |
| LLaVA-v1.5-7B | โ | โ | โ | 5.9 | 3.2 | 3.1 | 15.0 | 1.3 |
| LLaVA-v1.5-7B | โ | โ | โ | 4.8 | 2.7 | 2.8 | 13.8 | 1.2 |
| Qwen2-VL-Instruct-2B | โ | โ | โ | 2.5 | 1.6 | 2.7 | 17.5 | 0.8 |
| Qwen2-VL-Instruct-2B | โ | โ | โ | 3.3 | 2.6 | 2.7 | 22.5 | 1.1 |
| Qwen2-VL-Instruct-2B | โ | โ | โ | 3.5 | 2.3 | 3.2 | 18.6 | 1.0 |
| Qwen2-VL-Instruct-2B | โ | โ | โ | 1.6 | 1.0 | 2.7 | 16.1 | 0.9 |
Furthermore, applying Contextual Preference Calibration to other preference optimization schemes yields consistent gains: C2-SimPO cuts Object HalBench Rsp. from 3.9 to 3.3, and C2-RDPO drops Rsp. from 3.4 to 1.7 (a 50% relative reduction). In text-only instruction alignment with Qwen2.5-Instruct-1.5B on AlpacaEval 2, C2-DPO raises length-controlled win rate (LC WR) from 24.1% to 25.9%, and C2-SimPO improves from 33.5% to 34.1%.
Key Findings¶
- Essential synergy between calibration and anchoring: Activating either \(\mathcal{L}_c\) alone or \(\mathcal{L}_{\text{DPO}}(x')\) alone impairs hallucination rates due to margin imbalance (e.g., on Qwen2-VL, Rsp. degrades from 2.5 to 3.3 or 3.5). Strong improvements occur only when both are jointly trained (dropping to 1.6), proving the theoretical need for the lower-bound anchor.
- Continuous CPG growth during training: While baseline C-DPO stagnates near zero or negative CPG throughout optimization, C2-DPO demonstrates steady, monotonic CPG increases, tracking increased sensitivity to sentence- and word-level grounding cues.
- Robustness to noisy context with zero reasoning sacrifice: Under random caption masking up to 50%, C2-DPO retains substantially lower hallucination rates than C-DPO. Moreover, performance across ScienceQA, MM-Vet, and TextVQA remains completely intact, avoiding the traditional alignment tax.
Highlights & Insights¶
- Novel diagnostic insight into context blindness: Moves beyond the conventional paradigm of merely altering dataset prompts, identifying a structural flaw in the standard DPO objective that ignores contextual richness.
- Principled mathematical formulation: Translates the dual ordering constraint \(\Delta \hat{r}_\theta(x) > \Delta \hat{r}_\theta(x') > 0\) into a neat, differentiable combination of binary InfoNCE calibration and anchor regularization.
- Broad cross-objective and cross-modal generality: Seamlessly integrates with SimPO and RDPO, and proves equally effective for text-only LLMs without introducing test-time inference overhead.
Limitations & Future Work¶
- Context construction dependency: Relies on paired auxiliary captions during training; autonomous synthesis of fine-grained contextual pairs for open-ended queries warrants further exploration.
- Scope of degradation modes: Degraded inputs are primarily created by removing auxiliary text (\(x'=(v,q,\emptyset)\)); exploring visual degradations (e.g., partial image occlusion, lower resolution, visual perturbations) remains future work.
- Extension to online RL: The CPG formulation could be adapted to online trajectory-level RL (such as PPO or GRPO), measuring token-level preference gains with respect to evolving visual contexts.
Related Work & Insights¶
- vs C-DPO (ICCV 2025): C-DPO introduces auxiliary captions during training but relies on standard DPO, resulting in context blindness; C2-DPO maximizes CPG explicitly, yielding a 36% to 60% relative reduction in hallucination on Qwen2-VL using the exact same data.
- vs Contrastive Decoding (VCD / OPERA / DoLa): Contrastive decoding methods require multiple forward passes at inference time, adding heavy latency; C2-DPO operates purely during alignment training with zero test-time overhead.
- vs Dataset Synthesis (HA-DPO / POVID / RLAIF-V): Data-centric approaches incur heavy annotation costs and are constrained by external judge capabilities; C2-DPO introduces an objective-level improvement orthogonal to data quality.
Rating¶
- Novelty: โญโญโญโญโญ Formulates the context blindness phenomenon, defines the diagnostic CPG metric, and designs a clean dual-constraint objective.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluations across multiple models, benchmarks, loss ablations, noise robustness, and text-only extensions.
- Writing Quality: โญโญโญโญโญ Well-structured narrative transitioning smoothly from information-theoretic motivation to elegant mathematical formulation.
- Value: โญโญโญโญโญ Solves a critical alignment bottleneck in MLLMs; code is fully open-source and easily adaptable as a standard preference training objective.