Skip to content

Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation

Conference: ECCV 2026
Paper: ECCV Official
Cached Fulltext: PDF Link
Area: Medical Imaging
Keywords: Medical Report Generation, Direct Preference Optimization, Entity-level Clinical Diagnostic, Vision-Language Alignment, Counterfactual Uncertainty Mitigation

TL;DR

Addressing the issues of conventional Direct Preference Optimization (DPO) in medical report generation—namely the confounding of clinical facts with stylistic variations and the absence of visual grounding—this paper proposes DPO-Clin, which utilizes entity-level clinical diagnostics to craft linguistically-aligned report pairs, introduces context-triggered preference inversion (M2DPO), and suppresses latent clinical risks through counterfactual substitution.

Background & Motivation

Automatic Medical Report Generation (MRG) aims to accurately convert medical images (such as chest radiographs and endoscopic views) into structured, clinically sound diagnostic reports, substantially easing clinician workloads and curbing diagnostic discrepancies. With recent progress in vision-language models, the dominant paradigm has transitioned from bespoke multi-modal attention networks and explicit knowledge graph augmentations toward adapting general-purpose Large Language Models (LLMs) via Supervised Fine-Tuning (SFT). However, SFT fundamentally relies on uniform token-level cross-entropy loss, optimizing the imitation of surface lexical distributions rather than diagnostic validity. It cannot discriminate between essential pathological assertions and superficial connective phrasing, nor does it grasp clinical preference utility, inevitably yielding persistent hallucinations and critical medical errors.

To better align model outputs with clinical utility, post-training methods using Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have gained substantial traction. Although DPO removes the training instability and reward-hacking risks inherent in RLHF reward modeling, existing DPO-based MRG approaches naively construct preference pairs by directly setting the model's erroneous predictions as dispreferred completions and the ground-truth (GT) reports as preferred ones. This naive formulation introduces three major pitfalls: first, critical clinical findings (disease entities and assertion states) are inextricably interwoven with clinically irrelevant linguistic idiosyncrasies (connectives, syntactic styles, word order), prompting the model to favor stylistic proximity over clinical factuality; second, traditional DPO evaluates textual preference solely conditioned on a static reference image, lacking contrastive visual references and failing to ground diagnostic statements in fine-grained pathological evidence; third, it focuses strictly on overt textual errors while neglecting "latent risks"—predictions that are superficially accurate yet carry high uncertainty and low confidence.

The paper addresses these challenges by isolating clinical discrepancies from linguistic noise and anchoring preference shifts in discriminative visual cues. Core idea: construct linguistically-aligned preference reports via an Entity-level Clinical Diagnostic (ECD) module to isolate factual updates from phrasing style, introduce retrieval-augmented multi-modal preference inversion (M2DPO) triggered by visual context switches, and identify high-uncertainty correct predictions to construct counterfactual dispreferred reports for comprehensive latent risk mitigation.

Method

Overall Architecture

DPO-Clin is a post-training preference alignment framework tailored for supervised fine-tuned MRG baselines. Its architecture comprises three coordinated stages: First, the Entity-level Clinical Diagnostic (ECD) module parses both the predicted report and the ground-truth report into clinical entities and assertion states, forming structured discrepancy categories that guide an LLM to rewrite a preferred report \(\overline{y}^{\text{GT}}\) that mirrors the linguistic style of the prediction while strictly enforcing ground-truth clinical facts. Second, a two-stage retrieval protocol identifies a candidate image \(\hat{x}^{\text{pre}}\) from the training database that is clinically identical to the erroneous prediction \(y^{\text{pre}}\), enabling a multi-modal preference inversion objective (M2DPO) across the preference quadruplet \((x, \hat{x}^{\text{pre}}, \overline{y}^{\text{GT}}, y^{\text{pre}})\). Third, token-level entropy is evaluated to identify correct yet highly uncertain entities, which are modified via knowledge-graph-constrained counterfactual substitution to yield \(\overline{y}^{\text{unc}}\) and paired with retrieved image \(\hat{x}^{\text{unc}}\) to mitigate latent clinical risks.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image x and Model Prediction y_pre"] --> B["ECD Entity Diagnosis & Linguistically-Aligned Generation<br/>Extract entities and isolate purely clinical preference pairs"]
    B --> C["Retrieval-Augmented M2DPO Multi-Modal Alignment<br/>Two-stage retrieval matches images to enforce preference inversion"]
    A --> D["Uncertain Entity Recognition & Counterfactual Latent Risk Mitigation<br/>Compute token entropy and substitute entities under RadGraph constraints"]
    D --> C
    C --> E["Dynamic Masking Joint Preference Optimization<br/>Filter invalid samples and train end-to-end multi-modal DPO"]

Key Designs

1. ECD-Assisted Linguistically-Aligned Preference Construction: Isolating Clinical Facts from Linguistic Noise

Standard DPO pairs arbitrary reference reports \(y^{\text{GT}}\) against predictions \(y^{\text{pre}}\), where disparities in sentence structures and vocabulary dilute token-level loss gradients away from clinical errors. To resolve this, the Entity-level Clinical Diagnostic (ECD) module utilizes RaTE-NER to parse both \(y^{\text{pre}}\) and \(y^{\text{GT}}\), extracting entity text spans \(e\), semantic embeddings \(E\), and assertion states \(A \in \{\text{Present}, \text{Absent}\}\). Entity correspondence is formulated as a threshold-constrained linear assignment problem, optimizing the global cosine distance \(\text{cost}_{i,j} = 1 - \cos(E^{\text{pre}}_i, E^{\text{GT}}_j)\) subject to \(\text{cost}_{i,j} < 1 - \tau\). Solving the assignment matrix partitions all extracted entities into four clinically disjoint subsets: - Correct Matches (\(P_{\text{CM}}\)): Matched pairs sharing identical assertion states (\(A_i = A_j\)); - Extraneous Entities (\(P_{\text{EE}}\)): Unmatched predicted entities representing hallucinations; - Missing Entities (\(P_{\text{ME}}\)): Unmatched ground-truth entities denoting omitted findings; - False Assertions (\(P_{\text{FA}}\)): Matched anatomical entities with conflicting diagnostic assertions (e.g., misclassifying negative pleural effusion as positive).

These diagnostic findings are assembled into a structured prompt directing GPT-4o to generate \(\overline{y}^{\text{GT}}\). This generated report strictly preserves the sentence structure, transitional phrasing, and stylistic tone of \(y^{\text{pre}}\) while rectifying only the factual discrepancies identified in \(P_{\text{EE}}\), \(P_{\text{ME}}\), and \(P_{\text{FA}}\). Training with \((\overline{y}^{\text{GT}}, y^{\text{pre}})\) confines optimization entirely to clinical deviations.

2. Retrieval-Augmented M2DPO: Visual Context-Triggered Preference Inversion

Optimizing textual preferences over a single static visual input fails to incentivize cross-modal grounding, allowing language priors to dominate visual perception. To achieve fine-grained vision-language grounding, M2DPO establishes an inversion mechanism: given two factual image-report pairs \((x_1, y_1)\) and \((x_2, y_2)\), the policy must prefer \(y_1\) over \(y_2\) under the visual context of \(x_1\), yet invert its preference to favor \(y_2\) over \(y_1\) when switched to visual context \(x_2\).

To form the preference quadruplet \((x, \hat{x}^{\text{pre}}, \overline{y}^{\text{GT}}, y^{\text{pre}})\), a two-stage retrieval protocol locates an appropriate training image \(\hat{x}^{\text{pre}}\): 1. ECD-Driven Textual Filtering: Candidate database reports are matched against \(y^{\text{pre}}\) using ECD; only cases with valid correct matches (\(|P_{\text{CM}}| > 0\)) and completely empty error subsets (\(P_{\text{EE}} \cup P_{\text{ME}} \cup P_{\text{FA}} = \emptyset\)) are retained, ensuring strict clinical congruence with the prediction; 2. MedKLIP Spatial Visual Retrieval: Within the candidate pool, spatial image embeddings are extracted via MedKLIP, selecting the image with maximal cosine similarity to query image \(x\). This preserves lesion fidelity while minimizing non-pathological shifts (such as patient anatomy or device settings).

The M2DPO loss integrates two reciprocal DPO objectives: $\(\mathcal{L}_{\text{M}^2\text{DPO}}(x_1, x_2, y_1, y_2) = \mathcal{L}_{\text{DPO}}(x_1, y_1, y_2) + \mathcal{L}_{\text{DPO}}(x_2, y_2, y_1)\)$ This dual supervision explicitly forces the model to ground diagnostic claims in corresponding localized visual patterns.

3. Counterfactual Latent Risk Mitigation: Eliminating High-Uncertainty Predictions

A critical oversight in existing MRG training is the neglect of latent risks: entities that happen to match the reference annotations but stem from highly uncertain, flat token distributions. DPO-Clin introduces an active uncertainty localization technique. For each correctly matched entity \(e \in P_{\text{CM}}\) comprising \(K\) tokens, its uncertainty is quantified by the peak token-level Shannon entropy across the decoding trajectory: $\(\mathcal{H}(\mathbf{e} \mid x) = \max_{k=1}^K \left[ -\sum_{w \in \mathcal{V}} \pi_{\text{ref}}(w \mid x, T_{<k}) \log \pi_{\text{ref}}(w \mid x, T_{<k}) \right]\)$ Entities exceeding a threshold \(\theta\) are flagged as uncertain.

To suppress these fragile boundaries, counterfactual report \(\overline{y}^{\text{unc}}\) is synthesized from \(\overline{y}^{\text{GT}}\) by replacing the uncertain entity with the second most probable token candidate from the model distribution. To safeguard clinical credibility, the substituted entity is validated via a RadGraph knowledge graph schema: substitution is retained only if the altered finding maintains valid structural edges with neighboring entities in the sentence. Subsequently, the two-stage retrieval identifies an aligned visual counterpart \(\hat{x}^{\text{unc}}\), establishing the latent risk quadruplet \((x, \hat{x}^{\text{unc}}, \overline{y}^{\text{GT}}, \overline{y}^{\text{unc}})\) to reinforce decision certainty.

Loss & Training

The overall framework is fine-tuned using LoRA on all linear layer projections with temperature parameter \(\beta = 0.1\). To prevent noisy gradient updates caused by retrieval misses or invalid graph edges, two dynamic binary indicators \(\mathbb{1}_{\text{EE}}\) and \(\mathbb{1}_{\text{LR}}\) filter the training stream: \(\mathbb{1}_{\text{EE}} = 0\) if \(y^{\text{pre}}\) contains zero factual errors or if the first-stage candidate retrieval pool is empty; \(\mathbb{1}_{\text{LR}} = 0\) if the counterfactual candidate violates RadGraph connectivity constraints. The full objective is formulated as: $\(\mathcal{L}_{\text{total}} = \mathbb{1}_{\text{EE}} \mathcal{L}_{\text{M}^2\text{DPO}}(x, \hat{x}^{\text{pre}}, \overline{y}^{\text{GT}}, y^{\text{pre}}) + \mathbb{1}_{\text{LR}} \mathcal{L}_{\text{M}^2\text{DPO}}(x, \hat{x}^{\text{unc}}, \overline{y}^{\text{GT}}, \overline{y}^{\text{unc}})\)$ The pipeline is trained for 5 epochs on MIMIC-CXR (learning rate \(1 \times 10^{-5}\)) and 20 epochs on the endoscopy dataset (learning rate \(2 \times 10^{-5}\)) using AdamW, setting thresholds \(\tau = 0.4\) and \(\theta = 0.5\) for chest X-ray experiments.

Key Experimental Results

Main Results

The framework was evaluated on the in-domain benchmark MIMIC-CXR, the zero-shot cross-domain dataset IU X-Ray, and an in-house digestive endoscopy dataset across canonical (R2GenGPT) and expert-augmented (RADAR) backbones, benchmarked against generic VLM alignment (SIMA) and medical preference baselines (MMedPO, RRG-DPO).

Dataset Backbone & Method BLEU-1 ↑ BLEU-4 ↑ ROUGE-L ↑ 14Ma-F1 ↑ 14Mi-F1 ↑ RadGraph ↑ RadCliQ ↓ RaTEScore ↑
MIMIC-CXR R2GenGPT (SFT) 0.411 0.134 0.297 0.389 0.504 0.260 2.74 0.453
(In-domain Test) + SIMA 0.416 0.140 0.306 0.396 0.515 0.265 2.74 0.450
+ MMedPO 0.426 0.144 0.315 0.420 0.530 0.281 2.71 0.476
+ RRG-DPO 0.417 0.138 0.302 0.415 0.522 0.275 2.71 0.470
+ DPO-Clin (Ours) 0.423 0.140 0.321 0.437 0.545 0.297 2.67 0.494
RADAR (SFT) 0.509 0.262 0.397 0.460 0.627 0.346 2.61 0.531
+ SIMA 0.512 0.264 0.405 0.464 0.620 0.341 2.57 0.537
+ MMedPO 0.516 0.268 0.410 0.478 0.639 0.356 2.52 0.546
+ RRG-DPO 0.511 0.260 0.408 0.473 0.639 0.356 2.54 0.550
+ DPO-Clin (Ours) 0.522 0.271 0.419 0.490 0.654 0.372 2.48 0.567
IU X-Ray R2GenGPT (SFT) 0.356 0.103 0.270 0.303 0.446 0.204 2.86 0.405
(Cross-domain Eval) + DPO-Clin (Ours) 0.365 0.112 0.292 0.350 0.492 0.252 2.74 0.452
RADAR (SFT) 0.364 0.116 0.276 0.325 0.546 0.237 2.78 0.428
+ DPO-Clin (Ours) 0.377 0.126 0.298 0.362 0.580 0.275 2.66 0.461

On the endoscopic report generation benchmark, DPO-Clin achieves a binary clinical diagnosis score of 2F1 = 0.852 over the R2GenGPT backbone, significantly surpassing the SFT baseline (0.798) and outperforming the strongest competitor MMedPO (0.834).

Ablation Study

The table below reports ablation experiments on MIMIC-CXR using RADAR, investigating mitigating explicit errors (MEE), mitigating latent risks (MLR), ECD guidance, RadGraph verification, dynamic masking, and linguistically-aligned report synthesis:

Config / Variant BLEU-4 ROUGE-L 14Ma-F1 14Mi-F1 RadGraph RadCliQ ↓ RaTEScore Note
RADAR Baseline 0.262 0.397 0.460 0.627 0.346 2.61 0.531 Unaligned SFT baseline
+ MEE (w/o ECD) 0.256 0.390 0.452 0.615 0.337 2.64 0.518 Unconstrained LLM rewriting degrades accuracy
+ MEE 0.268 0.414 0.482 0.645 0.366 2.51 0.557 Explicit error mitigation with full ECD
+ MLR (w/o Verification) 0.265 0.406 0.472 0.635 0.354 2.55 0.546 Counterfactual without RadGraph consistency
+ MLR 0.266 0.409 0.475 0.640 0.360 2.55 0.550 Latent risk mitigation with verified graph edges
+ MEE & MLR (w/o Masking) 0.262 0.403 0.466 0.638 0.353 2.58 0.541 Unmasked invalid or empty retrieval pairs
+ MEE & MLR (\(y^{\text{GT}} \to y^{\text{GT}}\)) 0.273 0.417 0.482 0.642 0.360 2.52 0.552 Raw GT reports inflate n-grams but harm clinical F1
+ MEE & MLR (DPO-Clin Full) 0.271 0.419 0.490 0.654 0.372 2.48 0.567 Optimal clinical and RRG-specific alignment

Key Findings

  1. Linguistic Alignment Prevents Superficial Imitation: Substituting generated \(\overline{y}^{\text{GT}}\) with raw ground-truth reports \(y^{\text{GT}}\) slightly raises BLEU-4 (0.271 \(\to\) 0.273) but causes noticeable drops in 14Ma-F1 (0.490 \(\to\) 0.482) and RaTEScore (0.567 \(\to\) 0.552), proving that standard DPO learns superficial phrase matching at the cost of diagnostic precision.
  2. ECD is Imperative for Clinical Truthfulness: Removing ECD and relying purely on raw GPT-4o prompting leads to performance inferior even to the SFT baseline (14Ma-F1 dropping to 0.452), underscoring that general LLMs lack nuanced medical entity boundary awareness without diagnostic anchoring.
  3. M2DPO Establishes Genuine Lesion Grounding: Cross-modal attention maps demonstrate that standard DPO and SimPO suffer from dispersed attention over irrelevant anatomical structures, producing omission errors; M2DPO sharpens attention specifically around lesion centers (e.g., subtle polyps and pneumothorax margins).
  4. Substantial Shift Toward Certainty: Latent risk mitigation shifts the entropy distribution of correct predictions markedly leftward, reducing mean uncertainty from 0.514 to 0.282 and median score from 0.316 to 0.105 while expanding total correct predictions.

Highlights & Insights

  • Linguistic-Clinical Decoupling: Rather than comparing whole unaligned documents, ECD isolates entity-level discrepancies, allowing GPT-4o to rewrite facts while preserving prediction phrasing to remove stylistic confounders.
  • Context-Triggered M2DPO Inversion: By coupling dual-stage retrieval with cross-modal preference inversion, M2DPO enforces grounding directly through visual context switching, overcoming the text-dominant bias of prior methods.
  • Active Latent Risk Defense: By identifying highly uncertain but technically correct predictions via token entropy and evaluating counterfactual substitutions under RadGraph verification, DPO-Clin actively refines decision boundaries against silent clinical failures.

Limitations & Future Work

  • Dependency on Multi-Stage Pipelines: The offline curation requires several external tools (RaTE-NER, GPT-4o, MedKLIP, RadGraph), introducing computational latency and proprietary LLM costs.
  • Retrieval Bottlenecks in Long-Tail Pathologies: Demanding zero-error candidate reports (\(P_{\text{EE}} \cup P_{\text{ME}} \cup P_{\text{FA}} = \emptyset\)) can exhaust candidate pools for rare diseases, triggering the dynamic mask (\(\mathbb{1}_{\text{EE}} = 0\)) and attenuating preference learning on long-tail cases.
  • Future Directions: Exploring unified, self-contained architectures that perform entity-level self-diagnosis and counterfactual generation online without multi-step retrieval dependencies.
  • vs RRG-DPO / MMedPO: Existing medical DPO methods employ scalar clinical weighting or database retrieval over raw report pairs, failing to isolate stylistic interference or establish contrastive multi-modal visual inversion. DPO-Clin achieves fine-grained linguistic decoupling alongside bidirectional visual grounding.
  • vs SIMA / SimPO: General multi-modal alignment focuses on general instruction-following without medical entity validation or factual consistency guarantees; DPO-Clin leverages RadGraph validation and uncertainty minimization tailored to high-stakes healthcare scenarios.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant integration of entity-level clinical diagnosis, context-triggered multi-modal preference inversion, and counterfactual uncertainty mitigation.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarking over two architectures, three diverse datasets across CXR and endoscopy, complemented by detailed ablations and attention maps.
  • Writing Quality: ⭐⭐⭐ harmonization of concepts, mathematically sound formulations, and lucid diagrams.
  • Value: ⭐⭐⭐⭐⭐ Sets a rigorous methodological benchmark for transitioning medical VLMs from surface-level imitation to verifiable clinical reliability.