MedRepBench: Benchmarking Structured Understanding of Medical Report Images¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: Medical Imaging
Keywords: Medical Report Understanding, Structured Extraction, Vision-Language Models, Benchmarking, RL Alignment
TL;DR¶
MedRepBench introduces an end-to-end benchmark of 1,925 de-identified real-world Chinese medical report images, establishes dual objective (field-level recall) and subjective (LLM judge) evaluation protocols, and demonstrates that a lightweight GRPO reinforcement learning baseline using average field recall as a reward boosts end-to-end VLM recall by 6.14%.
Background & Motivation¶
Automated interpretation of medical laboratory and examination reports serves as a vital cornerstone for both intelligent clinical workflows and patient self-management. In everyday outpatient consultations, chronic disease management, and telemedicine triage, patients regularly present physical report sheets photographed with handheld smartphones, mobile app screenshots, or exported digital PDFs. Practical AI solutions must not only translate dense clinical terminology into accessible layperson-facing explanations, but also extract fine-grained structured records (test names, quantitative measurements, units, reference intervals, and abnormality indicators) to drive automated health records and inter-hospital data exchange.
However, existing approaches face a notable benchmark gap and architectural trade-offs in real-world report understanding. Conventional OCR+LLM multi-stage pipelines are widely deployed, yet they require heavy template engineering, discard rich 2D spatial layouts, and suffer from irreversible error propagation when OCR missegments text or confuses adjacent table columns. Conversely, modern vision-language models (VLMs) natively perceive visual layout and typography, but current multimodal medical benchmarks primarily target narrative radiology report generation (e.g., chest X-rays) or synthetic VQA datasets, leaving the end-to-end structured understanding of noisy, heterogeneous, real-world document images systematically under-evaluated.
Consequently, assessing whether modern VLMs can reliably extract and interpret structured medical findings directly from raw pixels is an urgent research priority. This paper bridges this gap by collecting a diverse, multi-departmental corpus of authentic Chinese medical reports across distinct capture modes and layout styles, rigorously analyzing the gap between OCR-assisted and end-to-end vision pipelines. The core idea is to establish MedRepBench, comprising 1,925 curated medical report images with a dual objective field-level recall and LLM-as-a-judge evaluation protocol, demonstrating that non-differentiable field recall directly drives lightweight GRPO reinforcement learning to boost mid-sized VLM structured extraction.
Method¶
Overall Architecture¶
MedRepBench establishes an end-to-end system encompassing real-world clinical data governance, dual-track fine-grained evaluation protocols, and downstream policy alignment via reinforcement learning. The pipeline accepts unconstrained, non-standardized medical report images (handheld photos, mobile screenshots, and digital PDFs), passing them through heuristic filtering, perceptual deduplication, and multi-tier privacy anonymization to form a trustworthy evaluation benchmark. The benchmark separates itemized laboratory reports (subjected to objective field-level recall) from narrative examination reports (evaluated via reference-grounded LLM judgment). Finally, the benchmark's objective recall metric is harnessed as a scalar reward to guide GRPO policy optimization for end-to-end vision-language models.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Real-World Medical Report Images<br/>Photos / Screenshots / e-Docs"] --> B["Multi-Source Data Curation & Privacy Filtering<br/>Heuristic quality gates & ROI blackout"]
B --> C["Dual Evaluation Benchmark"]
C --> D["Objective Field-Level Recall Protocol<br/>Name / Value / Unit / Range / Abnormality"]
C --> E["LLM-as-a-Judge Subjective Protocol<br/>Factuality / Interpretability / Reasoning"]
D --> F["GRPO Reinforcement Learning Alignment<br/>Optimize VLM policy via average field recall"]
E --> G["Patient-Facing Clarification & Clinical Utility"]
F --> G
Key Designs¶
1. Multi-Source Data Curation & Privacy Filtering: Robust Data Governance and Golden Schema Construction Authentic medical reports exhibit extreme cross-institutional layout variance, erratic lighting, and sensitive personal health information. Starting from 8,000 candidate reports from a real-world online consultation service, the authors enforce a multi-stage filtering regimen. Low-quality images with short-edge resolution under 800 px, severe motion blur, or cut-off margins are pruned, with quality thresholds audited on 500 random samples; perceptual hashing (pHash) paired with OCR text similarity subsequently removes 12.3% of near-duplicate images. To create golden annotations, a specialized medical OCR engine coupled with DeepSeek-R1 extracts structured JSON under constrained schemas, followed by unit canonicalization and reference range regex validation. A manual audit of 200 random instances revealed over 97% field-level agreement, with residual misalignments rectified prior to release. For privacy, an automated OCR pipeline detects personally identifiable information (PII) including names, telephone numbers, and addresses, applying pixel-level ROI blackouts, followed by 100% manual review across all 1,925 images.
2. Objective Field-Level Recall Protocol: Five-Attribute Alignment with Top-K Budget Truncation Conventional text generation metrics (BLEU, ROUGE, CIDEr) only quantify surface n-gram overlap and fail to measure clinical field accuracy. MedRepBench models structured laboratory reports as sets of five-attribute tuples: test name, measured value, measurement unit, reference range, and abnormality flag. Abnormality identification requires reasoning over numeric values against intervals or categorizing qualitative readings (positive/negative), preventing pure rote copying. In evaluation, a generated item must strictly match the ground-truth item's normalized name before the remaining four fields are scored. Crucially, to prevent models from artificially inflating recall through unbounded over-generation or hallucinated item dumping, the protocol enforces Top-K budget truncation, setting \(K\) to the exact number of ground-truth items in that report. Spurious predictions consume finite slots and displace correct items, effectively incorporating precision penalties into a single metric. Field-level recall and overall macro-average recall are defined as: $\(Recall_f = \frac{\# \text{ correctly extracted field } f}{\# \text{ ground truth field } f}\)$ $\(\mathrm{AvgRecall} = \frac{1}{|F|} \sum_{f \in F} Recall_f, \quad |F|=5\)$
3. LLM-as-a-Judge Subjective Protocol: Tri-Dimensional Usability Scoring with Expert Calibration Narrative clinical reports (such as ultrasound, electrocardiogram, and endoscopy records) feature free-form descriptive text where rigid five-field extraction yields brittle pseudo-labels. MedRepBench addresses these through an automated patient-facing evaluation protocol judged by DeepSeek-R1. The evaluator is provided with OCR-derived ground truth as a factual anchor while scoring candidate interpretations generated by VLMs in the native no-OCR visual setting. The evaluation assesses Factuality (no fabricated findings or inverted qualitative results), Interpretability (lay-friendly clarity without unauthorized diagnostic or therapeutic advice), and Reasoning Quality (logically derived abnormalities without internal contradictions). The judge issues discrete scores of 0, 1, or 2, from which Acceptability Rate (score \(\ge 1\)) and Excellence Rate (score \(= 2\)) are derived. In a calibration study across 60 sampled cases independently reviewed by 3 licensed clinicians with majority voting, the LLM judge achieved 88.3% agreement and a Cohen's \(\kappa\) of 0.82, validating the metric's reliability.
4. GRPO Policy Optimization: Direct End-to-End Multimodal Alignment via Non-Differentiable Recall Reward To tackle the substantial 10%–20% performance deficit of end-to-end VLMs compared to OCR-assisted pipelines, the authors design an alignment baseline utilizing Group Relative Policy Optimization (GRPO). Starting from an InternVL3-8B base policy supervised on 15k medical interpretation instruction samples (Ours-SFT), the average field-level recall \(\mathrm{AvgRecall}\) between predicted and reference JSON structures is directly designated as the non-differentiable scalar reward \(R\): $\(R = \frac{1}{|F|} \sum_{f \in F} Recall_f\)$ Across sampled generation candidates per image prompt, GRPO computes relative group advantage scores and updates model weights via a clipped surrogate objective with KL regularization against the SFT policy. This formulation allows the multimodal policy to autonomously discover visual alignment patterns across table columns and numeric rows purely from task-level reward signals, without requiring expensive token-level fine-grained spatial bounding box annotations.
Loss & Training¶
The GRPO alignment baseline adopts parameter-efficient fine-tuning to preserve foundational visual-language representations. The vision encoder and cross-modal projector of InternVL3-8B are entirely frozen; LoRA adapters (rank = 128) are inserted into all query, key, and value attention projection layers within the language backbone. Training proceeds for 2 epochs over a curated split of 800 structured laboratory reports matching the benchmark schema. The optimization objective minimizes the clipped policy gradient loss with a KL divergence penalty: $\(\mathcal{L}_{GRPO}(\theta) = -\mathbb{E} \left[ \min \left( \frac{\pi_\theta(y|x)}{\pi_{old}(y|x)} \hat{A}_i, \, \text{clip}\left(\frac{\pi_\theta(y|x)}{\pi_{old}(y|x)}, 1-\epsilon, 1+\epsilon\right) \hat{A}_i \right) - \beta \, \mathbb{D}_{KL}(\pi_\theta \parallel \pi_{ref}) \right]\)$ where the advantage \(\hat{A}_i\) is computed by standardizing rewards \(R\) across grouped rollouts. This setup delivers rapid, sample-efficient alignment directly targeting structured clinical accuracy.
Key Experimental Results¶
Main Results¶
The primary evaluation investigates representative open-source VLMs on laboratory reports under both the end-to-end visual (no-OCR) and the text-assisted (OCR-asst.) settings across all five target attributes.
| Model | Inference Setting | \(R_{name}\) (%) | \(R_{value}\) (%) | \(R_{unit}\) (%) | \(R_{range}\) (%) | \(R_{abnormal}\) (%) |
|---|---|---|---|---|---|---|
| InternVL2.5-8B | no-OCR | 71.38 | 55.71 | 60.63 | 58.88 | 45.06 |
| InternVL2.5-8B | OCR-asst. | 85.64 | 70.14 | 77.85 | 79.40 | 58.59 |
| InternVL2.5-38B | no-OCR | 84.83 | 68.85 | 74.93 | 72.37 | 66.01 |
| InternVL2.5-38B | OCR-asst. | 95.20 | 85.02 | 89.00 | 89.44 | 81.52 |
| InternVL3-8B | no-OCR | 88.21 | 71.29 | 75.56 | 70.63 | 67.92 |
| InternVL3-8B | OCR-asst. | 95.25 | 84.87 | 88.00 | 88.80 | 77.97 |
| InternVL3-14B | no-OCR | 88.70 | 67.69 | 74.63 | 67.96 | 70.12 |
| InternVL3-14B | OCR-asst. | 95.25 | 86.53 | 89.52 | 88.81 | 83.56 |
| InternVL3-38B | no-OCR | 88.72 | 64.83 | 75.63 | 69.36 | 66.74 |
| InternVL3-38B | OCR-asst. | 94.69 | 85.64 | 88.79 | 89.88 | 77.55 |
| Qwen2.5-VL-7B | no-OCR | 89.90 | 73.17 | 77.31 | 78.33 | 68.36 |
| Qwen2.5-VL-7B | OCR-asst. | 93.94 | 82.46 | 81.61 | 85.07 | 76.39 |
| Qwen2.5-VL-32B | no-OCR | 90.13 | 76.32 | 79.30 | 81.61 | 59.66 |
| Qwen2.5-VL-32B | OCR-asst. | 95.89 | 84.72 | 87.71 | 89.15 | 83.95 |
| LLaMA-4-Scout | no-OCR | 87.90 | 65.69 | 74.29 | 67.87 | 63.24 |
| LLaMA-4-Scout | OCR-asst. | 92.68 | 78.86 | 82.94 | 77.91 | 69.23 |
| LLaMA-4-Maverick | no-OCR | 88.51 | 72.26 | 78.83 | 77.58 | 75.83 |
| LLaMA-4-Maverick | OCR-asst. | 96.68 | 87.35 | 90.23 | 90.74 | 86.56 |
In text-only LLM benchmarks under OCR assistance, DeepSeek-V3 attained 96.48% \(R_{name}\) and 88.83% \(R_{value}\), establishing the upper bound among open-weight architectures; Qwen3-235B-A22B-2507 achieved competitive extraction marks (\(R_{name}\) 96.42%, \(R_{range}\) 92.48%).
Ablation Study¶
The ablation analysis assesses the effect of SFT and GRPO reinforcement learning alignment on InternVL3-8B against baselines across end-to-end and OCR-assisted modes, reporting objective Average Recall alongside subjective Acceptability and Excellence.
| Model | OCR Input | Param | Avg. Recall (%) | Accept. (%) | Excel. (%) |
|---|---|---|---|---|---|
| InternVL2.5-8B | No | 8B | 58.33 | 33.67 | 17.30 |
| Qwen2.5-VL-7B | No | 7B | 77.41 | 26.21 | 5.58 |
| InternVL3-8B | No | 8B | 74.72 | 48.24 | 31.83 |
| Ours-SFT | No | 8B | 73.31 | 56.27 | 38.86 |
| Ours-GRPO | No | 8B | 79.45 | 60.64 | 42.45 |
| LLaMA-4-Maverick | No | 17B | 78.60 | 60.78 | 42.74 |
| Qwen2.5-VL-32B | No | 32B | 77.41 | 67.83 | 51.40 |
| InternVL3-38B | No | 38B | 73.06 | 51.15 | 32.70 |
| InternVL2.5-8B | Yes | 8B | 74.32 | - | - |
| Qwen2.5-VL-7B | Yes | 7B | 83.89 | - | - |
| Qwen3-8B | Yes | 8B | 83.90 | - | - |
| InternVL3-8B | Yes | 8B | 86.98 | - | - |
| Ours-SFT | Yes | 8B | 82.37 | - | - |
| Ours-GRPO | Yes | 8B | 87.41 | - | - |
Key Findings¶
- Substantial Multimodal Perception Disparity: In the no-OCR setting, open-source VLMs suffer an average 10%–20% recall drop compared to their OCR-assisted configurations. The drop is most severe on numeric values and reference ranges, demonstrating that current general vision encoders still struggle with dense, tiny printed alphanumeric tokens.
- Three Core End-to-End Failure Modes: Qualitative error analysis indicates failures stem from: (1) character-level substitutions on small font numbers and medical unit symbols; (2) layout misbinding across multi-column or nested report grids; and (3) missed abnormality indicators when visual flags (e.g., arrows) are detached from numeric values.
- GRPO Reinforcement Learning Scales Efficiently: Using only 800 training reports with LoRA fine-tuning, Ours-GRPO elevates end-to-end recall from 73.31% (SFT) to 79.45% (+6.14%), outperforming larger models like LLaMA-4-Maverick (78.60%) and Qwen2.5-VL-32B (77.41%), confirming that benchmark-grounded RL alignment unlocks dense structural extraction.
Highlights & Insights¶
- Decoupled Evaluation with Budget Truncation: Segregating structured laboratory reports from narrative exam descriptions prevents pseudo-label noise, while the Top-K matching budget inherently penalizes hallucinated and over-generated fields without needing complex heuristic penalty weights.
- Standardized Clinical Data Curation: A complete pipeline incorporating heuristic resolution filtering, pHash deduplication, schema-constrained parsing, expert clinician verification, and multi-stage PII blackout sets a rigorous standard for real-world medical image benchmarks.
- Direct RL Optimization from Discrete Metrics: Packaging non-differentiable structured JSON match recall directly as a GRPO reward demonstrates a clean, sample-efficient paradigm for aligning multimodal models on structured document extraction without requiring dense token-level localization loss.
Limitations & Future Work¶
- Language and Institutional Specificity: MedRepBench comprises Chinese clinical reports exclusively; formatting rules, unit conventions, and abbreviations reflect regional healthcare standards, and validation on multi-lingual or Western clinical records remains future work.
- Absence of Proprietary Frontier Models: In accordance with medical privacy standards and open reproducibility, proprietary commercial APIs (such as GPT-4o or Gemini 1.5 Pro) were not evaluated.
- Extension to Medical-Domain VLMs: The current study primarily benchmarks general-purpose foundation VLMs; extending evaluations to specialized medical multimodal architectures and multimodal document agents presents a natural next step.
Related Work & Insights¶
- vs. Medical Report Generation (Argus, MedVAG): Prior literature focuses predominantly on generating free-text radiology narratives from clinical images (X-rays, CTs); MedRepBench addresses the inverse interpretive task of extracting discrete structured clinical values and providing lay-friendly explanations from non-standard document photos.
- vs. General Document Benchmarks (MMDocBench, BenchX): General document benchmarks concentrate on invoices, receipts, and general academic papers, missing clinical reference interval comparisons and patient communication boundaries; MedRepBench fills this gap with real-world medical documents.
- vs. Traditional OCR+LLM Pipelines: While OCR-assisted pipelines remain an accurate upper bound when text recognition succeeds, they incur high engineering maintenance overhead and fail under severe layout skew; MedRepBench establishes that end-to-end VLMs, when RL-aligned, provide a viable, unified alternative.
Rating¶
- Novelty: ⭐⭐⭐⭐ [Introduces the first end-to-end benchmark for structured understanding of real-world noisy medical report images with dual evaluation protocols]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive comparisons across 9 VLMs and 12 LLMs in dual OCR/no-OCR settings, validated by clinician blind review and RL alignment]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical definitions of evaluation metrics, well-structured prose, and comprehensive error taxonomy]
- Value: ⭐⭐⭐⭐⭐ [Releases 1,925 de-identified real clinical report images and evaluation scripts, providing an essential diagnostic benchmark for text-token-free multimodal models]