Mimicking Radiologists: A Coarse-to-Fine Framework with Structural Sparse Tokens for Dual-LLM Computed Tomography Report Generation¶
Conference: ECCV 2026
Paper: CVF Open Access
Code: https://github.com/ccarliu/CTRG
Area: Medical Imaging
Keywords: CT report generation, dual-LLM framework, sparse token filtering, attention sink elimination, reinforcement learning alignment
TL;DR¶
This paper proposes a coarse-to-fine dual-LLM framework mimicking the visual search workflow of radiologists for 3D computed tomography report generation, combining structural global token screening, abnormality-prompted multi-source local token filtering, and GRPO collaborative alignment to compress visual tokens by 96.1% while boosting clinical efficacy (CE-F1) by 32.0%.
Background & Motivation¶
Three-dimensional computed tomography (CT) plays an indispensable role in clinical diagnosis and treatment planning. However, reading volumetric CT scans requires intensive visual search across hundreds of high-resolution slices to inspect subtle, localized lesions. In clinical practice under heavy reporting burdens and radiologist shortages, this manual workflow is time-consuming and prone to fatigue, leading to missed findings and delayed diagnoses. While multimodal large language models (MLLMs) offer strong potential for automated CT report generation (CTRG), they face a fundamental physical challenge: volumetric CT data contain immense spatial dimensions and severe information redundancy, whereas pathologically critical lesions are exceptionally sparse and confined to localized anatomical subvolumes. Early methods connected 3D spatial pooling encoders to LLMs, but suffered catastrophic loss of localized pathology. Recent region-guided approaches employ organ segmentation masks to crop subvolumes, yet completely overlook substantial intra-region redundancy. Furthermore, selecting localized patch tokens purely based on cross-attention weights is severely undermined by "attention sinks"—where low-semantic boundary or background tokens absorb disproportionate attention mass, crowding out fine-grained lesion evidence.
These computational bottlenecks contrast sharply with how expert radiologists interpret scans. Radiologists never visually ingest an entire volumetric stack at uniform density, nor do they rely on isolated local intuition; rather, they perform a disciplined coarse-to-fine visual search. They first scan across the entire volumetric series guided by clinical metadata (such as chief complaint and medical history) to establish a coarse shortlist of suspicious abnormalities across anatomical compartments. Subsequently, they scrutinize flagged regions in detail to verify positive findings and rule out negatives, before composing an authoritative structured report. Single-stage or single-LLM frameworks that force global context comprehension and granular lesion reporting into a single representation space either choke under computational token limits or fail to capture localized pathologies.
This paper addresses this challenge by decomposing the radiologist's coarse-to-fine visual search into a collaborative dual-agent architecture: an abnormality-proposal LLM first predicts candidate abnormalities from structural global tokens, prompting a multi-source local token filter to extract high-yield patch tokens, which are then passed alongside global context to a dedicated report-generation LLM. Core idea: emulate the radiologist's coarse-to-fine visual search via a dual-LLM paradigm, using mask-guided sparse negative-entropy loss for structure-wise feature decoupling, abnormality-prompted local token filtering to counter attention sinks, and GRPO reinforcement learning to bridge the inter-LLM proposal gap.
Method¶
Overall Architecture¶
The framework directly mirrors the visual search and reporting workflow practiced by radiologists, consisting of four interdependent stages: structure-wise visual token extraction, coarse candidate abnormality proposal, abnormality-prompted local token filtering (AP-LTF), and fine-grained diagnostic report generation. Initially, a CT-ViT encodes the input 3D CT volume into dense patch tokens. A set of learnable structural queries extracts a global summary token and a candidate local patch pool for each predefined anatomical compartment via cross-attention, supervised by a mask-guided sparse negative-entropy loss to prevent cross-organ feature leakage. Next, an abnormality-proposal LLM (\(D_\text{pro}\)) takes only the single global token per structure alongside clinical metadata to generate a patient-specific shortlist of candidate abnormalities. A clinical BERT encoder transforms this shortlist into structural text queries, which guide AP-LTF in evaluating candidate local tokens by fusing semantic similarity, a learned projection score, and geometric attention weights, performing hard sparse filtering to isolate the most discriminative patch tokens. Finally, a report-generation LLM (\(D_\text{ful}\)) integrates the global tokens, filtered local tokens, and clinical metadata to produce a complete radiology report. To resolve exposure bias between ground-truth and predicted abnormality inputs, group relative policy optimization (GRPO) refines \(D_\text{pro}\) with a composite reward balancing clinical efficacy and linguistic fidelity.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["3D CT Volume + Clinical Metadata"] --> B["Mask-guided sparse negative-entropy decoupling<br/>CT-ViT extracts structural global & local token pools"]
B --> C["Coarse-to-fine dual-LLM collaborative generation<br/>Abnormality-proposal LLM predicts candidate shortlist"]
C --> D["Abnormality-prompted local token filtering<br/>Fuses similarity, linear projection, and attention weight"]
D --> E["Coarse-to-fine dual-LLM collaborative generation<br/>Report-generation LLM decodes complete report"]
E --> F["GRPO collaborative reinforcement alignment<br/>Optimizes proposal policy via CE-F1 and BLEU-4 rewards"]
Key Designs¶
1. Mask-guided sparse negative-entropy decoupling: suppressing inter-organ attention leakage
Standard cross-modal contrastive learning provides only indirect supervision from unstructured report text, which frequently causes attention distributions to bleed into adjacent anatomical organs. To enforce strict anatomical boundaries without incurring inference latency, the framework generates offline 3D masks for 10 anatomical regions using the Segment Anything in Medical Images (SAT) foundation model. A patch is defined as belonging to a structure if at least 30% of its volume intersects the mask. A mask-guided sparse negative-entropy loss penalizes outlier tokens located outside the anatomical boundary that nevertheless exhibit high attention weights: $\(\mathcal{L}_\text{sparse} = \frac{1}{|\mathcal{H}|} \sum_{(i,j)\in\mathcal{H}} -\log(1 - \mathbf{A}^v_{i,j})\)$ where \(\mathcal{H} = \{(i,j) \mid \mathbf{M}_{i,j} = 0 \land \mathbf{A}^v_{i,j} > \eta\}\) isolates tokens that fall outside structure \(i\) yet exceed attention threshold \(\eta\). This constraint concentrates patch representations strictly within the designated organ geometry during pretraining, requiring no segmentation models during clinical inference.
2. Abnormality-prompted local token filtering: eliminating attention sinks
In multi-head attention networks, background and boundary tokens frequently act as "attention sinks," accumulating high attention weights despite carrying minimal pathological semantics. Selecting tokens purely by top attention weights risks filling the context with uninformative background. The abnormality-prompted local token filter (AP-LTF) circumvents this vulnerability by evaluating each candidate token \(\mathbf{g}_{i,j}\) in an expanded pool of \(K=200\) tokens per structure against a multi-source scoring function: $\(s_{i,j} = \alpha_0 \operatorname{sim}(\mathbf{q}^t_i, \mathbf{g}_{i,j}) + \alpha_1 \operatorname{lin}(\operatorname{cat}(\mathbf{g}_{i,j}, \mathbf{q}^t_i)) + \alpha_2 \mathbf{A}_{i,j}^v\)$ where \(\mathbf{q}^t_i\) denotes the text query vector derived from the candidate abnormality shortlist via clinical BERT. This scoring balances semantic text similarity, learned cross-modal projection, and spatial attention prominence. Softmax-normalized probabilities over the \(K\) candidates drive the hard selection of the top \(K'=15\) tokens, while a straight-through estimator (STE) allows gradients to flow backwards through the discrete selection during end-to-end training with structural contrastive loss \(\mathcal{L}_\text{filt}\).
3. Coarse-to-fine dual-LLM collaborative generation: decoupling discovery from description
Tasking a single model with simultaneously localizing lesions and drafting extensive diagnostic prose strains model capacity and provokes hallucination. The framework decouples these roles across two specialized Llama-3.2-3B models. The proposal model \(D_\text{pro}\) handles anomaly detection, operating over concise inputs (a single global visual token \(\mathbf{s}^v_i\) per structure and clinical text) to output a concise abnormality shortlist averaging just 45 tokens per volume. The report-generation model \(D_\text{ful}\) handles synthesis, receiving the global tokens, filtered local tokens \(\{\mathcal{G}'_i\}\), and clinical metadata. During supervised fine-tuning, ground-truth-parsed abnormalities are prefixed to the target report, grounding \(D_\text{ful}\) in clinical findings and negations. This reduces the initial 4,096 CT-ViT tokens to just 160 tokens (10 structures \(\times\) [1 global + 15 local]), achieving a 96.1% compression rate and enabling efficient training and inference on standard GPU hardware.
4. GRPO collaborative reinforcement alignment: closing distribution gaps and mitigating cascade errors
A significant exposure bias exists between training (where \(D_\text{ful}\) is conditioned on clean ground-truth abnormality prefixes) and inference (where it relies on imperfect proposals from \(D_\text{pro}\)). Furthermore, false negatives in \(D_\text{pro}\) risk propagating irrevocably downstream. To align the dual-LLM system, \(D_\text{ful}\) and AP-LTF are frozen while \(D_\text{pro}\) is treated as a proposal policy \(\pi_\theta\) optimized via group relative policy optimization (GRPO). For each scan, GRPO samples a group of candidate proposal outputs and optimizes the policy using relative group advantages without requiring a separate critic network, driven by a composite reward: $\(r = \lambda \cdot \text{CE-F1} + (1-\lambda) \cdot \text{BLEU-4}\)$ Crucially, this collaborative optimization enables the framework to overcome strictly unidirectional cascade errors: because AP-LTF retains geometric attention weights \(\mathbf{A}^v_{i,j}\) in its scoring mechanism, \(D_\text{ful}\) successfully recovers 27.2% of the abnormality findings initially omitted by \(D_\text{pro}\).
Loss & Training¶
Training proceeds across four systematic stages: 1. Stage 1 (Abnormality Proposal Pre-tuning): SFT of \(D_\text{pro}\) via next-token prediction on global visual tokens and clinical metadata; 2. Stage 2 (Vision-Text Alignment and Token Filter Training): Freezing \(D_\text{pro}\) to optimize CT-ViT, structural queries \(\mathbf{Q}_v\), and AP-LTF using the combined objective: $\(\mathcal{L}_\text{com} = \mathcal{L}_\text{so-pre} + \mathcal{L}_\text{sparse} + \mathcal{L}_\text{filt}\)$ 3. Stage 3 (Full Report SFT): Freezing visual encoders and AP-LTF to fine-tune \(D_\text{ful}\) on ground-truth abnormality-prefixed reports using cross-entropy loss \(\mathcal{L}_\text{rg}\); 4. Stage 4 (GRPO Collaborative Refinement): Freezing \(D_\text{ful}\) and AP-LTF while optimizing \(D_\text{pro}\) policy parameters via \(\mathcal{J}_\text{GRPO}\) based on relative advantages and KL-divergence regularization.
Key Experimental Results¶
Main Results¶
The model was evaluated against leading 2D and 3D medical report generation methods across two benchmark datasets: CT-RATE (25,692 scans) and CTRG-Chest-548K (1,804 scans). Evaluation metrics include clinical efficacy (CE-Precision, CE-Recall, CE-F1) alongside natural language generation (NLG) metrics (BLEU-1, BLEU-4, METEOR, ROUGE-L).
| Dataset | Method | CE-Pre. | CE-Rec. | CE-F1 | BL-1 | BL-4 | MTR | RG-L |
|---|---|---|---|---|---|---|---|---|
| CT-RATE | R2Gen (EMNLP'20) | 0.158 | 0.057 | 0.066 | 0.418 | 0.228 | 0.414 | 0.327 |
| CT-RATE | PromptMRG (AAAI'24) | 0.364 | 0.297 | 0.299 | 0.486 | 0.214 | 0.438 | 0.389 |
| CT-RATE | 3D-CT-GPT (2024) | 0.242 | 0.153 | 0.166 | 0.459 | 0.253 | 0.422 | 0.361 |
| CT-RATE | Reg2RG (TMI'25) | 0.423 | 0.181 | 0.253 | 0.473 | 0.249 | 0.441 | 0.367 |
| CT-RATE | Liu et al. (IPMI'25) | 0.337 | 0.366 | 0.315 | 0.540 | 0.289 | 0.472 | 0.385 |
| CT-RATE | Ours | 0.467 | 0.429 | 0.415 | 0.535 | 0.305 | 0.481 | 0.396 |
| CTRG-548K | PromptMRG (AAAI'24) | 0.311 | 0.347 | 0.301 | 0.477 | 0.283 | 0.485 | 0.494 |
| CTRG-548K | Liu et al. (IPMI'25) | 0.434 | 0.413 | 0.393 | 0.545 | 0.326 | 0.509 | 0.497 |
| CTRG-548K | Ours | 0.457 | 0.420 | 0.399 | 0.532 | 0.335 | 0.512 | 0.504 |
In specialized radiology NLP benchmarks on the CT-RATE test split, our method achieved a medical named entity recognition RaTEScore of 0.679 (surpassing Liu et al.'s 0.569) and an LLM-adjudicated GREEN score of 0.433, exceeding \(\mu^2\text{tokenizer}\) (0.429), a model explicitly optimized for the GREEN metric using direct preference optimization (DPO).
Ablation Study¶
A step-by-step ablation study on CT-RATE evaluates the individual contributions of each core architectural component:
| Config | Core Addition | CE-Pre. | CE-Rec. | CE-F1 | BL-1 | BL-4 | MTR | RG-L |
|---|---|---|---|---|---|---|---|---|
| (a) Baseline | Single-LLM with top-10 attention tokens | 0.439 | 0.343 | 0.356 | 0.379 | 0.199 | 0.317 | 0.327 |
| (b) + \(\mathcal{L}_\text{sparse}\) | Mask-guided sparse negative-entropy loss | 0.464 | 0.367 | 0.377 | 0.375 | 0.192 | 0.317 | 0.322 |
| (c) + AP-SFT | Abnormality-prefixed SFT on reference reports | 0.468 | 0.397 | 0.394 | 0.524 | 0.282 | 0.477 | 0.389 |
| (d) + AP-LTF | Filter \(K'=15\) from \(K=200\) via proposed shortlist | 0.461 | 0.418 | 0.391 | 0.514 | 0.289 | 0.447 | 0.377 |
| (e) + GRPO (Full) | Policy alignment closing proposal-generation gap | 0.467 | 0.429 | 0.415 | 0.535 | 0.305 | 0.481 | 0.396 |
| (f) Single-LLM | Shared multitask LLM for proposal & reporting | 0.451 | 0.390 | 0.381 | 0.424 | 0.302 | 0.393 | 0.366 |
Key Findings¶
- Substantial Clinical Diagnostic Improvements: On the CT-RATE benchmark, our framework improves the prior state-of-the-art CE-F1 score by 32.0% (from 0.315 to 0.415), while raising recall by 17.2% and precision by 10.4%, demonstrating dramatic reductions in both hallucinated abnormalities and missed clinical findings.
- Inter-LLM Alignment Unlocks Token Filtering Potential: Introducing AP-LTF in isolation (row d) increases recall (0.397 to 0.418) but degrades CE-F1 and NLG metrics due to exposure bias between training ground truths and inference proposals. Incorporating GRPO refinement (row e) successfully bridges this distributional gap, producing top performance across six of seven metrics.
- Specialized Dual-LLMs Outperform Multitask Single-LLMs: Forcing a single 3B LLM to perform both anomaly proposal and full report generation (row f) leads to severe performance degradation across all clinical and text metrics, validating that explicit separation of anomaly proposal from descriptive composition avoids negative task interference.
- Token Compression Sweet Spot: Hyperparameter exploration reveals clear parabolic behavior for both pool size \(K\) and selection count \(K'\). The optimal trade-off occurs at \(K=200\) and \(K'=15\), feeding exactly 160 visual tokens to the LLM (a 96.1% compression rate over 4,096 raw patch tokens) and enabling full inference on a single 32GB V100 GPU within 87 seconds.
- Robust Recovery from Cascade Omissions: Error analysis indicates that 27.2% of pathological abnormalities missed by the proposal network (\(D_\text{pro}\)) are subsequently recovered by the report generator (\(D_\text{ful}\)), facilitated by the geometric attention component of AP-LTF retaining informative image cues even during textual proposal misfires.
Highlights & Insights¶
- Targeting Attention Sinks in Medical Multimodal Learning: Identifies the critical failure mode where low-semantic tokens monopolize attention weights in 3D volumes, demonstrating that semantic text queries combined with linear projection successfully redirect feature selection to localized pathology.
- Cooperative Reinforcement Learning for Modular MLLMs: Provides an effective blueprint for aligning multi-stage language models using GRPO without dedicated value critics, closing the inference-training discrepancy caused by intermediate text representations.
- Zero-Latency Visual Priors via Training-Only Supervision: Cleverly harnesses prompt-based foundation segmentation models (SAT) to generate coarse anatomical masks during training to constrain feature entropy, discarding the segmentation network entirely at test time for streamlined deployment.
Limitations & Future Work¶
- Coarse Resolution Beyond Anatomical Compartments: The framework partitions features across 10 macroscopic anatomical structures. While effective at organ-level containment, micro-scale lesions (such as sub-centimeter lung nodules or hair-line bone fissures) remain susceptible to intra-compartment token averaging. Future extensions should incorporate adaptive lesion-level anchors.
- Absence of Longitudinal Prior Scan Comparison: In real clinical practice, radiologists heavily rely on comparisons with prior historical CT studies. The current framework processes single volumetric examinations in isolation and lacks multi-temporal comparative reasoning.
- Broad Transferability Across 3D Volumetric Imaging: The principles of coarse-to-fine visual search, attention-sink-resistant filtering, and reinforcement-aligned dual-agent processing are broadly modality-agnostic, offering strong promise for volumetric MRI and PET-CT applications.
Related Work & Insights¶
- vs Liu et al. (IPMI 2025): While Liu et al. pioneered structure-wise visual queries, their token selection relied strictly on raw attention top-K, making it highly vulnerable to attention sinks. Our framework adds semantic abnormality proposal and multi-source filtering, boosting CE-F1 by over 30%.
- vs Reg2RG / MedRegion-CT: These methods crop organ subvolumes using dedicated 3D segmentation models at inference time, imposing significant computational overhead and deployment latency. Our approach employs masks exclusively during pretraining via sparse negative-entropy loss, ensuring zero segmentation overhead at test time.
- vs 3D-CT-GPT / M3D: Earlier 3D MLLMs relied on uniform 3D spatial pooling or full-volume downsampling, obliterating localized pathology. Our framework demonstrates that structure-guided non-uniform compression achieves superior feature density and clinical fidelity.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Faithfully mimics radiologists' visual search through a dual-LLM pipeline, solves attention sinks in CTRG via AP-LTF, and harmonizes the architecture with GRPO]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive validation across two public benchmarks, featuring both standard NLG metrics and rigorous clinical evaluations including RaTEScore and GREEN]
- Writing Quality: ⭐⭐⭐⭐⭐ [Exceptionally clear presentation with seamless integration of clinical rationale, mathematical formulations, and empirical ablation insights]
- Value: ⭐⭐⭐⭐⭐ [Enables high-fidelity 3D CT reporting on accessible hardware constraints (2×V100) via 96.1% token compression, offering high clinical utility]