Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/sajeedmehrab/op-hrg
Area: Multimodal VLM
Keywords: Part-Level Visual Grounding / Multimodal Large Language Models / Hierarchical Reasoning / Reinforcement Learning / GRPO
TL;DR¶
To overcome the severe object-level bias in multimodal large language models (MLLMs) where part queries collapse into parent object detections, this paper proposes Object-Part Hierarchical Reflective Grounding (OP-HRG) trained with part-aware GRPO, enabling a 4B model to surpass 7B baselines and SAM3 across cross-dataset zero-shot benchmarks.
Background & Motivation¶
Localizing fine-grained parts of objects (such as a mug's handle or a car's wheel) is fundamental for physical interaction, robotic manipulation, and anatomical understanding. When humans search for an object part, they intuitively follow a coarse-to-fine cognitive hierarchy: first identifying the whole entity that contains the part, and then isolating the constituent sub-region within that context. While recent multimodal large language models (MLLMs) such as Qwen-VL have demonstrated strong zero-shot grounding performance for whole objects described by natural language queries, they struggle when queries target sub-object parts. When prompted to ground "a mug's handle," conventional MLLMs overwhelmingly return bounding boxes covering the entire mug.
This failure stems from two fundamental tensions. First, vision-language pretraining is heavily dominated by coarse object-level descriptions; dense and accurate part annotations are rare and costly, instilling a strong object-level prior into foundational vision-language representations. Second, grounding MLLMs treat every prompt identically through an unconstrained, single-step coordinate regression. Prior visual grounding pipelines—even recent reinforcement learning frameworks such as Seg-Zero and VisionReasoner—lack an explicit architectural mechanism or supervision signal to model the topological containment hierarchy between parent objects and constituent parts. Although MLLMs exhibit latent spatial reasoning capacity, it remains dormant without structured prompting and dedicated reinforcement signals.
The paper attacks this challenge by decomposing part grounding into an explicit coarse-to-fine reasoning chain paired with stage-wise reward supervision. Core idea: propose Object-Part Hierarchical Reflective Grounding (OP-HRG) to structure visual grounding into parent-object anchoring followed by part localization and crop-re-encoded visual self-reflection, optimized end-to-end under a part-aware GRPO framework with anti-exploitation baseline rewards.
Method¶
Overall Architecture¶
The framework follows a decoupled reasoning and segmentation design. Given an input RGB image \(I \in \mathbb{R}^{H \times W \times 3}\) and a natural language query \(q\), the MLLM (built upon Qwen3-VL-Instruct) first executes structured reasoning to generate a set of geometric localization primitives \(\mathcal{O} = \{(b_i, p_i)\}_{i=1}^K\), consisting of tight 2D bounding boxes \(b_i\) and interior representative points \(p_i\). These geometric primitives then serve as spatial prompts to condition a frozen, off-the-shelf mask decoder (such as SAM2 or SAM3), which produces the final binary pixel segmentation mask \(S \in \{0, 1\}^{H \times W}\).
The MLLM reasoning process strictly adheres to the OP-HRG paradigm across three interconnected stages: target category identification and hierarchical spatial anchoring (localizing parent objects before predicting candidate parts), active visual perception (cropping the predicted region, re-encoding it via the vision backbone, and performing self-reflective critique), and verifiable policy alignment through part-aware GRPO with stage-wise geometric rewards.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Image and Natural Language Query (I, q)"] --> B["Object-Part Hierarchical Prompting<br/>Target classification + Parent object anchoring + Initial part proposal"]
B --> C["Active Visual Perception & Self-Reflection<br/>Bounding box cropping + Visual feature reinjection + Criticism & refinement"]
C --> D["Part-Aware Verifiable RL Rewards<br/>Compactness penalty + Part containment + Anti-exploitation improvement reward"]
D --> E["Frozen Mask Decoder (SAM2 / SAM3)<br/>Dense high-precision part segmentation mask"]
Key Designs¶
1. Object-Part Hierarchical Prompting: Enforcing Coarse-to-Fine Spatial Anchoring
To resolve the object-level bias inherent in single-step coordinate regression, OP-HRG introduces a strict tagged output sequence. The model must systematically output intermediate reasoning states: <locate> (task reasoning), <target> (binary classification: object or part), <object_hint> (parent object bounding boxes), <first_answer> (initial part bounding boxes and representative points), <criticism> (self-reflection and adjustment decision), and <answer> (final refined primitives). When a query is identified as a part, the model is compelled to localize the parent object first as a spatial anchor. The initial part prediction is then structurally constrained to reside within the anchor box, converting an unconstrained global search into a contained, coarse-to-fine localization problem.
2. Active Visual Perception & Self-Reflection: Grounding Criticism in Fresh Visual Evidence
Conventional self-reflection in language models operates purely in text coordinate space, forcing models to infer spatial precision from raw numbers without directly inspecting the enclosed visual content. The active visual perception (AP) mechanism bridges this gap: immediately following the generation of <first_answer>, token generation pauses. The system automatically crops the predicted regions from the input image \(I\), passes them through the model's pretrained visual encoder, and reinjects the resulting visual tokens directly into the ongoing context. Guided by this fresh visual evidence, the model reviews whether the initial box is too loose, overlaps adjacent limbs, or encloses unwanted background, explicitly outputting ADJUSTMENT: YES or NO before writing the finalized coordinates in <answer>.
3. Part-Aware Verifiable RL Rewards: Eliminating Gaming and Degenerate Solutions
Optimizing complex multi-step reasoning with standard terminal IoU rewards frequently induces degenerate policy gaming, such as deliberately generating defective initial answers to maximize improvement margins or replicating parent object boxes to trivially pass containment checks. The framework introduces a modular, verifiable reward structure:
- Compactness Reward: Evaluated over Hungarian-matched prediction-ground-truth pairs via a precision term \(\rho_{ij} = |\hat{b}_i \cap b^*_j| / |\hat{b}_i|\) and an over-prediction penalty \(\omega_{ij} = -\min\left(1, \max(0, |\hat{b}_i| / |b^*_j| - 1)\right)\), which saturates at \(-1\) once the predicted box area doubles the ground-truth area to strongly penalize bloated boxes.
- Part Containment Reward: Enforces three explicit conditions on predicted part boxes relative to parent object boxes: spatial containment, non-identity, and strictly smaller area (\(|\hat{b}_{\text{part}}| < |\hat{b}_{\text{obj}}|\)), preventing trivial copying of parent boxes.
- Anti-Exploitation Improvement Reward: To prevent the model from gaming the self-correction mechanism:
$\(R^{\text{IoU}}_{\text{improv}} = \max\left(0, R^{\text{IoU}}_{\text{final}} - \max\left(R^{\text{IoU}}_{\text{initial}}, \lambda_{\text{IoU}} \cdot \text{IoU}_{\text{baseline}}\right)\right)\)$
where \(\text{IoU}_{\text{baseline}}\) is an offline precomputed score from SAM3. The model earns reward only when its final output surpasses both its own initial prediction and the strong external baseline, while degradation from initial to final answer triggers an explicit penalty \(R^{\text{IoU}}_{\text{final}} - R^{\text{IoU}}_{\text{initial}}\).
- Adjustment Consistency Penalty: Imposes a hard penalty of \(-1\) if the model exhibits cognitive dissonance—declaring ADJUSTMENT: YES without altering coordinates, or stating ADJUSTMENT: NO while shifting predictions.
Loss & Training¶
The framework bypasses supervised fine-tuning (SFT) and directly optimizes the pretrained Qwen3-VL-Instruct checkpoint using Group Relative Policy Optimization (GRPO). For each image-query pair \((I, q)\), GRPO samples a candidate group of structured trajectories \(G = \{y_j\}_{j=1}^{|G|}\) and computes normalized advantages \(A_j\) relative to the group mean: $\(\mathcal{L}_{\text{GRPO}} = -\frac{1}{|G|}\sum_{j=1}^{|G|} \left[\min\left(\rho_j A_j, \text{clip}(\rho_j, 1-\epsilon, 1+\epsilon) A_j\right)\right] + \beta \mathbb{D}_{\text{KL}}(\pi_\theta \,\|\, \pi_{\theta_0})\)$ where \(\rho_j = \pi_\theta(y_j|I, q) / \pi_{\theta_0}(y_j|I, q)\), and \(\pi_{\theta_0}\) is the frozen reference policy. The training data incorporates a general 7k multi-object dataset alongside 1,200 curated images from the InstructPart training split. Representative interior points are extracted automatically from binary masks using Euclidean distance transform extrema.
Key Experimental Results¶
Main Results¶
Evaluation is reported in gIoU (mean IoU across all test queries). InstructPart serves as the in-domain benchmark (600 queries), whereas PascalPart-116 (10k part queries) and PartImageNet (14k part queries) represent challenging cross-dataset zero-shot benchmarks where neither images nor part labels were seen during training.
| Paradigm | Model Architecture | Parameters | InstructPart (Parts) | PascalPart (Parts) | PartImageNet (Parts) | Pascal-Obj (Objects) |
|---|---|---|---|---|---|---|
| Token-based MLLM | LISA-7B | 7B | 43.26 | 13.82 | 29.91 | 83.50 |
| Token-based MLLM | Sa2VA-4B | 4B | 50.15 | 14.67 | 38.13 | 77.02 |
| Token-based MLLM | PixelLM-7B | 7B | 44.41 | 16.57 | 35.25 | 81.71 |
| Token-based MLLM | UniPixel-3B | 3B | 59.68 | 30.64 | 46.18 | 85.54 |
| Decoupled MLLM-SAM | Molmo + SAM3 (point) | 7B+ | 51.02 | 8.87 | 17.30 | 57.57 |
| Decoupled MLLM-SAM | Grounding DINO + SAM3 (box) | - | 33.74 | 13.13 | 31.38 | 79.25 |
| Decoupled MLLM-SAM | VisionReasoner-7B | 7B | 59.38 | 27.44 | 45.59 | 86.68 |
| Text-prompted Segmentation | SAM3 (text direct) | - | 71.06 | 33.05 | 53.89 | 85.32 |
| Decoupled MLLM-SAM | Qwen3-VL-4B + SAM2 (Zero-shot Prompt) | 4B | 31.45 | 12.83 | 21.95 | 36.91 |
| Decoupled MLLM-SAM | OP-HRG (Qwen3-VL-4B + SAM2, Ours) | 4B | 75.56 | 38.59 | 56.87 | 87.50 |
| - | Gain vs. best baseline (\(\Delta\)) | - | +4.50 | +5.54 | +2.98 | +0.82 |
Ablation Study¶
Ablations on the InstructPart test set validate the necessity of each architectural component and training signal:
| Category | Configuration / Variant | gIoU (%) | Description & Impact |
|---|---|---|---|
| Data & Baseline Comparison | VisionReasoner-7B (original) | 59.38 | 7B baseline trained with standard localization rewards |
| VisionReasoner-7B + InstructPart RL | 62.33 | Exposing baseline to part data without OP-HRG yields only modest gains (+2.95) | |
| Ablation of Mechanisms | Full model w/o hierarchy & refinement | 70.32 | Standard localization rewards without OP-HRG structure (-5.24) |
| Full model w/o hierarchical grounding | 73.77 | Removing parent anchoring and containment rewards (-1.79) | |
| Full model w/o reflective refinement | 73.98 | Removing self-reflection loop and improvement reward (-1.58) | |
| Full Architecture | Full OP-HRG (4B) | 75.56 | Both hierarchy and reflection provide complementary gains |
Key Findings¶
- Compact 4B Model Outperforms 7B Baselines and Dedicated SAM3: The 4B OP-HRG model surpasses the 7B VisionReasoner by massive margins (+16.18 on InstructPart, +11.15 on PascalPart) and outstrips the text-conditioned SAM3 across both in-domain and zero-shot splits, demonstrating that structured hierarchical reasoning is vastly more effective than parameter scaling alone.
- Reflection Acts as a Train-Time Regularizer and Becomes Internalized: Tracking checkpoints across training reveals that active coordinate revisions occur frequently early on with net positive IoU gains. Over training, the initial answer converges to near-optimal quality, turning the reflection step into a passive verification. Evaluating the trained model with a single-step inference prompt retains 75.40 gIoU (vs. 75.56), showing that the reflective benefit is largely internalized into the model weights.
- Preserved Whole-Object Grounding and Strong Transferability: On referring expression benchmarks (RefCOCO/+/g), the model maintains 86.5% average [email protected] (closely matching the un-tuned base model's 89.7%). Furthermore, it transfers seamlessly to reasoning segmentation (ReasonSeg), achieving 69.6 gIoU and outperforming specialized baselines like GenSeg-R1-4B (68.4) and VisionReasoner-7B (63.6).
Highlights & Insights¶
- Decoupled Prompting Bridges Semantics and Pixel Precision: Instead of fine-tuning expensive mask decoders within LLMs, the method leverages bounding boxes and interior points as lightweight interfaces, unlocking the full zero-shot segmentation power of frozen SAM models.
- Baseline-Anchored Improvement Reward Solves RL Exploitation: By anchoring the improvement reward against an external SAM3 baseline, the formulation mathematically closes the vulnerability where RL agents intentionally emit degraded first answers to harvest easy delta rewards.
- Visual Crop Re-injection Grounds LLM Self-Critique: Unlike token-only reflection models that hallucinate spatial adjustments, physically re-encoding the cropped prediction allows the MLLM to inspect genuine visual evidence before confirming adjustments.
Limitations & Future Work¶
- Late-Stage Verification Collapse: As the initial predictions approach peak accuracy during training, the policy almost universally selects
ADJUSTMENT: NO, causing the self-correction mechanism to behave as a static verifier rather than an active refiner. - Multi-Instance Part Attribution Ambiguity: The part containment reward accepts inclusion in any parent object box of the correct category. In dense scenes with multiple identical instances, a part localized within the wrong parent instance is not explicitly penalized.
- Inference Latency Overhead: The two-stage generation and intermediate image cropping add approximately one second of wall-clock latency per query, motivating future exploration into dynamic, test-time adaptive execution.
Related Work & Insights¶
- vs. Seg-Zero / VisionReasoner: While sharing the decoupled MLLM-to-SAM paradigm, prior works treat all queries as flat entities in a single step; OP-HRG explicitly models parent-part hierarchies and containment constraints, drastically closing the part-grounding gap.
- vs. LISA / Sa2VA / PixelLM: Special-token models train mask decoders jointly with LLMs, incurring high compute costs and risking visual degradation; OP-HRG keeps the decoder frozen and focuses purely on spatial reasoning.
- vs. SAM3 (Text Prompt): Direct text-prompted segmentation in SAM3 lacks multi-step spatial reasoning for complex nested expressions; OP-HRG acts as an expert spatial planner that provides focused, coarse-to-fine visual prompts to the decoder.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Pioneers an object-part hierarchical reflective paradigm with crop-based active perception and anti-exploitation RL rewards.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorously evaluated on cross-dataset zero-shot part benchmarks, decoder swap ablations, referring expression preservation, and reasoning segmentation transfer.
- Writing Quality: ⭐⭐⭐⭐⭐ Well-structured narrative with mathematically precise reward definitions and clear ablations.
- Value: ⭐⭐⭐⭐⭐ Highly practical for robotic manipulation, embodied AI, and fine-grained visual understanding pipelines.