Skip to content

RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models

Conference: ECCV 2026
Paper: ECCV Paper
PDF: ECCV PDF
Area: Multimodal VLM
Keywords: region-level retrieval, large multimodal models, region-aware encoder, contrastive learning, REGMB benchmark

TL;DR

Addressing the trade-off between fine-grained local details and global background context in large multimodal models for region-level retrieval, RegRet introduces a layer-wise cross-attention Region-Aware Encoder with decoupled projectors alongside a three-stage training pipeline and the 225k REGMB benchmark, boosting regional retrieval accuracy by over 20% while preserving global retrieval capability.

Background & Motivation

Multimodal information retrieval plays a critical role in real-world applications including e-commerce product search, medical image retrieval, and multimodal retrieval-augmented generation (RAG). In practical industrial scenarios, user queries frequently rely on regions of interest (ROIs). For instance, a shopper might upload an indoor photo of a bedroom and draw a bounding box around a carpet to search for items with identical texture and pattern. Without explicit ROI modeling, conventional systems relying exclusively on global representations are easily misled by dominant background features, such as room layout or wall color, incorrectly retrieving images with matching scenes but completely different carpets. Despite its practical significance, existing multimodal retrieval models primarily optimize holistic image embeddings, leaving the problem of pushing retrieval accuracy when region-level guidance is available largely underexplored.

Directly applying current large multimodal models (LMMs) to region-level retrieval encounters two fundamental tensions. First, balancing fine-grained regional details against background context remains challenging. Existing regional visual prompting heuristics exhibit severe deficiencies: visual markers rely heavily on grounding capability and allow the background to dominate; direct cropping or ROIAlign discards essential contextual cues needed to interpret the region (e.g., misclassifying an outdoor umbrella as white fabric without shop facades); and auxiliary image concatenation often injects distractor features from the broader image into the final embedding. Moreover, naively fine-tuning a shared encoder on regional tasks leads to catastrophic forgetting and performance degradation on global-level retrieval. Second, both large-scale contrastive training datasets and diverse evaluation benchmarks with explicit region-level annotations are notably scarce, with prior benchmarks predominantly restricted to simple image-to-image matching.

The core intuition is that rather than relying on external visual prompting or a monolithic feature space, models require a dedicated region-aware pathway with decoupled representations that learn to adaptively ingest context via cross-attention. The core idea is to design a layer-wise Region-Aware Encoder (RAE) with decoupled projection layers that coordinates with a frozen Context Encoder (CE) to adaptively balance background cues and local features, supported by a three-stage training pipeline and a comprehensive 225k-pair benchmark (REGMB).

Method

Overall Architecture

RegRet is constructed upon an LMM backbone. The input consists of a query image paired with an ROI bounding box \(r_q\), along with optional text instructions. The image is processed simultaneously by two pathways: the native Context Encoder (CE) processes the entire scene to capture global context, while the Region-Aware Encoder (RAE) extracts fine-grained regional tokens from the cropped ROI. RAE features iteratively interact with corresponding CE layers via cross-attention across all depths. Finally, RAE tokens and CE tokens are projected into the LLM embedding space via separate projection layers. The LLM aggregates multimodal tokens, and the hidden state following a dedicated [EMB] query token is extracted as the final retrieval embedding.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Image + ROI Box + Instruction"] --> B["Stage 1: Region-Aware Encoder Pretraining<br/>Layer-wise context fusion via localized captioning"]
    B --> C["Stage 2: Pure-Text Contrastive Learning<br/>InfoNCE converts LLM into embedding model"]
    C --> D["Stage 3: Regional Contrastive Learning<br/>Mixed visual prompting and hard negative pairs"]
    D --> E["Decoupled Retrieval Embeddings<br/>[EMB] token supports regional & global retrieval"]

Key Designs

1. Layer-Wise Region-Aware Encoder (RAE): Hierarchical Multilevel Semantic Absorption Existing encoder-decoder approaches typically perform cross-attention only at the final layer of the visual backbone, losing granular low-level details such as fabric textures and physical material properties. RAE adopts a layer-wise coordination architecture where every layer refines the ROI tokens via self-attention and subsequently queries the corresponding layer of CE through cross-attention. Queries (\(Q\)) originate from RAE hidden states, while keys (\(K\)) and values (\(V\)) are supplied by CE hidden states at the identical layer. Self-attention weights are shared between CE and RAE to eliminate training from scratch and reduce parameters. The updated representation is governed by a learnable scalar \(\alpha\) initialized to zero: $\(h_{\mathrm{RAE}}^{i+1} = h_{\mathrm{RAE}}^{i} + \alpha \cdot \operatorname{xattn}\left(h_{\mathrm{RAE}}^{i},\, h_{\mathrm{CE}}^{i},\, h_{\mathrm{CE}}^{i}\right)\)$ This enables RAE to smoothly absorb low-level physical details from early layers alongside high-level contextual semantics from deeper layers.

2. Decoupled Projector Architecture: Preserving Global-Level Representation Integrity When RAE and CE visual tokens share a single projection layer into the LLM, gradients from region-level fine-tuning distort the shared embedding space, causing a catastrophic drop in global retrieval performance. RegRet resolves this semantic misalignment by explicitly freezing the CE visual encoder and its global projector, while introducing an independent dedicated projector for RAE. This decoupling isolates the regional optimization space from global features, enabling full specialization of the regional representations without sacrificing or degrading holistic image-level retrieval capabilities.

3. Three-Stage Progressive Training Pipeline: From Localized Captioning to Metric Alignment To structure the optimization difficulty progressively, RegRet follows a three-stage curriculum. Stage 1 executes RAE pretraining under a next-token prediction paradigm using detailed localized image-text datasets (DAM and PAM), optimizing only RAE cross-attention modules and the projector while freezing the LLM backbone. Stage 2 applies pure-text contrastive learning using large-scale text pairs to cost-effectively convert the LLM's generative capability into dense vector representations. Stage 3 conducts regional contrastive learning on a mixture of global and regional pairs with hard negatives. It adopts a mixed visual prompting strategy: local-centric tasks feed exclusively RAE tokens to the LLM, whereas context-dependent tasks concatenate CE and RAE tokens, establishing optimal discriminability.

Loss & Training

During Stage 1, the model is trained with an autoregressive next-token prediction cross-entropy loss over detailed localized descriptions: $\(\mathcal{L}_{\mathrm{rae}} = -\frac{1}{T}\sum_{t=1}^{T} \log P(x_t \mid x_{<t}; \theta)\)$ In Stages 2 and 3, optimization is driven by the InfoNCE contrastive loss over query-candidate pairs: $\(\mathcal{L}_{\mathrm{rcl}} = -\frac{1}{N} \sum_{i=1}^N \log \frac{\exp\left(\langle q_i \otimes r_{q_i},\, c_i^+ \otimes r_{c_i^+} \rangle / \tau\right)}{\sum_{j=1}^N \exp\left(\langle q_i \otimes r_{q_i},\, c_j \otimes r_{c_j} \rangle / \tau\right)}\)$ where \(\tau\) is the temperature parameter and \(N\) is batch size. To accommodate large batch sizes while preserving memory, LoRA is applied to the LLM backbone (rank 64 in Stage 2, rank 128 in Stage 3). The full three-stage training pipeline is completed within 20 hours on 24 NVIDIA H20 GPUs.

Key Experimental Results

Main Results

The REGMB benchmark evaluates four core settings: Task 1 (Text-to-Image, T2I), Task 2 (Image-to-Text, I2T), Task 3 (Image-to-Image, I2I), and Task 4 (Image-Text-to-Image, IT2I). Recall@5 (R@5) is reported for most splits, while Recall@1 (R@1) is used for VisMin and ImgDiff to avoid metric saturation.

Category Method Size Task 1 (SAM) Task 2 (SAM) Task 3 (XGoods) Task 4 (VisMin) REGMB Avg.
CLIP-based FG-CLIP 0.2B 48.5 48.9 60.7 53.6 58.5
CLIP-based SigLIP2 0.4B 46.3 55.9 79.0 30.5 62.9
LMM-based MM-EMBED 8B 42.5 35.2 80.5 82.9 61.3
LMM-based LamRA 8B 41.3 46.7 83.3 82.6 65.7
LMM-based mmE5 11B 50.4 42.2 82.0 80.9 67.5
Ours RegRet-3Bโ€  3B 58.1 70.6 91.7 83.1 79.9
Ours RegRet-8B-zs (Zero-Shot) 8B 59.5 72.4 86.3 81.1 78.8
Ours RegRet-8Bโ€  8B 69.1 86.5 93.6 83.4 86.2

On external out-of-domain benchmarks (ROxford-Hard, DeepFashion2, and ILIAS), RegRet demonstrates superior generalization:

Method ROxford-Hard (mAP) DeepFashion2 (R@1) ILIAS-I2I (mAP@50) ILIAS-T2I (mAP@50) Benchmark Avg.
LamRA 42.0 13.2 71.1 62.6 47.2
mmE5 36.9 11.1 50.1 39.0 34.2
RegRet-8B-zs 42.8 15.8 75.7 59.1 48.3
RegRet-8B 45.0 20.2 86.4 72.6 56.1

Ablation Study

Ablations on visual architecture components confirm the necessity of cross-attention, separate projectors, and layer-wise coordination:

Architecture Variant xattn Separate Proj. Layer-wise Task 1 Task 2 Task 3 Task 4 REGMB Avg. M-BEIR Avg.
Auxiliary Image (Baseline) โœ— โœ— โœ— 69.8 82.7 92.6 86.7 83.4 57.1
RAE-sharep (Shared Proj.) โœ“ โœ— โœ— 77.2 83.5 51.7 52.2 66.0 27.2
RAE-encdec (Single-layer + Sep. Proj.) โœ“ โœ“ โœ— 76.6 83.8 89.4 81.3 82.6 57.1
Full RAE โœ“ โœ“ โœ“ 79.0 86.9 92.9 86.0 86.2 57.3

Key Findings

  • Decoupled projectors prevent catastrophic interference: sharing the projection layer (RAE-sharep) collapses global M-BEIR retrieval performance from 57.1 to 27.2, whereas independent projectors completely restore global capability to 57.1 and rise to 57.3 with full RAE.
  • Layer-wise coordination provides vital low-level visual grounding: propagating intermediate features across all layers yields a 2.4% gain on Task 1 and 3.1% on Task 2 compared to the single-layer encoder-decoder design.
  • Superior zero-shot accuracy and inference efficiency: RegRet-8B-zs (78.8) significantly outperforms LamRA with auxiliary images (63.1) and cropping (55.9), while running 22% faster in inference time than the auxiliary image strategy (7.1s vs. 9.1s per batch).

Highlights & Insights

  • Layer-wise cross-attention with parameter sharing: RAE reuses pretrained self-attention weights from the context backbone, optimizing only lightweight cross-attention layers to achieve adaptive contextual grounding at minimal parameter cost.
  • Representation decoupling eliminates feature interference: dedicating separate projection matrices to global and regional features prevents fine-grained regional gradients from corrupting holistic global representations.
  • REGMB benchmark: provides a diverse 225k contrastive dataset spanning realistic cross-scene social e-commerce pairs with lighting and posture variations, establishing a rigorous benchmark for regional multimodal evaluation.

Limitations & Future Work

  • Context vs. detail trade-off in dense editing: on certain image-editing subtasks (e.g., Task 4 subsets) where retrieval hinges almost entirely on global background shifts rather than regional foregrounds, focusing on ROI features slightly compromises performance compared to purely global models.
  • Bounding box prompt precision: current ROI inputs rely on rectangular bounding boxes, which inevitably include background noise for non-convex or heavily occluded objects; integrating polygon or mask-level prompts could further improve boundary fidelity.
  • Scope extension: future iterations could extend RegRet's architecture to visual document retrieval and dense table cell alignment in multimodal RAG pipelines.
  • vs. FG-CLIP / FineCLIP: Early region-level methods rely on CLIP with ROIAlign cropping, missing LMM instruction-following flexibility and discarding critical background context; RegRet leverages an LMM backbone and learns adaptive background integration via RAE cross-attention.
  • vs. LamRA / mmE5: Generalist multimodal embedding models resort to heuristic prompting (visual markers or auxiliary images) that introduce background distractor noise; RegRet decouples regional representations via a dedicated encoder and separate projectors, improving regional retrieval by >20% while safeguarding global accuracy.

Rating

  • Novelty: โญโญโญโญ [The layer-wise cross-attention RAE and decoupled projector architecture elegantly reconcile regional and global retrieval]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluation across REGMB and multiple external benchmarks, backed by rigorous ablations and latency analyses]
  • Writing Quality: โญโญโญโญโญ [Clear motivation, well-structured methodology, and self-consistent empirical analysis]
  • Value: โญโญโญโญโญ [Directly tackles core bottlenecks in fine-grained e-commerce search and multimodal RAG with high practical utility]