Skip to content

CMDR: Contextual Multimodal Document Retrieval

Conference: ECCV 2026
Paper: ECCV Official
Code: https://cmdr-bench.github.io
Area: Multimodal VLM
Keywords: Multimodal Document Retrieval, Contextual Embeddings, Late Interaction, Contrastive Learning, Document VQA

TL;DR

Addressing the fundamental limitation where vision-based document retrievers encode pages in isolation and discard cross-page dependencies, this work introduces CMDR-Bench (the first benchmark requiring multi-page contextual reasoning) and CMDR-Embed, which combines chunk-then-split joint encoding with context-aware contrastive learning to significantly outperform non-contextual baselines.

Background & Motivation

Real-world multi-page documents—such as technical manuals, research reports, project proposals, and legal filings—are inherently multimodal. Text, structured tables, infographics, and graphical layouts are closely interleaved across successive pages. With the rapid evolution of Retrieval-Augmented Generation (RAG) and Large Vision-Language Models (LVLMs), vision-based multimodal document retrieval (e.g., ColPali, VisRAG) has emerged as a promising paradigm. By rendering document pages directly as visual screenshots and encoding them end-to-end, these systems bypass the layout breakage and character noise typical of traditional OCR pipelines.

However, existing multimodal document retrieval benchmarks and retrieval systems are uniformly built upon the assumption of isolated single-page encoding. Standard benchmarks (e.g., MMLB-Doc, SlideVQA, ViDoRe) evaluate direct lexical or shallow semantic matching between an explicit query and an individual target page. In lockstep, current multimodal retrievers split multi-page documents into independent page images and encode each page in isolation. This paradigm overlooks a core reality of document understanding: complex queries frequently require indirect, multi-page contextual reasoning. For instance, an entity might be defined on one page and referenced by a pronoun on another, a table header may reside on a previous page while relevant data rows continue on the next, or a multi-step inference chain might span several consecutive sections. Isolated page embeddings strip away document-level context, leaving retrievers unable to locate target pages that contain the answer but lack verbatim keyword matches with the query.

Bridging this gap requires both evaluating and modeling cross-page context in visually-rich document retrieval. This endeavor presents two primary challenges: the lack of rigorous multi-page retrieval benchmarks requiring context, and the risk of inter-page information leakage and representation collapse when jointly encoding multiple pages. Core idea: introduce CMDR-Bench, an expert-annotated benchmark specifically requiring indirect cross-page reasoning, and propose CMDR-Embed, a framework that leverages sliding-window joint encoding followed by page-level splitting, coupled with a Contextual Multimodal Contrastive Learning (CMCL) objective with in-chunk and in-document hard negatives to preserve sharp page discriminability while infusing document context.

Method

Overall Architecture

CMDR-Embed processes an \(N\)-page document using an end-to-end pipeline structured as: sliding-window chunked joint encoding \(\rightarrow\) page-level multi-vector extraction \(\rightarrow\) fine-grained query-page late interaction. To prevent joint attention from over-smoothing page representations across the same document, the model is trained with the Contextual Multimodal Contrastive Learning (CMCL) loss, which explicitly contrasts positive pages against hard negatives from the same chunk and document.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multi-page Document Images<br/>D = {I1, ..., IN}"] --> B["Chunked Joint Encoding & Splitting<br/>Sliding-window LVLM encoding into page vectors"]
    B --> C["Fine-grained Late Interaction<br/>MaxSim query token matching over visual patches"]
    C --> D["Contextual Contrastive Learning CMCL<br/>Joint optimization with in-chunk & in-doc hard negatives"]
    D --> E["Ranked Output<br/>Context-aware page retrieval ranking"]

Key Designs

1. Chunked Joint Encoding & Splitting: sliding-window cross-page interaction with page-level vector separation Conventional visual document retrievers encode each page as an independent image, severing coreferences, multi-page table structures, and cross-page narratives. To inject document context into page representations, CMDR-Embed introduces a chunk-then-split mechanism. Given a document \(D = \{I_1, \dots, I_N\}\), the model groups consecutive pages into chunks via a sliding window with window size \(w\) and stride \(s\) (defaulting to \(w=4, s=2\)). The \(t\)-th chunk \(\{I_{s(t-1)+1}, \dots, I_{s(t-1)+w}\}\) is processed jointly by the vision-language backbone, allowing full self-attention across adjacent pages. Following joint encoding, the token embeddings corresponding to each individual page \(I_n\) are extracted to yield a page-level multi-vector representation \(\mathbf{E}^{I_n} \in \mathbb{R}^{N^I \times D}\). When consecutive chunks overlap (\(s < w\)), representations from the earlier chunk are retained, ensuring that each page receives contextual grounding from preceding pages without duplicate computation or representation thrashing.

2. Fine-grained Late Interaction: preserving local visual tokens under contextualized embeddings Compressing contextualized page representations into a single vector via mean pooling or last-token pooling collapses fine-grained visual details, such as localized table cells, tiny footnotes, and diagrams. CMDR-Embed preserves token-level granularity by adopting Late Interaction (LI). Given query token vectors \(\mathbf{E}^q \in \mathbb{R}^{N^q \times D}\) and contextualized page vectors \(\mathbf{E}^{I_n} \in \mathbb{R}^{N^I \times D}\), the relevance score computes the sum of maximal dot products across all query tokens:

\[\text{LI}(q, I_n) = \sum_{i=1}^{N^q} \max_{j \in [1, N^I]} \langle \mathbf{E}^{q}_i, \mathbf{E}^{I_n}_j \rangle\]

This allows contextualized query tokens to attend directly to specific visual patch tokens, aligning global document reasoning with precise on-page evidence.

3. Contextual Contrastive Learning CMCL: dual-scale hard negatives to mitigate representation collapse Jointly encoding multiple pages within a shared Transformer window carries a substantial risk: unconstrained attention allows information from neighboring pages to bleed into one another, making pages within the same document excessively similar and degrading page discriminability. Standard contrastive learning relies exclusively on in-batch negatives from different documents, which provides no gradient signal to penalize intra-document collapse. To resolve this, the authors propose Contextual Multimodal Contrastive Learning (CMCL), incorporating two types of context-aware hard negatives: in-chunk negatives \(\mathcal{I}_{\text{chunk}}\) (non-target pages within the same window, countering local over-smoothing) and in-document negatives \(\mathcal{I}_{\text{doc}}\) (non-target pages from other chunks of the same document, countering document-level topic confusion). The total loss balances contextual and standard in-batch terms:

\[\mathcal{L}_{\text{CMCL}} = \lambda \mathcal{L}_{\text{Context}} + (1 - \lambda) \mathcal{L}_{\text{Batch}}\]

where the contextual term enforces discrimination among intra-document candidates:

\[\mathcal{L}_{\text{Context}} = -\log \frac{\exp(\text{LI}(q, I^+)/\tau)}{\exp(\text{LI}(q, I^+)/\tau) + \sum_{I^-_n \in \mathcal{I}_{\text{chunk}} \cup \mathcal{I}_{\text{doc}}} \exp(\text{LI}(q, I^-_n)/\tau)}\]

Setting \(\lambda=0.5\) optimally balances contextual semantic integration with sharp page-level boundary preservation.

Loss & Training

Training is supported by CMDR-Synth, a synthesized dataset of 39,796 query-page pairs constructed via a multi-stage pipeline using Qwen2.5-VL 72B and UniSE to identify context-dependent pages and synthesize cross-page queries. CMDR-Embed is initialized from pretrained ColPali or ColQwen checkpoints and fine-tuned using LoRA (\(r=32, \alpha=32\)) on the LLM layers and projection heads. Training is conducted on 8 A100-80G GPUs across 3 epochs using the AdamW optimizer, FlashAttention, batch size 192, learning rate \(2 \times 10^{-4}\), and temperature \(\tau=0.02\). At deployment time, training-free Hierarchical Token Pooling can optionally reduce stored vectors by 80.0% while accelerating retrieval by 3.57\(\times\) with minimal performance loss.

Key Experimental Results

Main Results

Evaluation is conducted on CMDR-Bench, consisting of 255 long documents (averaging 183.5 pages) and 800 high-quality, human-annotated queries across four categories: Text Completion (TC), Coreference Resolution (CR), Structured Understanding (SU), and Multi-hop Reasoning (MR). The evaluation metric is nDCG@5.

Retriever Category Model Backbone #Params TC CR SU MR Overall
Text Retrievers BM25 Lexical - 29.1 19.5 20.9 24.9 23.6
Text Retrievers Contriever BERT-base 109M 35.7 15.0 23.5 26.4 25.1
Text Retrievers BGE BERT-base 109M 42.8 18.4 28.5 29.5 29.8
Text Retrievers NV-Embed-v2 Mistral-7B 7.9B 36.2 16.9 28.6 30.4 28.0
General Multimodal SigLIP SOViT-400m 878M 24.5 8.7 15.5 22.2 17.8
General Multimodal E5-V LLaVA-1.6 8.4B 35.4 22.8 31.2 34.8 31.0
General Multimodal Qwen3-VL Embedding Qwen3-VL 8B 38.2 26.6 35.7 37.2 34.4
Multimodal Document DSE Phi-3-V 4.2B 33.7 23.0 29.1 30.4 29.1
Multimodal Document VisRAG-Ret MiniCPM-V 3.4B 35.3 19.8 27.8 33.9 29.2
Multimodal Document ColPali PaliGemma 2.9B 39.1 24.2 30.2 35.4 32.2
Multimodal Document ColPali + Finetune PaliGemma 2.9B 41.0 27.5 36.8 39.9 36.3
Multimodal Document ColQwen Qwen2-VL 2.2B 38.0 28.9 30.0 35.9 33.2
Multimodal Document ColQwen + Finetune Qwen2-VL 2.2B 45.8 34.8 38.7 42.2 40.4
Contextual (Ours) CMDR-Embed_Pali ColPali 2.9B 53.8 33.3 57.7 54.3 49.8 (+13.5)
Contextual (Ours) CMDR-Embed_Qwen ColQwen 2.2B 64.6 38.4 64.7 58.9 56.6 (+16.2)

Ablation Study

The ablation study systematically assesses the impact of context-aware hard negatives, the CMCL objective, and the multi-vector Late Interaction mechanism on overall nDCG@5.

Configuration Setting CMDR-Embed_Pali CMDR-Embed_Qwen Note
Full Model Full CMDR-Embed 49.8 56.6 Full model with sliding window, LI, and CMCL
w/o In-Chunk Negatives Exclude \(\mathcal{I}_{\text{chunk}}\) 48.7 (-1.1) 53.9 (-2.7) Reduces discriminability against adjacent pages
w/o In-Document Negatives Exclude \(\mathcal{I}_{\text{doc}}\) 48.6 (-1.2) 53.2 (-3.4) Degrades accuracy on distant context pages
w/o CMCL Loss Train with \(\mathcal{L}_{\text{Batch}}\) only 45.2 (-4.6) 49.4 (-7.2) Lacks intra-document negative supervision
w/o Late Interaction Mean Pooling + Cosine Sim. 23.3 (-26.5) 38.1 (-18.5) Single-vector compression destroys visual grounding

Key Findings

  • Contextual encoding unlocks massive gains over data scaling: Even when non-contextual baselines are fine-tuned on the identical CMDR-Synth dataset (ColQwen reaching 40.4), CMDR-Embed_Qwen achieves 56.6 (+16.2 points), proving that architectural context modeling is indispensable.
  • Structured understanding and text completion benefit most: Structured Understanding (SU) gains over 26 points (38.7 to 64.7 for Qwen), and Text Completion (TC) improves by 18.8 points, reflecting how frequently table columns, headers, and paragraphs break across page boundaries.
  • Complementary roles of dual hard negatives: In-chunk negatives are crucial when the context page is immediately adjacent (1–2 pages away), whereas in-document negatives provide essential discrimination when context is separated by 3 or more pages.
  • Late interaction is fundamental: Replacing multi-vector late interaction with single-vector mean pooling causes performance to plummet by 26.5 points on Pali and 18.5 points on Qwen, confirming that fine-grained token alignments are vital for resolving contextual document queries.

Highlights & Insights

  • Adapting late chunking principles to vision-first document retrieval: Extends the concept of late chunking from plain text into the multimodal document regime, shattering the historical constraint of treating pages as isolated images.
  • Effective mitigation of attention over-smoothing: Discovers that joint encoding induces inter-page feature contamination within documents, elegantly resolving it with localized in-chunk and global in-document contrastive negatives.
  • Establishing CMDR-Bench as a multi-page visual reasoning frontier: Provides an authentic 183.5-page average benchmark requiring indirect contextual reasoning, moving the community beyond trivial keyword-matching benchmarks.

Limitations & Future Work

  • Unidirectional causal attention constraints: The autoregressive LLM backbones enforce causal attention, meaning page representations only incorporate preceding context and cannot attend to subsequent pages during encoding.
  • Context window length and memory tradeoffs: A sliding window of size \(w=4\) remains computationally constrained; extremely long-range multi-hop queries across dozens of pages remain challenging.
  • Fine-grained visual parsing in academic papers: Retrieval accuracy remains relatively lower on research papers and dissertations with dense mathematical formulas and minute chart labels, indicating that contextual modeling must be paired with stronger high-resolution visual encoders.
  • vs ColPali / ColQwen: ColPali pioneered end-to-end multi-vector document retrieval using vision-language backbones, but treats every page as an isolated snapshot. CMDR-Embed inherits its effective Late Interaction paradigm while breaking page isolation via chunked joint encoding and the CMCL loss.
  • vs DAPR / ConTEB: These benchmarks identified the necessity of document context for passage retrieval, but are confined to plain text. CMDR-Bench and CMDR-Embed successfully bring contextual retrieval into multimodal, visually rich, multi-page document domains.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Introduces the first dedicated contextual multimodal document retrieval benchmark and combines chunk-then-split encoding with targeted contrastive hard negatives.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorously evaluated across 255 long documents with human-curated queries, paired with extensive ablations across context distance, loss weights, and efficiency metrics.
  • Writing Quality: ⭐⭐⭐⭐⭐ Lucid problem formulation, crisp diagrams, and meticulous empirical analyses.
  • Value: ⭐⭐⭐⭐⭐ Sets a new paradigm for visual document retrieval and multimodal RAG, shifting the focus from isolated page matching to document-level contextual reasoning.