PanoRec: Spatially-Structured Sequence Modeling for Multi-Granularity Panoramic Retrieval¶
Conference: ECCV 2026
Paper: ECCV Official
Area: 3D Vision / Multimodal VLM
Keywords: panoramic image retrieval, multi-granularity retrieval, spatial sequence modeling, semantic dilution, MaxSim routing
TL;DR¶
Addressing the severe semantic dilution caused by compressing 360-degree scenes into a single global embedding in conventional vision-language models, PanoRec serializes distortion-free cubemap faces, a downsampled global panorama, and spatial anchor tokens into a unified sequence, combining dynamic MaxSim routing with a joint spatial InfoNCE objective to achieve efficient multi-granularity panoramic retrieval in a single forward pass.
Background & Motivation¶
Panoramic images capture comprehensive 360-degree by 180-degree environmental visuals, serving as a critical cornerstone for virtual reality, autonomous navigation, and immersive scene perception. Following the dramatic advances in large-scale vision-language models (VLMs) across standard 2D visual reasoning tasks, adapting these foundational priors directly to panoramic cross-modal retrieval has become highly desirable. However, standard panoramic imagery in the equirectangular projection (ERP) format exhibits extreme spherical distortions along high-latitude regions. Applying specialized spherical operators or geometric attention layers introduces invasive architectural modifications that compromise the valuable 2D visual priors established during massive pre-training. Consequently, projection-based approaches that transform the spherical field into distortion-free planar perspective views, such as cubemaps, have become the established practice to bridge the gap.
Nevertheless, existing retrieval pipelines predominantly process panoramic content into a single, global holistic embedding vector. In pilot zero-shot evaluations using state-of-the-art Qwen3-VL-Embedding models across indoor Matterport3D and outdoor CVACT benchmarks, a severe granularity asymmetry emerged: while the models achieved solid retrieval under holistic scene-level captions, their performance plummeted when evaluated against fine-grained local queries describing specific cubemap views (for instance, on Matterport3D, Recall@1 for the 8B model collapsed from 62.1% down to 6.2%). The primary culprit is semantic dilution: an information-dense panorama contains vast contextual and background details, and forcing this extensive visual content into a single compressed vector inevitably submerges subtle local target signals beneath irrelevant background noise.
Simply processing individual perspectives in isolation breaks the holistic spatial continuity of the scene, whereas holistic pooling triggers semantic dilution. Core idea: abandon single-vector compression by interleaving distortion-free cubemap perspectives and a downsampled global panorama with learnable spatial anchor tokens in a unified sequence, extracting a multi-vector representation in a single forward pass, and employing dynamic MaxSim routing to simultaneously filter background noise and mine in-batch hard spatial negatives.
Method¶
Overall Architecture¶
PanoRec follows a dual-encoder retrieval paradigm comprising a Query Tower and a Panorama Tower. The Query Tower encodes diverse incoming queries (such as localized textual descriptions, holistic scene captions, or narrow field-of-view image crops) into dense query embeddings. The Panorama Tower processes panoramic imagery through a spatially-structured sequence modeling framework: six distortion-free cubemap faces and a downsampled global ERP image are interleaved with learnable spatial anchor tokens into a unified multi-modal sequence. In a single forward pass, the model extracts a decoupled multi-vector representation consisting of six view-specific local embeddings and one holistic scene embedding. A joint spatial InfoNCE objective with dynamic MaxSim routing supervises this space, balancing holistic scene understanding with fine-grained local discrimination.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Query & ERP Panorama<br/>(Query & ERP Image)"] --> B["Spatially-Structured Sequence Modeling<br/>Cubemap Faces + Global Image + Spatial Anchor Tokens"]
B --> C["Single-Pass Multi-Vector Extraction<br/>Single VLM forward pass yields 6 local + 1 global embeddings"]
C --> D["Dynamic MaxSim Routing & Joint Spatial Optimization<br/>Dynamic view alignment + Hard spatial negative mining + Joint InfoNCE"]
D --> E["Multi-Granularity Panoramic Retrieval Output<br/>(Global/Local Text & Image Retrieval)"]
Key Designs¶
1. Spatially-Structured Sequence Modeling: Preserving Local Distortion-Free Geometry and Global Topology
Directly feeding raw ERP images inevitably introduces heavy geometric distortion, whereas processing perspective patches as an unordered bag-of-views disrupts physical spatial continuity. To address this tension, PanoRec first projects the ERP image into six distortion-free planar cubemap faces (front, right, back, left, up, down), which natively align with standard 2D VLM visual tokens. To ensure the model retains macro-level environmental awareness without incurring prohibitive computational overhead, a downsampled version of the global ERP panorama is integrated. To establish rigorous spatial topology across views without modifying the underlying attention mechanisms of the VLM, seven learnable spatial anchor tokens are added to the vocabulary and interleaved directly following each corresponding visual input:
This structured sequence preserves the geometric scanning order of the surrounding physical environment, enabling the VLM to perform holistic cross-view reasoning natively within its self-attention layers.
2. Single-Pass Multi-Vector Extraction: Non-Invasive and Efficient Multi-Granularity Decoupling
Prior efforts to avoid semantic dilution via dense patch matching suffer from severe index storage inflation and high query latency, while multi-branch architectures induce heavy inference overhead. Leveraging the native interleaved multi-modal capabilities of foundation VLMs, PanoRec executes a single forward pass over the unified sequence \(\mathcal{S}\) and extracts the final normalized hidden states strictly at the positions corresponding to the spatial anchor tokens. This directly produces six decoupled local view embeddings \(E_\text{local} \in \mathbb{R}^{6 \times D}\) and one holistic global embedding \(E_\text{global} \in \mathbb{R}^D\). This design completely avoids destructive pooling operations, maintaining pristine localized visual features while preserving global context at minimal computational cost.
3. Dynamic MaxSim Routing & Joint Spatial Optimization: Noise Suppression and Hard Spatial Negative Mining
When evaluating a localized query \(q_\text{local}\) against the six view-specific representations \(E_\text{local}\), forced aggregation via mean or attention pooling fatally re-introduces background noise. PanoRec adopts a dynamic MaxSim routing mechanism where the retrieval score is determined by the maximum cosine similarity across all six cubemap faces:
At inference time, this dynamically aligns the query to its single most relevant perspective, discarding background noise from non-target faces. During training within the joint spatial InfoNCE objective, MaxSim introduces an implicit hard spatial negative mining mechanism: a local query is contrasted not only against its true view but also against the most confusing (highest-scoring) faces from non-target panoramas in the batch. This discourages the model from relying on superficial global shortcuts (such as overall lighting or ambient color palettes) and compels it to learn fine-grained spatial discrimination.
Loss & Training¶
The overall training objective combines global and local contrastive losses:
Here, \(\mathcal{L}_\text{global}\) aligns holistic captions with \(E_\text{global}\) via standard cosine similarity, while \(\mathcal{L}_\text{local}\) guides fine-grained alignment between localized queries and \(E_\text{local}\) via MaxSim routing. The visual encoder remains completely frozen throughout training to protect pre-trained 2D visual priors, while the seven spatial anchor tokens are trained from scratch. The language model layers can be fine-tuned via either parameter-efficient LoRA or full-parameter fine-tuning.
Key Experimental Results¶
Main Results¶
Evaluations on the indoor ZInD benchmark and outdoor CVACT benchmark encompass global caption retrieval, local caption retrieval, and localized narrow-FoV image crop retrieval.
| Dataset | Training | Method | Global Caption R@1 (%) | Global Caption R@5 (%) | Local Caption R@1 (%) | Local Caption R@5 (%) | Local Image R@1 (%) | Local Image R@5 (%) |
|---|---|---|---|---|---|---|---|---|
| ZInD (Indoor) | LoRA | Baseline-2B | 60.61 | 86.17 | 7.63 | 20.32 | 11.93 | 28.53 |
| ZInD (Indoor) | LoRA | PanoRec-2B (Ours) | 68.96 | 90.97 | 61.98 | 83.15 | 88.88 | 95.78 |
| ZInD (Indoor) | LoRA | Baseline-8B | 68.60 | 90.90 | 11.50 | 28.18 | 15.36 | 35.49 |
| ZInD (Indoor) | LoRA | PanoRec-8B (Ours) | 76.83 | 94.78 | 74.80 | 90.87 | 92.31 | 96.87 |
| ZInD (Indoor) | Full-FT | Baseline-8B | 71.49 | 92.86 | 17.33 | 39.55 | 16.16 | 38.55 |
| ZInD (Indoor) | Full-FT | PanoRec-8B (Ours) | 79.68 | 95.97 | 76.55 | 92.19 | 82.63 | 89.90 |
| CVACT (Outdoor) | LoRA | Baseline-8B | 59.98 | 83.78 | 15.77 | 33.43 | 30.55 | 50.90 |
| CVACT (Outdoor) | LoRA | PanoRec-8B (Ours) | 62.10 | 85.92 | 69.37 | 87.41 | 94.83 | 97.67 |
| CVACT (Outdoor) | Full-FT | Baseline-8B | 57.79 | 83.60 | 16.37 | 36.05 | 27.66 | 48.52 |
| CVACT (Outdoor) | Full-FT | PanoRec-8B (Ours) | 61.91 | 85.94 | 71.47 | 89.00 | 93.00 | 95.43 |
Ablation Study¶
Ablations on ZInD (using the 8B backbone under full fine-tuning) systematically evaluate input representation paradigms and similarity scoring interaction mechanisms.
| Category | Variant / Interaction | Global Caption R@1 (%) | Local Caption R@1 (%) | Local Image R@1 (%) | Note |
|---|---|---|---|---|---|
| Input & Representation | (i) ERP Baseline (Single Vector) | 71.49 | 17.33 | 16.16 | Standard high-resolution ERP, suffering from distortion and semantic dilution |
| Input & Representation | (ii) Single-Vector Cubemap | 65.86 | 18.95 | 24.23 | Resolves geometric distortion but aggregates into one vector; local retrieval remains stagnant |
| Input & Representation | (iii) Multi-Vector Cubemap (w/o Global) | 71.91 | 74.78 | 82.46 | Decoupled 6 local anchor tokens yield a dramatic surge (>55% gain in local R@1) |
| Input & Representation | (iv) Full PanoRec (Ours) | 79.68 | 76.55 | 82.63 | Integrates global downsampled panorama, boosting both global and local metrics |
| Scoring Interaction | Mean Pooling | 77.38 | 16.25 | 32.75 | Forcible fusion averages background noise, severely harming local discriminability |
| Scoring Interaction | Attention Pooling | 77.68 | 23.29 | 49.17 | Weighted fusion fails to fully insulate localized targets from background interference |
| Scoring Interaction | MaxSim Routing (Ours) | 79.68 | 76.55 | 82.63 | Dynamic routing eliminates background interference while mining hard spatial negatives |
Key Findings¶
- Semantic dilution, rather than geometric distortion, is the governing bottleneck for fine-grained retrieval: As evidenced in the ablation table, the Single-Vector Cubemap variant (ii) completely eliminates ERP distortion but achieves only 18.95% Local Caption R@1 (comparable to 17.33% for raw ERP). Transitioning to the Multi-Vector Cubemap representation (iii) propels Local Caption R@1 to 74.78%, confirming that single-vector compression is the root cause of local semantic failure.
- Global scene context and localized details provide mutual reinforcement: Introducing the downsampled global panorama in the full model (iv) not only boosts Global Caption R@1 from 71.91% to 79.68%, but also yields additional improvements in Local Caption R@1 (74.78% to 76.55%). Global context provides anchor topology for local perspectives, while rich local features enhance global discriminability.
- Forced feature pooling is catastrophic for localized matching: Replacing MaxSim routing with Mean Pooling drops Local Caption R@1 by 60.30% (from 76.55% down to 16.25%), and Attention Pooling only manages 23.29%. Dynamic routing is strictly superior to static or soft weighted fusion for panoramic multi-granularity retrieval.
- Broad generalization across VLM families and zero-shot transfers: Across eight distinct foundation architectures (ranging from SmolVLM-256M to InternVL3.5-8B), PanoRec consistently delivers substantial performance boosts. In zero-shot evaluation on Matterport3D without domain fine-tuning, PanoRec-8B achieves 82.78% Local Caption R@1, dramatically outperforming the direct training baseline of 17.13%.
Highlights & Insights¶
- Non-invasive sequence modeling paradigm: Rather than designing complex spherical attention layers or graph networks, PanoRec elegantly utilizes the native interleaved sequence modeling capabilities of modern VLMs, extracting decoupled multi-granularity representations using only seven spatial anchor tokens.
- Dual utility of MaxSim routing: The MaxSim operation serves as a dynamic noise filter during inference and seamlessly transforms into an in-batch hard spatial negative miner during contrastive optimization, preventing shortcut learning without additional computation.
- Lightweight multi-vector indexing: By compacting the entire 360-degree environment into six local vectors and one global vector, PanoRec offers an efficient indexing strategy for robotic navigation and spatial memory retrieval, avoiding the storage explosion of dense token grids.
Limitations & Future Work¶
- Boundary truncation across planar faces: The 90-degree field of view of standard cubemap projections can split objects situated across face seams, potentially causing incomplete semantic extraction for queries focused on boundary landmarks.
- Fixed orientation ordering versus rotational invariance: The spatial anchor sequence enforces a predetermined scanning order (front, right, back, left), which may exhibit slight sensitivity to continuous yaw rotations of the physical camera.
- Resolution ceiling on distant micro-objects: Constraining each cubemap face to 280x280 pixels caps visual acuity on fine distant targets; integrating adaptive multi-scale cropping could further enhance detection of miniature objects.
Related Work & Insights¶
- vs Spherical Operators & Deformable Attention (e.g., PanoSwin): Traditional methods introduce custom spherical convolution or attention layers that disrupt pre-trained VLM weights; PanoRec maintains strict architectural compatibility with standard 2D transformers.
- vs Dense Patch Matching (e.g., Dense360): Dense patch indexing requires thousands of token vectors per scene and incurs substantial search latency; PanoRec achieves comparable or superior granularity using a compact 7-vector footprint.
- vs Unordered Bag-of-Views Aggregation: Conventional multi-view techniques discard spatial topology and fuse features into a single vector; PanoRec preserves structured geometric sequence context while keeping local representations decoupled.
Rating¶
- Novelty: โญโญโญโญโญ Formulates the semantic dilution problem in panoramic retrieval and provides an elegant, non-invasive sequence modeling solution.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive cross-architecture validation across 8 VLM backbones, multiple indoor/outdoor datasets, zero-shot benchmarks, and detailed ablations.
- Writing Quality: โญโญโญโญโญ Lucid narrative flow, clear formulation, and solid empirical justifications.
- Value: โญโญโญโญโญ Offers an efficient and practical blueprint for multi-granularity spatial perception, embodied agent memory, and panoramic retrieval.