MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling¶
Conference: ECCV 2026
Paper: ECCV 2026 Oral/Poster
Code: https://co1dspring.github.io/MV-STRIDE/
Area: Multimodal VLM / 3D Vision
Keywords: Multi-View Spatial Reasoning, Multimodal Large Language Model, Hierarchical Capability Modeling, Chain-of-Thought Distillation, Reinforcement Learning
TL;DR¶
Addressing the lack of coherent 3D cognitive maps in current multimodal LLMs under multi-view observations, MV-STRIDE introduces a hierarchical capability modeling framework with cross-view dependency constraints, coupled with cognition question groups and a three-stage training pipeline (SFT-ColdStart-RL), boosting an 8B open-source model to 38.9% SOTA on MMSI-Bench.
Background & Motivation¶
In recent years, Multimodal Large Language Models (MLLMs) have made remarkable strides in 2D visual comprehension and basic single-perspective perception tasks. However, robust multi-view spatial reasoningโa fundamental requirement for embodied intelligence, robotics, and autonomous drivingโremains a severe bottleneck. Comprehensive evaluations on benchmarks like MMSI-Bench and ViewSpatial-Bench reveal a persistent, glaring performance gap between humans and MLLMs. Humans naturally consolidate localized objects and spatial layouts from disparate perspectives into a cohesive internal 3D cognitive map, effortlessly performing camera pose transformations and deriving spatial relationships. In contrast, existing multimodal models frequently fail in establishing cross-view correspondence, preserving 3D geometric consistency, and conducting stable viewpoint-dependent inference; diagnostic error analyses show that these failures predominantly stem from deficiencies in scene-level reconstruction and view transformation rather than isolated single-image recognition errors.
The root cause of this multi-view bottleneck lies in the design flaws of existing spatial reasoning datasets. The vast majority of datasets focus on single-view 2D planes, failing to foster the emergence of 3D spatial awareness. While several recent datasets introduce multi-view inputs, they predominantly focus on low-level quantitative perception (e.g., depth estimation or pose regression), falling short of cultivating structured high-level reasoning. More critically, prior datasets almost universally treat multi-view spatial tasks as flat, loosely defined capability collections, lacking explicit dependency transitions between foundational perception, scene modeling, and contextual reasoning. Consequently, models exploit local 2D appearance shortcuts, lacking a principled learning trajectory from basic geometry to complex, multi-step spatial deduction.
The angle of attack in this work is to mirror the progression of human spatial cognition, explicitly decomposing multi-view spatial reasoning into an interdependent hierarchical capability system with geometric safeguards against single-view shortcuts. The core idea is to introduce MV-STRIDE, organizing multi-view spatial reasoning into a three-level hierarchy ("Single-View Spatial Perception (Level I) โ 3D-Consistent Scene Understanding (Level II) โ Multi-View Contextual Reasoning (Level III)"), strictly enforcing cross-view dependency constraints and leveraging a three-stage training strategy (SFT, CoT cold-start, and GRPO reinforcement learning) to guide models toward internalizing authentic 3D spatial consistency.
Method¶
Overall Architecture¶
MV-STRIDE addresses multi-view spatial reasoning by explicitly modeling the dependencies between different spatial capabilities. The overall architecture encompasses two core pillars: automated hierarchical data generation and progressive multi-stage training. In the data generation phase, leveraging procedural 3D environments from Infinigen and high-fidelity real-world reconstructions from ScanNet++, the pipeline adopts a top-down decomposition strategy: it first synthesizes complex Level III contextual tasks governed by single-view unsolvability constraints, and subsequently back-engineers the necessary Level I and Level II prerequisite sub-questions. These multi-level tasks are organized into spatial cognition process question groups, which feed into an LLM synthesizer to produce verifiable Chain-of-Thought (CoT) supervision. In the training phase, the base model (Qwen3-VL-8B-Instruct) undergoes a three-stage progressive alignment consisting of foundational multi-task SFT, CoT cold-start tuning, and multi-view GRPO reinforcement learning.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}, 'subGraphTitleMargin': {'top': 8, 'bottom': 16}}}%%
flowchart TD
A["Multi-Source 3D Scene Input<br/>Infinigen Procedural Gen + ScanNet++ Real Reconstruction"] --> B["Hierarchical Capability & Question Group Generation<br/>Level I Perception / Level II Scene / Level III Reasoning"]
B --> C["Cross-View Dependency Constraint Filtering<br/>Overlap Anchor Selection + Shortcut Elimination"]
C --> D["Cognition-Grounded CoT Synthesis<br/>Sub-question Prerequisite Integration + LLM Refinement"]
D --> E["Progressive Three-Stage Training Strategy<br/>Stage 1 SFT โ Stage 2 Cold-Start โ Stage 3 GRPO"]
E --> F["Output: 3D-Consistent Spatial Reasoning<br/>Accurate Camera Pose Estimation + Cross-View Relational Reasoning"]
Key Designs¶
1. Hierarchical Capability & Question Group Generation: Structuring Cognitive Dependencies
To resolve the flat, disconnected task structures in prior datasets, this paper decomposes spatial reasoning into a structured, three-tiered cognitive hierarchy. Level I establishes egocentric single-view geometric perception, including 2D/3D bounding boxes, camera pitch, absolute depth, and camera-object yaw offsets, preventing inaccurate local perception from corrupting subsequent 3D modeling. Level II focuses on 3D-consistent scene understanding across two viewpoints, requiring cross-view object correspondence matching and 6-DoF relative camera motion estimation (both rotation and translation) to bridge disjoint views into a unified cognitive map. Level III addresses multi-view contextual reasoning, spanning object counting, size comparison, object orientation, and cross-reference-frame spatial positioning (e.g., object-to-object, camera-to-object, camera-to-region). Starting from Level III as the apex, the framework reverse-engineers the necessary perceptual and scene-modeling dependencies, instantiating them as Level I and Level II sub-questions. Together, tasks derived from the same scene and camera set form a coherent "Spatial Cognition Process Question Group", ensuring that high-level reasoning is strictly anchored upon verifiable geometric building blocks.
2. Cross-View Dependency Constraint Filtering: Eliminating Single-View Shortcuts
Because multi-view reasoning datasets are vulnerable to single-view heuristic shortcuts, the authors institute rigorous filtering rules during camera and object selection. First, camera pairs are filtered based on spatial proximity and orientation; viewpoint pairs with negligible or excessive displacement are removed, while requiring a minimum number of shared co-visible objects to serve as reliable spatial anchors. Second, all Level III tasks must strictly satisfy the cross-view dependency constraint: queries are designed such that they cannot be answered from any individual perspective. This is achieved by either evaluating relationships between entities that never co-occur in the same frame, or referencing an object visible only in View B relative to the camera coordinate frame of View A. Empirical validation confirms that on 1K sampled Level III queries, Gemini-3-Flash achieves only 31.5% accuracy with a single-view input versus 41.2% with multi-view inputs, demonstrating that single-view shortcuts are effectively blocked.
3. Cognition-Grounded CoT Synthesis: Reverse Logic Distillation
Standard CoT generation approaches rely on unconstrained LLM hallucinations, resulting in geometrically ungrounded reasoning traces. MV-STRIDE leverages the explicit dependency graph within each cognition question group: the ground-truth answers of Level I (object detection, angular offsets) and Level II (camera translation, rotation) serve as deterministic intermediate steps for Level III goals. An advanced LLM is employed purely as a linguistic synthesizer, threading these exact 3D ground-truth facts into a structured, three-stage natural-language narrative: Single-View Perception โ Cross-View 3D Modeling โ High-Level Contextual Reasoning. This automated pipeline generates scalable, highly transparent, and interpretable reasoning supervision without manual annotation, achieving 88.0% accuracy on QA pairs and 97.5% on CoT annotations under human/LLM verification.
4. Progressive Three-Stage Training Strategy: From Imitation to Active 3D Consistency
To enable models to systematically acquire multi-view reasoning proficiency, training is structured across three stages: Stage 1 conducts foundational multi-task SFT on 60% of scene data (313.8K QA pairs across Levels IโIII plus open-source spatial datasets), teaching Qwen3-VL-8B-Instruct direct spatial perception and answering formats. Stage 2 performs CoT cold-start training on 19.0K multi-stage reasoning chains constructed from ScanNet++, establishing the structured cognitive template for long-chain deduction. Stage 3 implements reinforcement learning via Group Relative Policy Optimization (GRPO) on 24.8K high-level reasoning samples with moderate difficulty (8 rollouts per prompt). This RL exploration empowers the model to verify spatial hypotheses dynamically, progressing beyond superficial imitation of SFT templates to independently formulate internal 3D coordinates and camera transformations.
Key Experimental Results¶
Main Results¶
Evaluated on four prominent spatial reasoning benchmarks (MMSI-Bench, ViewSpatial-Bench, 3DSR-Bench, and CV-Bench), MV-STRIDE significantly outperforms existing open-source baselines and rivals top proprietary models:
| Model | Size | MMSI-Bench (Avg) | ViewSpatial-Bench (Avg) | 3DSR-Bench (Avg) | CV-Bench (Avg) |
|---|---|---|---|---|---|
| GPT-5 [34] | Proprietary | 41.90% | - | 66.70% | - |
| GPT-4o [18] | Proprietary | 32.10% | 34.98% | 60.30% | 78.90% |
| Gemini2.5-Pro [9] | Proprietary | 37.60% | - | 64.30% | - |
| InternVL3-78B [50] | 78B | 28.90% | - | 61.30% | - |
| Qwen2-VL-72B-Instruct [37] | 72B | 31.50% | - | 57.50% | - |
| InternVL3-38B [50] | 38B | 29.00% | - | 59.10% | - |
| Qwen2.5-VL-7B-Instruct [4] | 7B | 26.80% | 36.85% | 53.20% | 73.00% |
| Base Model (Qwen3-VL-8B) [3] | 8B | 29.20% | 40.51% | 59.98% | 84.31% |
| MV-STRIDE-SFT (Stage 1) | 8B | 37.20% | 50.28% | 65.52% | 86.39% |
| MV-STRIDE-RL (Stage 3) | 8B | 37.50% | 49.30% | 59.43% | 85.94% |
| MV-STRIDE-SFT (Full) | 8B | 38.90% | 48.35% | 64.51% | 86.73% |
On MMSI-Bench, MV-STRIDE-SFT (Full) achieves 38.90%, yielding a remarkable 9.7% gain over the base model, outperforming the 78B parameter InternVL3 and closed-source GPT-4o (32.10%). On specific challenging sub-categories, such as Camera Motion (45.95% vs. GPT-5's 32.40%) and Camera-Object Positional Relationship (66.28% vs. GPT-5's 48.80%), it establishes new state-of-the-art records across all models. Furthermore, MV-STRIDE-SFT (Stage 1) attains the highest score overall on ViewSpatial-Bench (50.28%) and the best open-source score on 3DSR-Bench (65.52%).
Ablation Study¶
To evaluate the complementary nature of procedural synthetic environments and real-world captures, the authors analyze the impact of different Level III training data sources on MMSI-Bench:
| Data Source Config | # Infinigen | # ScanNet++ | MMSI-Bench (Acc) | Note |
|---|---|---|---|---|
| Base Model | - | - | 29.20% | Unfine-tuned baseline |
| Synthetic Only | 99.3k | - | 36.70% | High precision geometry, procedural domain gap |
| Real-world Only | - | 99.5k | 37.50% | Realistic textures, bounded spatial diversity |
| Hybrid | 49.7k | 49.8k | 38.00% | Total ~100k samples, exceeds either single source |
| Unified | 99.3k | 99.5k | 39.40% | Optimal performance through mutual synergy |
Key Findings¶
- Synthetic vs. Real-World Synergy: Procedural generation (Infinigen) supplies noise-free 3D coordinates and camera poses, establishing solid geometric foundations, while real-world reconstructions (ScanNet++) inject complex photorealistic visual distributions. Mixing them in equal parts (Hybrid) outstrips either source alone, and scaling both (Unified) elevates MMSI-Bench accuracy to 39.40%.
- Multi-Stage Necessity and the "Reasoning Tax": Omitting Stage 1 leaves the model without a viable geometric basis, causing RL in Stage 3 to diverge; omitting Stage 2 causes long-chain logic breakdown during RL. Transitioning from direct short-answer SFT (Stage 1) to CoT generation (Stage 2/3) incurs a slight "reasoning tax" on direct single-view metrics, but Stage 3 enables the model to actively calculate explicit camera translations (e.g., 0.85m shift) and angular rotations (e.g., 30.5ยฐ), achieving true white-box spatial deduction.
Highlights & Insights¶
- Prerequisite-Conditioned CoT Generation: Rather than prompting LLMs to hallucinate intermediate steps, MV-STRIDE utilizes verified Level I/II ground truths as logical milestones, leveraging the LLM purely for natural-language synthesis to produce scalable, factually grounded 3D CoT data.
- Cognitive 3D Map Emergence: Instead of treating multi-view spatial QA as a black-box pattern matching task, MV-STRIDE induces the autonomous emergence of a structured multi-view deduction paradigm: local detection โ relative camera pose calculation โ allocentric global reasoning.
Limitations & Future Work¶
- Static Scene Limitation: The current benchmark and training data operate exclusively on static indoor scenes, leaving dynamic environments with moving objects or changing illumination for future investigation.
- Inference Latency Overhead: The structured multi-stage CoT introduces significant token expansion, posing computational latency challenges in high-frequency real-time robotic control or autonomous driving settings.
Related Work & Insights¶
- vs. SpatialVLM / SpatialRGPT: Earlier approaches relied on single-view depth or 3D bounding box pseudo-labels; MV-STRIDE explicitly introduces multi-view rigid transformation constraints and enforces cross-view dependencies to eradicate single-view shortcuts.
- vs. MultiSPA / ViewSpatial: Previous multi-view benchmarks either treated tasks as isolated flat categories or focused heavily on low-level perceptual estimation; MV-STRIDE is the first to establish a top-down capability hierarchy uniting perception, modeling, and reasoning.
Rating¶
- Novelty: โญโญโญโญโญ Pioneering hierarchical capability modeling and cross-view dependency constraints for multi-view spatial reasoning.
- Experimental Thoroughness: โญโญโญโญโญ Evaluated across four diverse benchmarks, comparing numerous open-source and proprietary MLLMs with rigorous data and stage ablations.
- Writing Quality: โญโญโญโญโญ Exceptionally clear structure, rigorous formulation, and strong grounding in cognitive principles.
- Value: โญโญโญโญโญ Sets new state-of-the-art performance for open-source 8B models, delivering a scalable paradigm for 3D-aware embodied AI.