MonoSR: Open-Vocabulary Spatial Reasoning on Monocular Images¶
Conference: ECCV 2026
arXiv: 2511.19119
Code: None
Area: Multimodal VLM
Keywords: Monocular Spatial Reasoning, Open-World, VLM Benchmark, Spatial VQA, Single-view Answerability
TL;DR¶
MonoSR constructs the first large-scale open-world monocular spatial reasoning dataset (1M+ QA pairs, covering indoor, outdoor, and object-centric scenes with 98 semantic categories). It ensures that every QA pair is answerable from a single RGB image through a four-stage observability filtering mechanism, and systematically evaluates the impacts of seven open/closed-source VLMs as well as various auxiliary signals on spatial reasoning capabilities.
Background & Motivation¶
Background: Spatial Reasoning (SR) is a critical capability for applying VLMs to physical world scenarios such as embodied AI. However, existing spatial reasoning datasets and benchmarks almost exclusively rely on multi-view images, video sequences, or point clouds. These inputs implicitly provide geometric cues, sparing the model from inferring 3D structures from a single image. Meanwhile, existing benchmarks are heavily restricted to indoor scenes and focus on low-level perceptual tasks, lacking coverage of high-level spatial imagination and open-world generalization.
Limitations of Prior Work: Constructing monocular spatial reasoning datasets faces three fundamental barriers. First, spatial quantities must be defined relative to the camera coordinate system rather than the absolute world coordinate system; otherwise, the answers are underdetermined under single-view constraints. Second, a large number of unanswerable questions caused by occlusion, truncation, and annotation noise must be systematically identified and excluded. Third, annotating perspective-aware imagination tasks requires rigid geometric transformations rather than direct image observation, making data generation far more complex than existing multi-view schemes.
Why This Is Feasible Now: The emergence of large-scale monocular 3D detection datasets, such as Omni3D, provides high-quality multi-domain 3D annotations (230K images covering indoor, outdoor, and object-centric scenes). Meanwhile, the maturity of LLMs allows for automated and diverse paraphrasing of question texts to eliminate template biases. These two enabling factors make the construction of large-scale single-view spatial reasoning datasets feasible for the first time.
Design Choices of This Work: MonoSR is built directly upon the 3D annotations and RGB images of Omni3D, ensuring data quality through a rigorous pipeline: "Scene Preprocessing \(\rightarrow\) Scene Graph Construction \(\rightarrow\) Two-Stage QA Generation (Deterministic Derivation + LLM Paraphrasing) \(\rightarrow\) Four-Stage Observability Filtering". The tasks are structured into three levels according to human cognitive hierarchy (Foundational Perception, Perspective-Aware Imagination, Situational Reasoning), with each level designed with explicit geometric constraints and answerability guarantees. Building upon this, the paper systematically injects auxiliary information (scene category, 2D visual prompts, 3D bounding boxes) to quantify the geometric information gap of current VLMs in monocular spatial reasoning.
Core Idea: By deterministically deriving QA answers from ground-truth 3D annotations and applying strict observability filtering, this work constructs a large-scale monocular spatial reasoning dataset that covers indoor, outdoor, and object-centric scenes to support both training and evaluation. Based on this, it systematically quantifies the capability boundaries and key bottlenecks of current VLMs in monocular spatial reasoning.
Method¶
Overall Architecture¶
The overall construction pipeline of MonoSR consists of four stages. The inputs are RGB images and their corresponding 3D bounding box annotations from Omni3D; the outputs are over 1M QA pairs, each guaranteed to be answerable from a single RGB image.
The first stage is Scene Preprocessing and Filtering: Strict quality filtering is conducted on the raw Omni3D scenes, excluding instances with excessive truncation (\(>0.4\)), excessive occlusion (\(>0.6\)), or too small projection areas (\(<400\text{ px}^2\)). Geometrically inconsistent or degenerate 3D bounding boxes (e.g., dimensions \(>20\text{m}\), near-zero volume, 2D projection deviation from 3D footprint \(>30\%\)) are discarded. The number of valid instances per image is controlled between 3 and 40 to strike a balance between richness and noise.
The second stage is Scene Graph Construction: Fine-grained descriptive captions are generated for each salient object in the filtered scenes, and 3D relationships (including spatial relations, relative positions, and size comparisons) among objects are parsed to construct structured scene graphs as the foundation for subsequent QA generation.
The third stage is QA Data Generation: A two-stage hybrid strategy is adopted. In the first stage, all answers are deterministically derived from 3D annotations—answers for foundational perception tasks are computed directly from object coordinates, while perspective-aware imagination tasks obtain answers through rigid-body transformations and geometric consistency checks. These answers are then paired with questions using predefined templates. In the second stage, LLMs are utilized to paraphrase and diversify the question texts, preserving correct answers while generating a more natural linguistic style to eliminate template biases.
The fourth stage is Single-View Observability Filtering: A fourth-level quality check is performed on the generated QA pairs to exclude questions whose answers are visually indistinguishable (such as depth-ambiguous pairs with projected pixel distances \(<20\text{px}\) on the image), ensuring that remaining unanswerable questions in the final dataset are below \(1.4\%\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Omni3D<br/>RGB + 3D Annotations"] --> B["Scene Preprocessing & Filtering<br/>Visibility + Annotation Quality + Scene Complexity"]
B --> C["Scene Graph Construction<br/>Descriptive Captions + 3D Relation Parsing"]
C --> D["Two-Stage QA Generation<br/>Deterministic Derivation + LLM Paraphrasing"]
D --> E["Single-View Observability Filtering<br/>Visual Ambiguity Removal"]
E --> F["MonoSR Dataset<br/>1M+ QA Pairs"]
Key Designs¶
1. Three-Level Cognitive Task System: From Foundational Perception to Situational Reasoning
MonoSR designs 8 spatial reasoning tasks organized into three levels based on the human cognitive hierarchy to systematically evaluate monocular spatial understanding. The first level, Foundational 3D Perception (52% of QA pairs), evaluates basic geometric and spatial properties. It includes four subtasks: Spatial Relation (SR), Size Estimation (Size), Distance Measurement (Dist), and Dimension Comparison (Dim), where answers are deterministically calculated from object 3D coordinates. The second level, Perspective-Aware Imagination (38%), evaluates the ability to answer questions from unobserved viewpoints. It includes three subtasks: Occlusion Judgment (OJ), Perspective Relation (PR), and Object Grounding (OG)—requiring the model to possess implicit perspective imagination and spatial consistency retention capabilities. Answers are obtained via rigid-body transformations and geometric consistency checks. The third level, Situational Reasoning (10%), simulates real-world scenarios, requiring the model to integrate visual scenes with common-sense knowledge for complex reasoning. The solution assigns expert roles (e.g., safety officer, roboticist) to LLMs and provides motivational clauses to anchor questions in realistic operational contexts. Each task supports three formats: Yes/No, Multiple-Choice, and Numerical. This hierarchical design is used not only for training but also for fine-grained diagnostic evaluation—by comparing model performance across different levels, capability deficiencies can be precisely pinpointed.
2. Four-Stage Single-View Answerability Guarantee: The Quality Cornerstone of Data
This is the most critical design constraint of MonoSR—each QA pair must be answerable solely from a single RGB image. This guarantee is enforced progressively through a four-stage filtering pipeline: (i) instance visibility—truncation \(< 0.4\), occlusion \(< 0.6\), and projected area \(> 400\text{ px}^2\), eliminating at the source instances that cannot be judged due to viewpoint issues; (ii) 3D annotation quality—discarding degenerate bounding boxes (dimensions \(>20\text{m}\), near-zero volume, 2D-3D projection deviation \(>30\%\)) to avoid incorrect answers caused by annotation noise; (iii) scene complexity—maintaining 3 to 40 valid instances per image, as too low complexity lacks reasoning material while too high complexity introduces uncontrollable annotation noise; (iv) QA ambiguity—removing questions whose answers cannot be distinguished in 2D images (e.g., foreground and background objects where the depth difference corresponds to a projected pixel distance of \(<20\text{px}\)). Through manual auditing and validation, the remaining unanswerable questions constitute less than \(1.4\%\). This mechanism is the key structural difference between MonoSR and indoor multi-view benchmarks like ScanQA and SQA3D—which do not enforce this constraint, meaning their QA pairs may rely on multi-view contrastive information or depth cues, thereby failing to fairly evaluate monocular vision capabilities.
3. Auxiliary Information Injection Analysis Framework: Methodology for Quantifying the Geometric Gap
MonoSR is not only a dataset but also offers a systematic experimental paradigm to quantify the capability boundaries of current VLMs in monocular spatial reasoning. The concrete scheme injects three types of auxiliary signals into VLM inputs in a controlled manner to compare performance differences under different configurations: (a) Scene Information (SI)—informing the model of the input image's scene type (indoor/outdoor/object-centric) via text to provide high-level contextual priors; (b) 2D Visual Prompts (2D VP)—marking the 2D bounding boxes of detected objects on the image (input via visual channels) to provide local 2D spatial anchors; (c) 3D Bounding Boxes (3D Bbox)—providing the 3D center coordinates, sizes, and orientations of objects in textual format (serving as an oracle upper bound for geometric information—not a practical solution, but a controlled probe). The core insight of this framework is: by comparing the gap between "MonoSR fine-tuning alone" and "+3D Bbox", one can precisely quantify the amount of geometric information missing in current monocular VLMs, which defines the exact target that future monocular perception modules (e.g., metric depth estimation, monocular 3D detection) need to bridge. Experiments are verified across two different backbone architectures (Qwen-2.5-VL-3B and InternVL-3.5-2B) to ensure the consistency of trends across model families.
An Illustrative Example: Indoor Scene QA Generation¶
Taking an indoor image containing "a chair to the left of a table" as an example, the QA generation pipeline of MonoSR is as follows:
-
Scene Preprocessing: The Omni3D annotations record the 3D bounding box of the chair as \((x=1.2, y=0.5, z=2.3, w=0.6, h=0.9, l=0.6)\) and the table as \((x=2.8, y=0.4, z=2.1, w=1.2, h=0.8, l=1.0)\). Both objects have truncation \(<0.4\), occlusion \(<0.2\), and projected area \(>400\text{ px}^2\), with no geometric degeneracies—passing stage (i) and (ii) filters. The number of valid instances in the image is 15—passing the stage (iii) filter.
-
Scene Graph Construction: The chair is described as "a brown wooden chair on the left side of the room"; the table is described as "a white square table situated in the center-right of the room"; the scene graph records the relationship "\(chair\ x=1.2 < table\ x=2.8 \rightarrow chair\ is\ to\ the\ left\ of\ the\ table\)".
-
Stage 1 QA Generation: The spatial relationship between the two objects is computed from their 3D coordinates—comparing the x-coordinates (\(1.2 < 2.8\)), yielding the deterministic answer "left". A template-based question is generated: "Is the chair to the left or to the right of the table?" + answer "left".
-
Stage 2 LLM Paraphrasing: The LLM paraphrases the question to: "From the camera's perspective, on which side of the table is the chair in the room located?" while keeping the answer unchanged, rendering a more natural linguistic style.
-
Stage 4 Filtering: The distinguishability of the question on the 2D image is checked—the horizontal distance of the chair and table in the 2D projection is about \(80\text{px} > 20\text{px}\) threshold, which is visually clear and answerable—passed.
-
Insertion: The QA pair is formatted as a Multiple-Choice task (options: left/right/front/behind), annotated with the cognitive level "Foundational Perception - Spatial Relation", and entered into the training or evaluation set.
Loss & Training¶
In the auxiliary information experiments, MonoSR performs supervised fine-tuning on Qwen-2.5-VL-3B and InternVL-3.5-2B backbones under identical auxiliary input configurations. Each configuration only varies the injected auxiliary signals while keeping hyperparameters like batch size, learning rate, and training epochs consistent. Numerical questions are evaluated using an adaptive threshold: a predicted value is considered correct if its relative error is \(|\frac{\hat{d}-d}{d}| < 10\%\). A sensitivity analysis is also performed at a stricter \(5\%\) threshold, confirming that model rankings are robust to the threshold choice. Across all configurations, the remaining unanswerable questions on the validation set are \(<1.4\%\), ensuring that evaluations are not biased by data noise.
Key Experimental Results¶
Main Results: Comparison with Existing Spatial Reasoning Benchmarks¶
| Dataset | Indoor | Outdoor | Object-Centric | Images | QA Pairs | Monocular-Only | Training Support |
|---|---|---|---|---|---|---|---|
| ScanQA | ✓ | ✗ | ✗ | 800 | 41K | ✗ | ✓ |
| SQA3D | ✓ | ✗ | ✗ | 650 | 33K | ✗ | ✓ |
| EmbSpatial-Bench | ✓ | ✗ | ✗ | ~2K | 3.6K | ✗ | ✗ |
| SPHERE | ✓ | ✗ | ✗ | ~3K | ~6K | ✓ | ✗ |
| OmniSpatial | ✓ | ✗ | ✗ | 6.5K | 8.4K | ✓ | ✗ |
| VSI-Bench | ✓ | ✗ | ✗ | 288 | 5K | ✗ | ✗ |
| SpatialBot | ✓ | ✓ | ✗ | 743K | 743K | ✗ | ✓ |
| InternSpatial | ✓ | ✓ | ✓ | — | 12M | ✗ | ✓ |
| MonoSR (Ours) | ✓ | ✓ | ✓ | 230K | 1M+ | ✓ | ✓ |
MonoSR is the only large-scale, trainable dataset that covers all three scenarios—indoor, outdoor, and object-centric—and is specifically designed for monocular inputs. Unlike purely diagnostic benchmarks (such as SPHERE 6K and OmniSpatial 8.4K), MonoSR's scale (1M+ QA pairs) enables large-scale training and open-world evaluation of monocular spatial reasoning for the first time.
Ablation Study: Impact of Auxiliary Information on Spatial Reasoning¶
| Configuration | Indoor Overall | Outdoor Overall | Object-Centric Overall |
|---|---|---|---|
| Qwen-2.5-VL-3B baseline (w/o FT) | 0.283 | 0.302 | 0.052 |
| + MonoSR FT (w/o Aux) | 0.494 | 0.442 | 0.335 |
| + SI (Scene Info) | 0.566 | 0.558 | 0.381 |
| + 2D VP (2D Visual Prompt) | 0.585 | 0.576 | 0.437 |
| + 3D Bbox (3D Bbox, oracle upper bound) | 0.722 | 0.723 | 0.986 |
| + 2D VP + 3D Bbox (optimal combination) | 0.798 | 0.768 | 0.996 |
| + 2D VP + 3D Bbox (all objects) | 0.767 | 0.732 | 0.988 |
Key Findings¶
- MonoSR Fine-tuning Alone Significantly Boosts Performance: Qwen-2.5-VL-3B improves from 0.283 to 0.494 (\(+74.6\%\)) in indoor scenes, and InternVL-3.5-2B improves from 0.301 to 0.483. This demonstrates that MonoSR provides effective spatial supervision signals, activating latent reasoning capabilities already present in foundational models.
- 3D Bbox is the Largest Contributor: Adding 3D bounding boxes leads to a near-perfect performance jump (from 0.335 to 0.986 in object-centric scenes), indicating that the core bottleneck of current VLMs lies in their inability to recover fine-grained 3D geometric information from a single image, rather than in object recognition or semantic understanding.
- More Auxiliary Information is Not Always Better: Providing a "full" configuration of 2D VP + 3D Bbox for all detected objects in the scene actually degrades performance compared to providing it only for target objects (indoor dropping from 0.798 to 0.767). This suggests that irrelevant signals distract the model's attention, implying that future monocular perception modules should focus on selective geometric signal injection rather than full-scale inputs.
- Object-Centric Scenes are the Achilles' Heel of All VLMs: General VLMs (LLaVA-OV-72B drops to 0.087, ChatGPT-4 drops to 0.069) experience performance collapses in numerical estimation within object-centric scenes. In contrast, SpatialVLM (0.618) establishes a significant lead owing to its geometry-aware training pipeline, proving that precise metric reasoning is currently the weakest link for VLMs.
Highlights & Insights¶
- Single-View Answerability Guarantee is the Core Soul of MonoSR: The four-stage filtering mechanism fundamentally solves the core question of "whether this question can truly be answered from a single image," making MonoSR the only spatial reasoning dataset currently supporting true monocular evaluation. Many existing benchmarks claim to be "monocular" but feature questions and answers that actually rely on contrastive cues across multiple views.
- The Elegance of the 3D Oracle Experimental Framework: By treating 3D bounding boxes as the "upper bound" of geometric information, 2D visual prompts as "2D anchors", and scene info as "contextual priors," the framework allows the quantification of "how far current VLMs are from perfect geometric understanding." This quantitative guidance holds much greater engineering and practical value than general claims like "VLM performance is poor."
- The Counter-Intuitive Finding of "Selective over Full": Providing 3D bounding boxes for all objects degrades performance, proving that attention mechanisms are not immune to irrelevant information. This discovery offers direct engineering guidance for the design of future monocular perception modules.
- Validation Across Model Families Enhances Credibility: Repeating the auxiliary information experiments on two backbones with distinct architectures (Qwen and InternVL) yields consistent trends, showing that the main conclusions are not specific to a particular model.
Limitations & Future Work¶
- Solely Covering Static Scenes: MonoSR currently does not handle reasoning under viewpoint uncertainty (e.g., predicting spatial relations from new or occluded viewpoints that extend beyond perspective-aware geometric transformations). The authors suggest that combining neural rendering or scene completion as the data generation backbone could bridge this gap.
- The 3D Oracle Gap Remains Unfilled in Practical Systems: The auxiliary information experiments reveal the huge advantage of 3D bounding boxes acting as a geometric upper bound, but achieving performance near this upper bound in practical monocular systems remains an open challenge. The authors plan to integrate metric depth estimation and monocular 3D detection modules directly into VLM architectures, selectively injecting geometric signals using the auxiliary analysis results as a blueprint.
- RL-based Post-training is a Promising Unexplored Direction: Improving spatial reasoning capabilities without additional annotation data through reinforcement learning-based post-training is a path worth exploring. The authors have explicitly planned to conduct relevant work on the MonoSR training set.
- The "Dimensional Curse" of Object-Centric Scenes Awaits Breakthroughs: Performance in object-centric scenes relies heavily on the injection of 3D bounding boxes; once the oracle signals are removed, almost all models perform poorly. This indicates that precisely estimating the dimensions and distance of a single object from a single viewpoint is inherently challenging, potentially requiring new inductive biases (e.g., size priors, category-level physical constraints) to achieve breakthroughs.
Related Work & Insights¶
- vs SpatialVLM / VLM3R / SpatialBot: These methods or datasets are designed for multi-view or depth-enhanced inputs rather than prioritizing monocular settings. MonoSR fills the critical gap of spatial reasoning datasets explicitly constrained to monocular inputs. Its direct comparative value lies in revealing the vast performance chasm between "multi-view" and "monocular" settings.
- vs SPHERE / OmniSpatial: Although these two benchmarks are monocular-specific, their scales are extremely small (\(<10\text{K}\) QA, indoor only), making them suitable only for evaluation rather than training. MonoSR overcomes this shortcoming with a scale that is over 100 times larger, rendering the "train on monocular data \(\rightarrow\) evaluate on open world" pipeline a reality.
- vs SeeGround / Perception-Aware Reasoning: These methods address monocular spatial reasoning from the algorithmic side (e.g., rendering query-aligned views, explicit object-aware visual grounding), whereas MonoSR provides the infrastructure for training and evaluation from the data side. The two are naturally complementary—training models with MonoSR and implementing reasoning pipelines via these algorithmic methods is a highly promising direction.
- vs Qwen-2.5-VL / InternVL / Gemini: The systemic deficiencies of general VLMs in spatial reasoning are precisely localized via MonoSR's multidimensional evaluation to the weakest link: "metric estimation from a single object-centric view." This provides a clear target for the next capability breakthrough of VLMs.
Rating¶
- Novelty: ⭐⭐⭐⭐ [First large-scale open-world monocular spatial reasoning dataset. Both the four-stage observability filtering and the auxiliary information analysis framework possess originality, though the overall contribution falls under data and evaluation infrastructure rather than methodology.]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Spans 7 SOTA VLMs, 3 scene domains, 3 cognitive tiers, threshold sensitivity analysis, long-tail classification evaluation, and multi-combination ablations on auxiliary information. The experimental design is systematic and comprehensive.]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, detailed designs, in-depth experimental analyses, and well-supported figures and tables. Structure is coherent and the arguments are rigorous.]
- Value: ⭐⭐⭐⭐⭐ [Fills a critical gap in large-scale data for monocular spatial reasoning, directly benefiting key application fields like embodied AI and autonomous driving. In the long term, it is likely to act as a catalyst similar to VQA-v2 for visual question answering.]