Kilometer-Vision: A New Frontier for Large-Scale Spatial Awareness in VLMs¶
Conference: ECCV 2026
Paper: ECCV Paper
Project: https://perception-test-challenge.github.io/kilometervision.html
Area: Multimodal VLM
Keywords: spatial intelligence, city-scale, landmarks, routes, cognitive maps
TL;DR¶
KilometerVision introduces the first benchmark probing city-scale spatial intelligence in Vision-Language Models from real-world walking tour videos spanning up to 1km and 10 minutes, using a cognitive landmark-route-map framework across 1000 multiple-choice QAs to demonstrate that state-of-the-art models circumvent genuine geometric reasoning via 2D visual matching and OCR shortcuts.
Background & Motivation¶
Spatial intelligence serves as a foundational pillar of general intelligence across species. Biological agents evolve specialized spatial representations tuned to their perceptual systems: migratory birds track geomagnetic headings, echolocating bats infer depth and distance, and humans seamlessly fuse visual streaming with proprioceptive signals to build internal cognitive maps for long-range navigation. In contrast, modern Vision-Language Models (VLMs) have been pre-trained primarily on passive, web-scale image-text and short video datasets. Whether these models naturally develop valid, metric world models of large physical environments remains an unanswered empirical question.
Prior spatial evaluations have largely focused on localized indoor scenes, table-top object relationships, or synthetic environments like factories and driving simulators. Such benchmarks span relatively short temporal windows (1 to 3 minutes) and small physical footprints, failing to challenge models with cumulative navigational drift, prolonged visual memory decay, and large-scale geographic path integration. Generating high-precision 6-DoF camera poses and trajectory ground truth from hours of real-world outdoor walking tours has historically been cost-prohibitive: manually plotting continuous paths on maps requires roughly 6 hours of human labor per 1-hour video while entirely discarding rotation information, whereas wearing custom sensor arrays (such as ARIA glasses) strictly limits data collection to constrained geographic regions and hardware availability.
To overcome this bottleneck, this work leverages Google's Visual Positioning System (VPS) combined with minimal human start/end coordinates to automatically infer high-fidelity geographic routes and camera rotations across hundreds of hours of diverse YouTube city walking tours. Building on the classical landmark-route-map hierarchy from cognitive science, the authors translate spatial cognition into a rigorous, diagnostic multi-choice video question-answering benchmark. Core idea: evaluate city-scale spatial intelligence across three hierarchical stages of cognitive spatial awareness—landmark localization, egocentric route integration, and allocentric survey map formation—revealing whether multimodal frontier models construct genuine geometric mental maps or merely exploit 2D visual recognition and text-matching heuristics.
Method¶
Overall Architecture¶
The KilometerVision benchmark comprises two principal components: an automated, city-scale video-to-map grounding pipeline and a hierarchically structured multi-choice evaluation suite. The data collection ingests 235 high-resolution, hour-long walking tour videos from worldwide urban centers (totalling 288 video hours). Human annotators spend only ~5 minutes per video specifying starting and ending coordinates, which primes a bidirectional, confidence-weighted VPS solver operating at 1 FPS to extract metric trajectory coordinates and orientation quaternions.
From this grounded corpus, the authors extract video clips spanning up to 10 minutes (approximating 1 km of walking distance) to instantiate 1000 balanced, 5-way video multiple-choice questions. These tasks systematically interrogate spatial understanding across the three developmental stages of human spatial cognition: landmark recognition and line-of-sight distance estimation, route-level compass tracking and loop closure detection, and survey-level map trace matching and start-to-end Euclidean distance estimation. The standardized multiple-choice formulation provides an accessible interface across diverse open-source and proprietary model architectures while unlocking intermediate thinking traces for qualitative reasoning analysis.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Raw Walking Tour Videos<br/>235 videos / 288 hours / Start-End GPS"] --> B["Bidirectional Weighted VPS Pipeline<br/>1 FPS trajectory and quaternion extraction"]
B --> C["DTW Path Filtering & Routes API Verification<br/>Outlier removal & route trace validation"]
C --> D["Cognitive Hierarchical Task Construction<br/>1000 five-way video QA benchmarks"]
D --> E["Landmark Stage: Identification & Distance<br/>High-confidence keyframe matching / Metric distance"]
D --> F["Route Stage: Compass Heading & Loop Closure<br/>Path integration / 8-way orientation / Temporal recall"]
D --> G["Survey Map Stage: Topology & Euclidean Distance<br/>Labeled map / Blank map / Straight-line distance"]
Key Designs¶
1. Bidirectional Weighted VPS Pose Inversion: Scalable Continuous Trajectory Extraction
Obtaining ground-truth metric poses from uncalibrated internet videos at scale typically presents an insurmountable annotation hurdle. KilometerVision solves this by framing localization through Google Visual Positioning System (VPS) queries against StreetView embeddings. With human raters providing approximate start and end coordinates, a forward temporal pass sequentially processes frames at 1 FPS, updating estimated positions whenever the VPS confidence exceeds a strict threshold. A complementary backward pass initiates from the final frame, and both directional trajectories are subsequently reconciled via confidence-weighted interpolation followed by a final refinement run. To guarantee metric reliability in feature-sparse corridors, Dynamic Time Warping (DTW) metrics against calibrated human trajectory ground truth establish an acceptance envelope, discarding drifting segments and compressing full-video grounding costs from 6 hours of manual tracing down to 5 minutes of setup.
2. Cognitive Hierarchical Task Suite: Bridging Landmarks, Routes, and Survey Knowledge
Drawing from spatial cognitive psychology, the evaluation protocol operationalizes three distinct stages of spatial comprehension into multiple-choice formats: - Landmark Stage: Evaluates the model's capacity to anchor itself via salient environmental entities. The model receives a video clip alongside a candidate landmark crop (mined via non-maximal suppression on peak-confidence VPS frames, or non-overlapping negative distractors). The prompt asks whether the landmark appeared in the clip, and if so, what the line-of-sight Euclidean distance is between the landmark capture pose and the final video frame, probing visual retrieval alongside metric 3D scale estimation. - Route Stage: Tests egocentric movement tracking and path integration without external global reference frames. Tasks include Compass Tracking (given the camera's initial 8-way cardinal orientation, predict the terminal orientation after navigating cumulative turns); Loop Closure Detection (determine whether the final frame revisits an earlier location within a 4-meter radius, specifying the exact timestamp of initial passage); and Route Summary (select the correct sequence of turn-by-turn navigation actions generated via Google Maps Routes API). - Survey Knowledge (Map) Stage: Represents the highest tier of spatial awareness—the formation of an allocentric, bird's-eye mental map. Tasks include Map Trace Identification (matching the traversed path against five visual overhead map candidates, evaluated both with full text/icon annotations and over in-painted blank maps stripped of text/landmarks) and Start-End Euclidean Distance Estimation (computing straight-line distance across non-linear street routes, where distractors are guaranteed to differ by at least 100 meters).
3. Multimodal Cue Decoupling and Thinking-Trace Diagnostics
To rigorously separate authentic physical geometric inference from trivial semantic pattern matching, the benchmark deploys controlled ablation probes. A blind baseline replaces RGB video inputs with blank frames to confirm that questions cannot be answered purely through linguistic priors. Furthermore, Gemini 2.5 Flash is evaluated under privileged telemetric context (explicit camera yaw and absolute VPS coordinates) to assess zero-shot sensor integration. Crucially, visual overhead map questions are contrasted between fully labeled cartographic displays and text-free basemaps, while model thinking traces are systematically mined to expose whether answers stem from continuous path integration or brittle, brute-force 2D frame matching.
Key Experimental Results¶
Main Results¶
Evaluations encompass five representative vision-language models—PLM-8B, Qwen2.5-VL-72B, Gemini 2.5 Flash, Claude Opus, and GPT-5—under a uniform 32-frame sampling budget, accompanied by a blind baseline and an 18-participant human baseline evaluated at 30 FPS across 337 randomly sampled questions (with an inter-rater agreement of 0.76).
| Model / Baseline | Landmark (Acc %) | Compass (Acc %) | Route Summary (Acc %) | Map Trace w/ Text (Acc %) | Map Trace w/o Text (Acc %) | Loop Closure (Acc %) | Euclidean Distance (Acc %) | Overall Mean Acc (%) |
|---|---|---|---|---|---|---|---|---|
| Random Chance | 20.0 | 20.0 | 20.0 | 20.0 | 20.0 | 20.0 | 20.0 | 20.0 |
| Gemini 2.5 Flash (Blind) | ~20.0 | ~20.0 | ~20.0 | ~20.0 | ~20.0 | ~20.0 | ~20.0 | ~20.0 |
| PLM-8B (32 frames) | ~22.0 | ~21.0 | ~23.0 | ~20.5 | ~19.5 | ~24.0 | ~21.0 | ~21.6 |
| Qwen2.5-VL-72B (32 frames) | ~23.5 | ~22.0 | ~26.0 | ~22.0 | ~20.0 | ~28.0 | ~22.5 | ~23.4 |
| Claude Opus (32 frames) | ~34.0 | ~28.0 | ~42.0 | ~31.0 | ~22.0 | ~58.0 | ~25.0 | ~34.9 |
| Gemini 2.5 Flash (32 frames) | ~38.0 | ~31.0 | ~42.2 | ~28.6 | ~28.6 | ~62.0 | ~28.0 | ~36.9 |
| GPT-5 (32 frames) | ~45.0 | ~36.0 | ~54.0 | ~38.0 | ~29.0 | ~78.0 | ~32.0 | ~44.6 |
| GPT-5 (500 frames) | ~52.0 | ~48.0 | ~67.0 | ~49.0 | ~38.0 | ~88.0 | ~42.0 | ~54.9 |
| Human Baseline | 78.4 | 68.2 | 74.5 | 76.0 | 71.8 | 52.1 | 78.0 | 71.3 |
Note: In loop closure detection, human evaluators suffered from fatigue across 10-minute videos, often failing to notice subtle returns and achieving 52.1%. In contrast, GPT-5 and Gemini 2.5 Flash achieved super-human loop closure scores via exhaustive keyframe retrieval, while lagging dramatically behind human performance on tasks requiring geometric path integration and metric mental mapping.
Ablation Study¶
Ablation investigations using Gemini 2.5 Flash examine the impact of privileged orientation/coordinate injection, cartographic text annotations, and the complete omission of video inputs:
| Experimental Configuration (Gemini 2.5 Flash) | Route Summary (Acc %) | Map Trace w/ Text (Acc %) | Map Trace w/o Text (Acc %) | Core Observation / Mechanism |
|---|---|---|---|---|
| Default (Video input only) | 42.2 | 28.6 | 28.6 | Baseline multi-modal performance |
| + Privileged Camera Yaw (+Yaw) | - | ~29.2 | ~28.0 | Raw angular telemetry yields negligible benefit |
| + Privileged Trajectory (+Yaw +VPS) | - | ~29.5 | ~28.8 | Zero-shot models fail to ground coordinate strings |
| + Labeled Map Context | 49.0 (+6.8) | - | - | Text labels on map provide cross-modal shortcuts |
| + Unlabeled Map Context | 37.4 (-4.8) | - | - | Bare geometric shapes confuse cross-modal reasoning |
| + Text Route Summary Context | - | 34.7 (+6.1) | 26.5 (-2.1) | Route text helps only when map contains street names |
| Video Omitted Entirely (No video) | 53.1 (+10.9) | 46.3 (+17.7) | 38.1 (+9.5) | Eliminating video forces pure text/map OCR alignment |
Key Findings¶
- Robust 2D Recognition vs. Absent Metric Grounding: Decomposing performance on landmark and loop closure queries reveals that models possess strong 2D visual recognition (>80-90% accuracy in verifying whether an entity was observed), but suffer catastrophic degradation when computing metric distance or initial revisit timing (~25% accuracy). VLMs readily classify what an object is, but exhibit virtually no awareness of where it stands in 3D metric space.
- Super-Human Loop Closure via Brute-Force Retrieval: GPT-5 (78-88%) and Gemini 2.5 Flash (62%) outperform humans on loop closure detection. However, analysis of their reasoning traces demonstrates that this stems entirely from dense, exhaustive frame-to-frame visual retrieval rather than topological path integration. This shortcut proves fragile: visual false positives (such as observing a distant statue with similar architecture) routinely derail model deduction.
- Collapse of Angular Path Integration: Tracking cumulative turns in the Compass Heading task proved disastrous for 32-frame models, languishing between 28% and 36%. Increasing the frame count from 32 to 600 in Gemini 2.5 Flash produced no observable gain (~38%), indicating that simple token expansion does not induce spatial odometry; only GPT-5 leveraged dense 500-frame inputs to reach 57%.
- The "Video Distraction" Paradox: Stripping away the input video entirely and prompting the model to reconcile textual route summaries directly against labeled maps boosted accuracy from 42.2% to 53.1% on Route Summary, and from 28.6% to 46.3% on Map Trace. This demonstrates that VLMs excel at reading maps via OCR text comprehension, but lack the grounding mechanisms required to fuse continuous egocentric visual flow into an allocentric topological frame.
Highlights & Insights¶
- Operationalizing Cognitive Spatial Hierarchy for Large Multimodality: KilometerVision bridges cognitive psychology and AI benchmarking by systematically decomposing spatial intelligence into landmark, route, and survey stages, exposing the stark divergence between surface-level vision-language correlation and genuine physical embodiment.
- High-Precision Video Grounding at Negligible Overhead: Combining Google's StreetView-backed VPS API with forward-backward confidence weighting and DTW error pruning reduces the human annotation burden from 6 hours per video to 5 minutes, establishing a practical blueprint for turning raw web video archives into metrically calibrated spatial datasets.
- Unmasking the Illusion of VLM World Models: Empirical findings decisively prove that contemporary frontier VLMs do not construct metric mental maps. Apparent navigation capabilities in smaller benchmarks are largely artifacts of high-frequency OCR text spotting and 2D appearance matching, underscoring critical research directions for spatial foundation models.
Limitations & Future Work¶
- Passive Video Consumption vs. Active Embodied Agency: The benchmark relies exclusively on passive human walking footage, lacking motor action feedback, vestibular sensory streams, and active closed-loop navigation policies. Human spatial mapping also degrades under purely passive observation; extending this protocol to embodied interactive environments remains a key horizon.
- Discretized Multiple-Choice Evaluation: Formulating spatial tasks as 5-way multiple-choice questions enables standardized evaluation and thinking-trace analysis, but inevitably quantizes continuous spatial coordinates and topological graphs into discrete bins.
- In-Context Sensory Integration Barriers: Zero-shot conditioning on raw coordinate and orientation tokens fails to improve performance, demonstrating that future spatial foundation models require specialized architectural inductive biases or supervised geometric fine-tuning to effectively ingest odometry streams.
Related Work & Insights¶
- vs. VSI-Bench / Thinking in Space: Prior spatial intelligence evaluations focused on indoor table-top or factory-scale scenes spanning dozens of meters over 1 to 3 minutes. KilometerVision scales temporal and spatial horizons by an order of magnitude, probing continuous kilometer-long trajectories and complex urban topologies.
- vs. Aria Everyday Activities / Traditional Visual SLAM Datasets: Conventional visual-inertial SLAM benchmarks require dedicated wearable sensor rigs and focus purely on low-level trajectory drift errors. In contrast, KilometerVision scales across diverse global cities using uncalibrated web videos, evaluating holistic spatial reasoning through a flexible multimodal language interface.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering kilometer-scale, 10-minute video spatial intelligence benchmark rooted in classical cognitive developmental hierarchy.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive cross-model evaluation, frame-rate scaling, human alignment, privileged data ablations, and detailed qualitative trace audits.
- Writing Quality: ⭐⭐⭐⭐⭐ Compelling motivation, rigorous experimental formulations, and insightful dissection of spatial shortcuts versus true path integration.
- Value: ⭐⭐⭐⭐⭐ A landmark diagnostic resource guiding the evolution of Vision-Language Models toward genuine spatial intelligence, world models, and embodied robotics.