Skip to content

Stand Up and Move: Benchmarking Interactive Spatial Intelligence in WalkerBench

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/lalayang123456-ctrl/WalkerBench
Area: Robotics & Embodied AI
Keywords: Interactive Spatial Intelligence, Spatial Benchmarking, Explicit Topological Memory, Anterograde Spatial Amnesia, Humanoid Robot Navigation

TL;DR

To overcome the limitations of static "Spectator View" benchmarks that ignore active parallax for disambiguation, this paper presents WalkerBench—a global interactive benchmark spanning 161 cities—reveals a fundamental representation misalignment between VLMs' 1D linear context and 3D topology, and introduces Spatial-IDE, a training-free scaffold with explicit topological memory and cognitive decoupling that achieves more than twofold gains on benchmarks and zero-shot transfer on a Unitree humanoid robot.

Background & Motivation

The ultimate frontier of spatial intelligence lies in embodied interaction within the physical world, which demands a paradigm shift from semantic pattern recognition over frozen images to purposeful engagement with dynamic 3D environments. However, existing spatial benchmarks overwhelmingly adopt a "Spectator View"—evaluating models using either single static frames or pre-recorded passive videos. A static first-person view is geometrically degenerate: a VLM estimating distance from a single frame inevitably hallucinates scale via unreliable semantic priors (e.g., misjudging a 500m distant object as 300m), whereas a deliberate lateral displacement of several meters immediately resolves the ambiguity via geometric parallax. By stripping agents of the ability to move, existing benchmarks probe only static visual commonsense rather than genuine interactive spatial intelligence.

Furthermore, conventional vision-and-language navigation benchmarks remain largely confined to synthetic indoor environments or a handful of manually annotated urban corridors, lacking global architectural and topological diversity. When state-of-the-art vision-language models (VLMs) are evaluated in open real-world outdoor settings, their performance collapses dramatically: while human navigators achieve 69.43%, the strongest proprietary models achieve only 24.49%, degrading sharply as trajectory depth increases. The paper diagnoses this failure mode—strong local perception at each timestep accompanied by catastrophic loss of global coherence—as "Anterograde Spatial Amnesia."

The underlying root cause is a Representation Mismatch: the 1D linear context stream of autoregressive transformers is fundamentally incommensurate with the non-Euclidean nature of 3D topological space. Over multi-turn interactions, linear dialogue history tangles immediate perception, long-term memory, and spatial planning, causing salient landmarks to be diluted by growing textual descriptions and triggering irrecoverable spatial forgetting. To resolve this tension, the core idea is to break the linear context accumulation by externalizing persistent state into a program-level Explicit Topological Memory (ETM) graph and decoupling per-step execution into stateless goal-directed perception and pure topological reasoning, validated on WalkerBench across 161 global cities and transferred zero-shot to a physical humanoid robot.

Method

Overall Architecture

Spatial-IDE is designed around the foundational principle of "breaking the linear context," eliminating unbounded historical image and text accumulation. The system architecture integrates three decoupled pillars: a Goal-Directed Perception module (Perception Core), an Explicit Topological Memory (ETM), and a high-level Spatial Reasoning Engine (Reasoning Core). At each decision step \(t\), the agent executes two separate, stateless single-turn VLM calls mediated by the ETM. First, the perception module inspects the current image using a targeted query generated in the previous step to extract an ultra-concise semantic note. Next, the updated ETM is projected into a compact, fixed-size working memory snapshot (encoding coordinates, traversed path, and available topological transitions) that feeds into the spatial reasoning engine for action selection and next-query formulation. When fine-grained metric estimation demands multi-view parallax, the engine selectively retrieves geometrically complementary historical viewpoints from the ETM rather than unrestrictedly flooding the context window.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Current Viewpoint Panorama & Environment Input"] --> B["Goal-Directed Visual Perception<br/>Extract minimal factual note from targeted query"]
    B --> C["Explicit Topological Memory Update<br/>Maintain metric-topological graph & semantic annotations"]
    C --> D["Spatial Reasoning Engine Decision<br/>Fixed-length memory snapshot guides action & query"]
    D -->|Action Execution| E["Physical Action Execution & Humanoid Pose Transition"]
    D -.->|On-demand Geometric Retrieval| C

Key Designs

1. Goal-Directed Visual Perception: Transforming Open Scene Description into Targeted Inquiry Standard visual navigation agents prompt VLMs to generate open-ended scene narrations at every step, which introduces redundant descriptions, dilutes key landmarks, and squanders the model's finite attention budget on irrelevant background clutter. Goal-Directed Perception (GDP) replaces generic captioning with targeted interrogation: guided by a precise, decision-relevant inquiry formulated by the reasoning engine at the previous step (e.g., "Is the pharmacy sign visible on the right façade, and at what bearing?"), the perception core analyzes the current RGB image. It returns exclusively a compact Semantic Note—a short factual proposition strictly limited to at most 25 words. Because this call carries only the current image and a single question, its execution cost and token count remain invariant across arbitrarily long navigation trajectories.

2. Explicit Topological Memory Update: Program-Level Metric Graph Storage Outside the Context Window To fundamentally conquer anterograde spatial amnesia caused by the mismatch between 1D text streams and 3D geometry, the framework maintains an Explicit Topological Memory (ETM). ETM is an external dynamic topological graph \(\mathcal{G}_t = (\mathcal{V}_t, \mathcal{E}_t)\) stored and manipulated as a program-level data structure rather than an LLM context buffer. Each node \(v \in \mathcal{V}_t\) records ground-referenced metric coordinates (in integer meters), camera yaw, and pitch, annotated with the semantic note list extracted at that viewpoint; each edge \(e \in \mathcal{E}_t\) stores physical displacement vectors and relative bearings. Because coordinates and topology reside in program memory, they never suffer from attention attenuation, hallucinated drift, or context overflow even after fifty steps. At each turn, ETM serializes into a compact snapshot ([POSITION], [PATH], [OBSERVATION LOG], [AVAILABLE MOVES]), ensuring bounded, uniformly high-quality spatial inputs.

3. Spatial Reasoning Engine Decision: Cognitive Decoupling for High-Level Planning and On-Demand Parallax Retrieval The spatial reasoning engine receives only the current panorama and the serialized ETM snapshot, completely free of past conversational turns or historical visual streams. Decoupling low-level visual measurement from high-level decision-making allows the VLM to operate purely as an executive graph planner. Furthermore, for fine-grained metric tasks requiring multi-view parallax (such as distance or angle triangulation), the engine executes "on-demand geometric retrieval": by inspecting ETM node poses and baseline distances, it calculates the most geometrically complementary historical viewpoint and requests only that specific frame. This achieves surgical multi-view geometric triangulation while keeping execution bounded and efficient.

Key Experimental Results

Main Results

WalkerBench evaluates two progressive tiers: Tier-I coarse-grained spatial tasks (Navigation Nav and Visual Localisation Vis) and Tier-II fine-grained metric tasks (Building Height, Distance Dis, and Relative Bearing Angle), evaluated via accuracy (%) across 161 cities. Table 3 presents the evaluation across nine leading proprietary and open-source models under the standard multi-turn dialogue baseline (Original) versus the Spatial-IDE framework, benchmarked against human reference performance.

Evaluated Agent Setting Nav (%) Vis (%) Height (%) Dis (%) Angle (%) Total (%) Relative Gain
Claude-Opus-4.6 Original 52.86 18.57 14.0 9.0 18.0 24.49 —
Claude-Opus-4.6 w/ Spatial-IDE 68.29 38.57 35.0 38.0 36.0 43.17 +76.28%
Kimi-K2.5 Original 54.86 9.71 17.0 12.0 15.0 21.71 —
Kimi-K2.5 w/ Spatial-IDE 69.14 28.57 37.0 31.0 33.0 39.74 +83.05%
Gemini-3-Pro Original 50.86 10.29 12.0 16.0 11.0 20.03 —
Gemini-3-Pro w/ Spatial-IDE 65.43 29.14 33.0 35.0 28.0 38.11 +90.26%
GPT-5-chat-latest Original 20.00 10.86 16.0 19.0 20.0 17.17 —
GPT-5-chat-latest w/ Spatial-IDE 45.14 29.71 37.0 38.0 39.0 37.77 +119.98%
Qwen3-VL-235B Original 16.00 10.86 46.0 1.0 13.0 17.37 —
Qwen3-VL-235B w/ Spatial-IDE 40.57 30.00 58.0 20.0 31.0 35.91 +106.74%
Qwen3-VL-8B Original 10.29 10.00 39.0 6.0 8.0 14.66 —
Qwen3-VL-8B w/ Spatial-IDE 32.29 28.57 52.0 23.0 25.0 32.17 +119.44%
GLM-4.6V Original 29.14 11.43 1.0 13.0 12.0 13.31 —
GLM-4.6V w/ Spatial-IDE 46.57 30.29 21.0 31.0 30.0 31.77 +138.69%
Average (Original) Baseline 34.67 11.43 18.33 12.33 13.00 18.17 —
Average (Spatial-IDE) Full Framework 52.92 30.41 37.11 32.11 30.78 36.66 +104.97%
Human (Reference) Baseline 78.29 68.86 73.0 66.0 61.0 69.43 —

Ablation Study

To isolate the quantitative contributions of Explicit Topological Memory (ETM) and Goal-Directed Perception (GDP), systematic ablations were conducted across all nine agents over the complete WalkerBench test suite.

Configuration Nav (%) Vis (%) Height (%) Dis (%) Angle (%) Total (%) Absolute Gain (\(\Delta\))
Original (baseline) 34.67 11.43 18.33 12.33 13.00 18.17 —
+ ETM only 47.14 21.00 25.00 23.00 23.00 27.83 +9.66
+ GDP only 39.00 26.57 33.00 26.57 25.57 30.14 +11.97
+ ETM + GDP (Full Spatial-IDE) 52.92 30.41 37.11 32.11 30.78 36.66 +18.49

Key Findings

  • Distinct Division of Labor and Positive Synergies: ETM primarily unlocks long-horizon spatial coherence, yielding a +12.47 point jump in Nav by converting transient steps into an \(O(1)\) queryable graph. Conversely, GDP drives fine-grained perceptual precision, elevating Vis by +15.14 points and each metric task by over +12 points by filtering irrelevant scene clutter. Combined, ETM + GDP produces an absolute improvement of +18.49 points, surpassing simple additive predictions.
  • Structural Universality Across Model Scales: Gains are consistent across all architectures and scales: Qwen3-VL-8B and Qwen3-VL-235B achieve relative improvements of +119.44% and +106.74% respectively. Equipped with Spatial-IDE, the lightweight 8B model achieves 32.17% overall accuracy, outperforming the unaugmented performance of top-tier proprietary models like Claude-Opus (24.49%) and GPT-5 (17.17%).
  • Zero-Shot Real-World Humanoid Deployment: Deployed on a physical Unitree G1 humanoid robot, the high-level planner operated directly on onboard RGB imagery and local traversability candidates, successfully executing multi-block autonomous outdoor navigation across brick pavements, crosswalks, and signalized intersections without any physical fine-tuning.

Highlights & Insights

  • From Spectator to Actor Paradigm: Demonstrates that passive single-image or pre-recorded video benchmarks are geometrically degenerate, establishing active movement and parallax as mandatory prerequisites for evaluating genuine spatial intelligence.
  • Pathological Diagnosis of Anterograde Spatial Amnesia: Uncovers the architectural mismatch between 1D autoregressive context and 3D spatial topology, pinpointing why standard LLM/VLM multi-turn agent frameworks collapse on extended physical paths.
  • Zero-Shot, Training-Free Cognitive Scaffold: Requires no weight modification or gradient updates; simply externalizing memory into an explicit graph and decoupling perception from reasoning doubles overall spatial benchmark accuracy while drastically cutting per-step inference tokens.

Limitations & Future Work

  • Admitted Limitations: Real-world robot transfer abstracts dynamic motor stability and bipedal locomotion to proprietary controllers; active dynamic obstacle avoidance in high-density pedestrian flow remains preliminary. Furthermore, while Spatial-IDE doubles agent capability, a substantial 26-percentage-point performance gap remains compared to human spatial reasoning (69.43%).
  • Future Directions: Integrating lightweight neural rendering primitives (such as continuous 3D Gaussian Splatting) into ETM node representations could allow agents to internally synthesize novel unvisited viewpoints, pushing spatial planning and parallax triangulation to higher precision.
  • vs. Static Spatial Benchmarks (Spatial-Bench, SpatialRGPT, VSR): Prior work confines models to passive single frames without the ability to move; WalkerBench forces active spatial displacement across 161 cities to resolve geometric ambiguities.
  • vs. Indoor VLN Simulators (R2R, RxR, Habitat): Indoor simulators evaluate within limited synthetic scenes relying on privileged depth or GPS sensor data; WalkerBench operates on real global outdoor street topologies with pure RGB input.
  • vs. Agentic Scaffolds (ReAct, Swe-agent): Conventional agentic architectures append full observation traces to conversational context; Spatial-IDE demonstrates that spatial navigation fundamentally requires external geometric graph memory to halt attention dilution.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the concept of interactive spatial intelligence via active parallax, diagnoses spatial amnesia in VLMs, and constructs a global-scale real-world benchmark.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 9 leading commercial and open-source VLMs, detailed component ablations, and physical humanoid robot field deployment.
  • Writing Quality: ⭐⭐⭐⭐⭐ Exceptionally articulate narrative connecting geometric theory, cognitive architecture, and robotic execution.
  • Value: ⭐⭐⭐⭐⭐ Offers a standard benchmark and plug-and-play methodology for embodied AI, spatial reasoning, and humanoid robotics.