Skip to content

STEP: Spatial Thinking and Egocentric Pointing for Embodied Instruction Following

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/VIPL-VSU/STEP
Area: Multimodal VLM / Robotics & Embodied AI
Keywords: Embodied Instruction Following, Spatial Thinking, Egocentric Pointing, BEV Semantic Map, Chain-of-Thought Planning

TL;DR

STEP integrates global Bird's-Eye-View (BEV) semantic maps with egocentric visual waypoint prompting, driven by an automated 400k+ multimodal Chain-of-Thought data engine to establish an interpretable point-level planning paradigm: "Look at the View, Think with the Map, and Point to the Goal."

Background & Motivation

Embodied Instruction Following (EIF) demands that an agent carry out long-horizon navigation and fine-grained manipulation in complex environments based on natural language instructions. With the rapid evolution of large language models (LLMs) and vision-language models (VLMs), recent literature has widely embraced foundation models as embodied decision planners. Nonetheless, current embodied agent paradigms remain polarized along the perceptual and planning axes, failing to bridge high-level semantic intent and low-level physical actuation.

On the perception front, existing agents suffer severely from "Spatial Myopia": end-to-end executors rely almost exclusively on egocentric camera feeds, lacking long-term topological memory and awareness of occluded spaces; conversely, high-level textual planners operate on symbolic object lists, detached from direct, continuous visual grounding. On the planning front, agents face an acute "Granularity Imbalance": atomic action outputs (e.g., move_ahead, rotate_left) entail low decision efficiency, bloated step counts, and poor interpretability; in contrast, abstract subgoal outputs (e.g., go to the kitchen) are excessively vague, creating a substantial "execution gap" between high-level reasoning and low-level motion control.

Drawing inspiration from Tolman's classical Cognitive Maps theory and human semantic-spatial anchoring—where humans navigate by coupling an internal mental layout with explicit physical target points in their visual field—this paper adopts a balanced perspective. Rather than choosing between myopic egocentric control and abstract symbolic planning, the agent can jointly perceive a global BEV map and egocentric frames, projecting discrete visual markers onto navigable surfaces and generating explicit Chain-of-Thought (CoT) traces before selecting waypoints. Core idea: introduce the STEP framework uniting Spatial Thinking and Egocentric Pointing under the paradigm of "Look at the View, Think with the Map, and Point to the Goal," empowered by a scalable STEP-CoT data engine that bridges global layout reasoning with point-level physical grounding.

Method

Overall Architecture

The STEP framework is structured to translate multimodal environmental observations and textual goals into verifiable, physically compliant point-level decisions. The overall system takes three inputs: a natural language instruction \(\mathcal{T}\), an Annotated Semantic Map (ASM) \(\mathcal{G}_t\) dynamically updated from RGB-D history, and an egocentric frame \(I_t\) visually prompted with discrete waypoint markers on traversable surfaces.

The pipeline comprises two synchronized phases: perception and planning. In perception, raw observations are distilled via Environment Filtering into an abstracted 2D BEV map, while Egocentric Traversability Grounding segments reachable floors and superimposes Set-of-Mark (SoM) identifiers. A tri-encoder backbone built on Qwen2.5-VL-7B projects these modalities into a unified latent space via dedicated linear projectors. In planning, the multimodal LLM first synthesizes an explicit Chain-of-Thought (CoT) reasoning sequence analyzing spatial relationships, and subsequently outputs a categorical marker identifier or a primitive rotation command. The Fast Marching Method (FMM) then translates the predicted point into continuous, smooth trajectory steps.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    IN["RGB-D Streams + Task Instruction"] --> A["Egocentric Traversability Grounding<br/>Grounded-SAM segmentation & grid projection"]
    IN --> B["Environment Filtering & Map Abstraction<br/>3D voxel projection & goal-aware filtering"]
    A --> C["Tri-Encoder Cross-Modal Alignment<br/>View, map, and text feature projection"]
    B --> C
    C --> D["CoT & Point-Level Action Prediction<br/>Reasoning trace & waypoint selection"]
    D --> E["Low-Level Trajectory Synthesis<br/>Fast Marching Method FMM execution"]

Key Designs

1. Egocentric Traversability Grounding: visual spatial anchoring on homogeneous ground

Standard multimodal models excel at identifying distinct foreground items like lamps or chairs, but they struggle to predict accurate spatial coordinates on visually uniform floor or carpet areas. STEP circumvents continuous coordinate regression by transforming spatial landing into a discrete visual prompting problem. The system first deploys Grounded-SAM with generic floor prompts \(\mathcal{T}_{ground} = \{\text{ground}, \text{carpet}\}\) on the egocentric frame \(I_t\) to produce a binary traversability mask \(\mathcal{M}_{sem} \in \{0, 1\}^{H \times W}\).

A local 3D horizontal ground grid \(\mathcal{P}_w = \{ (n \cdot s, m \cdot s, 0, 1)^T \mid n \in \mathbb{Z}^+, m \in \mathbb{Z} \}\) in front of the robot is projected onto the 2D image plane using camera intrinsics \(\mathbf{K}\) and extrinsics \(\mathbf{T}_c\), retaining only the points falling within the navigable mask \(\mathcal{M}_{sem}\): $$ \lambda_i \mathbf{u}i = \mathbf{K}\mathbf{T}_c \mathbf{P}_i, \quad \mathcal{V}} = { \mathbf{ui \mid \mathcal{M}_w } $$ Using a Set-of-Mark (SoM) approach, unique numerical tags (e.g., "Marker 25") are overlaid on each candidate pixel }(\mathbf{P}_i) = 1, \mathbf{P}_i \in \mathcal{P\(\mathbf{u}_i\). This formulation reframes complex geometric waypoint regression into a discrete categorical selection task, directly tapping into the VLM's native object grounding capabilities while guaranteeing kinematic navigability.

2. Environment Filtering & Map Abstraction: resolving spatial myopia without cognitive overload

Relying strictly on egocentric visual streams leaves the agent vulnerable to long-term memory loss and disorientation behind obstacles. STEP incorporates an Environment Filtering pipeline to dynamically construct a 2D Annotated Semantic Map (ASM). Point clouds back-projected from semantic segmentation and depth maps are voxelized and orthographically projected along the height axis into a multi-channel grid \(M_t \in \{0, 1\}^{(C+2) \times H \times W}\), capturing \(C\) object classes alongside obstacle and exploration masks.

To prevent dense spatial grids from flooding the VLM with redundant tokens, STEP abstracts the grid into an annotated graphic: obstacle boundaries are dilated by a safety kernel \(K_{dila}\); an instruction-conditioned aggregator \(A(\cdot)\) suppresses task-irrelevant background categories to avoid visual clutter; and a force-directed label placement renderer \(R(\cdot)\) draws clean semantic labels directly onto the 2D map: $$ \mathcal{G}t = \mathcal{R}\left(A(K)\right) $$ This linguistically indexed topological layout enables the VLM to parse global scene context with minimal perceptual overhead.}(M_t), \mathcal{I

3. Tri-Encoder Cross-Modal Alignment: unifying view, map, and language in a shared latent space

To simultaneously reason over localized perspectives, global layouts, and language instructions, STEP adopts a tri-encoder structure built on Qwen2.5-VL-7B. Text tokens are mapped to language representations \(\mathbf{X}_{inst} \in \mathbb{R}^{L \times d}\) via text encoder \(f_\theta\), while the egocentric frame \(I_t\) and the rendered semantic map \(\mathcal{G}_t\) pass through visual encoders \(f_\phi\) and \(f_\psi\) to extract visual features \(\mathbf{X}_{vis}\) and topological map features \(\mathbf{X}_{map}\).

Modality-specific projection layers \(g_v\) and \(g_m\) map the heterogeneous visual and spatial representations into the shared \(d\)-dimensional embedding space: $$ \mathbf{Z}{vis} = g_v(\mathbf{X}}) \in \mathbb{R}^{V \times d}, \quad \mathbf{Z{map} = g_m(\mathbf{X} $$ The inputs are interleaved with learnable boundary markers to form the synchronized sequence }) \in \mathbb{R}^{K \times d\(\mathbf{H}_t = [s_{inst}; \mathbf{X}_{inst}; s_{map}; \mathbf{Z}_{map}; s_{view}; \mathbf{Z}_{vis}]\). During training, the visual and map backbones are kept frozen to prevent catastrophic forgetting of general visual priors, while the LLM backbone and the linear projection heads undergo full fine-tuning.

4. Scalable STEP-CoT Data Engine: dual-stream synthesis for cross-modal spatial reasoning

End-to-end point-level planning requires rich supervisory signals linking instructions, BEV maps, camera observations, and waypoint choices. STEP establishes the STEP-CoT data engine, automatically generating 420k+ multimodal reasoning samples via two distinct streams: - Stream A (Action-Grounded Reasoning-Trace Annotation): For each step in expert demonstrations, a look-ahead window \(\Delta\) selects the furthest reachable pose \(\delta^*\) projected onto candidate waypoints as ground truth \(v^*\). Heading Normalization eliminates ego-versus-world frame ambiguity, while object filtering suppresses hallucination. A powerful teacher model (Qwen3-VL-235B-Instruct) synthesizes structured CoT reasoning traces \(C_t\) explaining why marker \(v^*\) must be chosen. Rigorous dual-stage validation—combining symbolic goal consistency checks \(\Phi_{cons}\) with a binary VLM audit \(\Phi_{eval}\)—ensures high fidelity, yielding 229,544 validated traces. - Stream B (Primitive Skills Data Synthesis): Spatial navigation is decomposed into fundamental sub-capabilities: Pose Estimation relative to the map (45,623 samples), Relative Direction Inference of target objects (18,172 samples), Semantic Perception inventorying visible entities (9,069 samples), and Topological Reasoning predicting layouts (72,552 samples), supplemented by Task Breakdown sequences (45,321 samples). These primitive tasks instill robust geometric grounding and cross-view alignment, preventing the model from exploiting superficial shortcuts.

Key Experimental Results

Main Results

On the long-horizon ALFRED benchmark, STEP-7B was comprehensively evaluated across seen and unseen validation splits against specialist robotic learning models and generalist LLM/VLM systems.

Method Training Mode Val Seen SR (%) ↑ Val Seen GC (%) ↑ Val Unseen SR (%) ↑ Val Unseen GC (%) ↑ \(\Delta\text{SR}\) (%) ↓
E.T. from scratch (Specialist) 46.59 52.92 7.32 20.87 39.27
FILM modular (Specialist) 24.63 37.20 20.10 32.45 4.53
LEBP modular (Specialist) 27.63 35.76 22.36 29.58 5.27
SayCan few-shot (Generalist) 12.30 24.52 9.88 22.54 2.42
LLM-P (GPT) few-shot (Generalist) 16.45 30.11 15.36 29.88 1.09
R2C-Mistral-7B fine-tune (Generalist) 22.31 32.40 22.35 31.97 0.04
Qwen2.5-VL-7B (Base) zero-shot (Generalist) 9.96 18.59 7.87 18.26 2.09
Qwen3-VL-32B zero-shot (Generalist) 15.20 26.69 11.76 24.54 3.44
GPT-4o zero-shot (Generalist) 16.00 27.13 16.00 25.00 0.00
GPT-5.4 zero-shot (Generalist) 18.00 28.68 16.00 25.38 2.00
STEP-7B (Ours) fine-tune (Generalist) 30.28 39.69 26.77 39.36 3.51

Note: SR denotes Success Rate, GC denotes Goal-Conditioned success rate, and \(\Delta\text{SR}\) indicates the performance drop from seen to unseen environments.

In zero-shot cross-task transfer on the AI2-THOR ObjectNav benchmark, STEP-7B significantly outperforms prior methods without any task-specific fine-tuning:

Method Kitchen SR / SPL (%) Living Room SR / SPL (%) Bedroom SR / SPL (%) Bathroom SR / SPL (%) Average SR / SPL (%)
Random 10.40 / 4.80 11.47 / 3.40 13.07 / 8.20 21.60 / 11.13 14.13 / 6.88
SAVN 43.60 / 17.80 21.60 / 7.71 29.20 / 8.65 69.60 / 28.49 40.86 / 16.15
S2P 45.50 / 31.05 28.50 / 19.00 50.40 / 25.29 60.25 / 36.70 46.16 / 28.01
R2C 45.00 / 27.10 46.00 / 26.27 33.00 / 17.31 85.00 / 47.38 52.25 / 29.52
STEP-7B (Ours) 50.00 / 34.38 56.00 / 34.87 55.00 / 33.41 82.00 / 52.43 60.75 / 38.77

Ablation Study

Incremental ablation on the ALFRED validation splits reveals the individual contribution of each component:

Config Val Seen SR (%) Val Seen GC (%) Val Unseen SR (%) Val Unseen GC (%) Note
Qwen2.5-VL-7B (Base) 9.96 18.59 7.87 18.26 Low zero-shot baseline lacking physical control grounding
+ RGB Data 12.75 23.44 13.92 22.38 Direct action mapping initiates primitive motion response
+ Marker 19.68 30.66 18.43 30.61 Large jump (+6.93% / +4.51% SR) validating visual prompt grounding
+ Map Modality 21.12 31.25 19.22 29.85 Moderate gain; simple feature concatenation lacks cross-modal reasoning
+ Chain-of-Thought 28.40 40.35 23.72 33.33 Substantial surge (+7.28% / +4.50% SR); CoT unlocks map-view alignment
+ Primitive (STEP-7B) 30.28 39.69 26.77 39.36 Full model; disentangled spatial primitives bolster unseen generalization
+ Oracle Seg. 38.55 46.76 41.57 50.45 Eliminates mapping noise and segmentation errors
+ Oracle Goal 45.20 52.50 55.69 61.52 Provides ground-truth subgoals, removing semantic parsing ambiguity

Key Findings

  • Chain-of-Thought is the critical catalyst for cross-view alignment: Appending the map modality alone (+Map) yields modest gains (+1.44% Seen SR), showing that VLMs cannot spontaneously infer correspondence between 2D orthographic maps and perspective visual frames. Incorporating explicit CoT reasoning traces (+CoT) drives a massive +7.28% Seen SR and +4.50% Unseen SR surge.
  • Superior execution efficiency and lower computational latency: In comparison to R2C (~1.5k generated tokens, 3.79s latency) and atomic action planners requiring ~50 calls per episode, STEP-7B produces focused CoT explanations (~0.5k tokens) at 0.978 seconds per call with only ~23 decisions per episode, sharply cutting overall computational overhead.
  • Remarkable robustness against environment shifts: While specialist models like E.T. overfit to training environments (dropping by 39.27% in unseen splits), STEP-7B exhibits a minimal performance gap (\(\Delta\text{SR} = 3.51\%\)). Treating each frame as an independent Markov decision state prevents rote memorization of visual layouts.

Highlights & Insights

  • Discrete visual prompting on reachable surfaces: Bypassing brittle continuous pose regression by combining Grounded-SAM floor segmentation with Set-of-Mark waypoint overlays frames navigation as a categorical selection task, ensuring strict kinematic feasibility.
  • Cognitive decoupling of global mapping and local pointing: "Look at the View, Think with the Map, Point to the Goal" elegantly resolves the tension between spatial myopia and cognitive overload, allocating global topology to the abstracted map and obstacle avoidance to egocentric observations.
  • Modular data engine for embodied reasoning: Deconstructing embodied intelligence into action-grounded reasoning traces (Stream A) and primitive spatial competencies (Stream B) provides a reproducible methodology for training generalist embodied agents across diverse robot embodiments.

Limitations & Future Work

  • Performance degradation in dense, cluttered environments: Heavily cluttered indoor scenes with stacked items compromise the 2D orthographic projection of the semantic map due to severe multi-to-one occlusions.
  • High cost of teacher-model distillation: Generating high-fidelity reasoning traces relies on frontier models like Qwen3-VL-235B and GPT-4, incurring significant compute and token costs.
  • Future Directions: The authors plan to investigate richer 3D topological representations for complex layouts and explore reinforcement learning (RL) post-training to diminish reliance on synthetic CoT traces while boosting autonomous decision discovery.
  • vs End-to-end action executors (e.g., RT-2, NaVid): Direct visual-to-atomic action mapping suffers from opaque decision logic and bloated execution steps. STEP employs intermediate point-level planning, where each waypoint encapsulates multi-step locomotion with full CoT interpretability.
  • vs Purely symbolic LLM planners (e.g., SayCan, LLM-Planner): Abstract high-level subgoals create an execution gap without fine-grained spatial awareness. STEP anchors language intent into metric space through dynamic BEV mapping and visual waypoint markers.
  • vs Semantic map navigators (e.g., MapNav, VLFM): Prior map-based systems utilize maps in isolation or rely on handcrafted heuristics. STEP establishes tight cross-attention and CoT alignment between the global map and the first-person perspective.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneers a hybrid BEV-egocentric perception and point-level CoT planning paradigm for embodied agents]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations on ALFRED and AI2-THOR ObjectNav, accompanied by successful real-world LoCoBot deployment]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous formal definitions, cohesive technical narrative, and clean, informative visual illustrations]
  • Value: ⭐⭐⭐⭐⭐ [Sets a benchmark for integrating multimodal foundation models into physically grounded, long-horizon embodied robotics]