SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts¶
Conference: ECCV 2026
Paper: ECCV Official
Area: 3D Vision / Multimodal VLM / LLM Reasoning
Keywords: 3D spatial reasoning, sparse views, Mixture-of-Experts, multimodal large language models, spatiotemporal manifold sampling
TL;DR¶
To tackle topological disruption and cross-modal contention in sparse-view 3D reasoning, SpaR3D-MoE introduces adaptive spatiotemporal manifold sampling alongside an instruction-pose-aware heterogeneous Mixture-of-Experts, setting a new state of the art on VSI-Bench, ScanQA, and SQA3D from sparse RGB-only inputs.
Background & Motivation¶
Three-dimensional spatial reasoning is a cornerstone of embodied AI, empowering autonomous agents to interpret physical environments, grasp 3D spatial topology, and ground natural language instructions in physical reality. Although Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in 2D visual and video semantics, extending them into the physical 3D realm remains fundamentally bottlenecked by a severe representational gap. Standard visual backbones inherently compress fine-grained metric geometry during semantic alignment, causing current MLLMs to severely hallucinate when estimating physical distances, discerning relative orientations, or planning multi-step navigation routes.
To bridge this representational gap, prior literature has diverged into two principal directions. The first paradigm incorporates explicit 3D structures (such as depth maps, point clouds, or reconstructed meshes) into LLMs using specialized point encoders like PointNet++. While offering explicit geometric grounding, these methods depend rigidly on physical depth sensors or compute-intensive 3D reconstruction, and the intrinsic sparsity of point clouds inevitably discards rich semantic visual details. The second paradigm seeks scalable RGB-only 3D reasoning by deploying pre-trained geometric foundation models (such as VGGT) to extract implicit 3D structures from monocular inputs. However, existing RGB-based approaches suffer from two core limitations: in frame selection, they rely on topology-agnostic heuristics such as uniform downsampling or discrete voxel occupancy, which disrupts the continuous spatiotemporal manifold and overlooks critical transition views; in multimodal fusion, monolithic and static projections indiscriminately merge appearance textures and geometric cues into a shared latent space, triggering destructive cross-modal contention across diverse spatial tasks.
Core idea: reformulate sparse-view 3D spatial reasoning as a dual-stage closed loop of "manifold-preserving sampling + task-adaptive heterogeneous routing," leveraging a quality-gated spatiotemporal graph to extract topologically coherent keyframes and dynamic instruction-pose-conditioned Mixture-of-Experts (MoE) to disentangle cross-modal geometric interactions.
Method¶
Overall Architecture¶
SpaR3D-MoE decomposes the unified sparse-view 3D spatial reasoning problem into two synergistic sub-problems: first, extracting an informative, topologically connected sparse keyframe sequence from continuous observations via information maximization on the spatiotemporal manifold; second, conditionally routing multimodal tokens through an instruction- and pose-guided gating mechanism into specialized geometry-inductive experts for adaptive cross-modal fusion and autoregressive generation.
The input comprises an embodied RGB video stream paired with a spatial language instruction. The framework first applies Adaptive Spatiotemporal Manifold Sampling (ASMS) to distill the long video sequence into a sparse set of informative keyframes. Subsequently, Qwen3-VL encodes the instruction and 2D visual tokens, while VGGT extracts corresponding implicit 3D geometric features along with 6-DoF camera poses. The Instruction-Pose Aware Router (IPAR) synthesizes task semantic intent with camera motion bias to compute Top-K expert activation probabilities. Four architecturally heterogeneous experts then perform distinct operations: holistic additive integration, geometric-semantic cross-attention, hypernetwork-driven dynamic pose adaptation, and gravity-aligned structural filtering. Finally, the aggregated multimodal representations are fed into the LLM decoder to autoregressively generate physical answers, jointly optimized by language modeling cross-entropy and a routing load-balancing penalty.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Long Video Stream + Instruction"] --> B["Adaptive Spatiotemporal Manifold Sampling<br/>Graph Metric & Motion-Aware Quality Gate"]
B --> C["Multimodal Feature Extraction<br/>Visual / Geometric / Pose / Instruction"]
C --> D["Instruction-Pose Aware Router<br/>Cross-Modal Intent Mixing & Pose Bias Injection"]
D --> E["Heterogeneous Geometry-Inductive Experts<br/>E0 Additive / E1 Cross-Attn / E2 Pose Adapter / E3 Gravity Struct"]
E --> F["Adaptive Multimodal Feature Aggregation"]
F --> G["Autoregressive LLM Decoder<br/>Spatially-Grounded Textual Response"]
Key Designs¶
1. Adaptive Spatiotemporal Manifold Sampling: eliminating redundancy while preserving topology To overcome the limitations of uniform sampling, which breaks topological continuity, and heuristic sampling, which ignores trajectory transitions, this design constructs a geometric graph approximating the underlying low-dimensional manifold. After downsampling raw dense frames into candidate frames, a composite distance metric \(D(i, j)\) quantifies edge weights by integrating camera displacement, geometric viewing variations, and temporal intervals: $\(D(i, j) = \bar{D}_{trans}(i,j) + \gamma \bar{D}_{geo}(i,j) + \omega \bar{D}_{tmp}(i,j)\)$ where \(\bar{D}_{trans}\) is the normalized L2 distance between camera optical centers predicted by VGGT to guarantee broad trajectory coverage; \(\bar{D}_{geo}\) calculates the cosine dissimilarity of VGGT implicit geometric embeddings to prevent over-sampling stationary vantage points; and \(\bar{D}_{tmp}\) preserves temporal progression. To prevent farthest point sampling (FPS) from clustering on textureless surfaces such as blank walls, a node-level quality score \(S_i\) evaluates the average L2 norm of the Top-\(K_v\) visual patch tokens, penalizing high instantaneous velocities via an exponential motion-blur decay. The extraction is formulated as quality-gated FPS, filtering noisy observations via a motion-aware quality gate \(M(i)\) while retaining topologically vital keyframes.
2. Instruction-Pose Aware Router: disentangling cross-modal interactions via dynamic dispatching To resolve cross-modal contention caused by static monolithic projection, the router dynamically steers tokens based on task semantics and physical camera motion. The router takes instruction tokens \(q\), 2D visual tokens \(v\), 3D geometric tokens \(g\), and camera pose features \(p\). The semantic components (\(q, v, g\)) are linearly projected and concatenated into an Intent Mixer MLP to capture high-order cross-modal correlations, whereas camera pose features \(p\) are linearly projected directly into the routing logit space as a spatial bias: $\(L = \text{MLP}([\mathbf{W}_q q\,;\,\mathbf{W}_v v\,;\,\mathbf{W}_g g]) + \mathbf{W}_p p\)$ This formulation ensures that token routing is simultaneously informed by whether the user instruction demands fine-grained metric measurement or broad topological planning, as well as the geometric vantage shift between sparse viewpoints. The resulting Top-K activation scores direct tokens into distinct, specialized expert pathways.
3. Heterogeneous Geometry-Inductive Experts: specialized multi-tier cross-modal fusion Rather than utilizing homogeneous feed-forward experts, the system establishes four architecturally distinct experts tailored with distinct inductive biases: - Holistic Representation Expert (E0): Employs a primitive additive transformation \(E_0(v, g) = \text{RMSNorm}(v + g)\). Preserving this simple fusion provides an efficient path for generic background tokens that do not require complex geometric manipulation, shielding higher-order experts from representational saturation. - Geometric-Semantic Cross-Attention Expert (E1): Sets visual tokens as Query and 3D geometric tokens as Key and Value in a multi-head cross-attention layer: \(E_1(v, g) = \text{RMSNorm}(v + \text{CrossAttn}(Q=v, K=g, V=g))\). This establishes explicit spatial grounding between 2D visual regions and 3D geometric descriptors, crucial for complex multi-object spatial queries. - Pose-Conditioned Dynamic Adapter Expert (E2): Addresses large camera baseline transformations across sparse views by utilizing a HyperNet to dynamically map camera pose features \(p\) into low-rank bottleneck weights. These weights adaptively project and restore geometric features \(g\) in dynamic subspaces before injecting them into visual tokens with a learnable residual scaling \(\alpha\). This acts as an implicit coordinate transformer aligning disjointed views. - Gravity-Aligned Structural Expert (E3): Deploys learnable structural probes \(\Phi \in \mathbb{R}^{N_p \times D}\) to represent canonical gravity-aligned surfaces (e.g., floor planes). Geometric features are projected into a structural subspace, filtered through probe-affinity masks to remove unaligned noise, and fused with visual tokens via relational mapping to anchor reasoning in a unified physical coordinate system.
Loss & Training¶
The framework is optimized end-to-end using a joint objective combining language generation cross-entropy and a routing load-balancing loss: $\(\mathcal{L} = \mathcal{L}_{gen} + \lambda_{moe} \mathcal{L}_{moe}\)$ where the auxiliary penalty is formulated as \(\mathcal{L}_{moe} = N_e \sum_{i=1}^{N_e} \bar{P}_i \rho_i\) with \(N_e = 4\) experts. \(\bar{P}_i\) denotes the batch-wide mean routing probability for expert \(i\), and \(\rho_i\) represents the actual dispatch frequency to expert \(i\). This objective penalizes uneven routing distributions, preventing expert collapse and ensuring stable convergence across all specialized pathways.
Key Experimental Results¶
Main Results¶
On VSI-Bench, which spans 9 spatial reasoning tasks, SpaR3D-MoE achieves an overall average of 63.5 using only 32 sparse frames, surpassing all proprietary and open-source baselines:
| Model | Input Modality / Frames | Avg. | Route Plan | Rel. Dir. | Abs. Dist. | Room Size |
|---|---|---|---|---|---|---|
| GPT-4o | API / Dense | 34.0 | 31.5 | 41.3 | 5.3 | 38.2 |
| Gemini-1.5-Pro | API / ~85 frames | 45.4 | 36.0 | 46.3 | 30.9 | 43.6 |
| InternVL3-78B | Open-source / Dense | 48.5 | 28.9 | 39.5 | 53.7 | 39.5 |
| Qwen3VL-8B | Open-source / Dense | 55.7 | 32.5 | 46.3 | 45.8 | 60.3 |
| Spatial-MLLM-4B | Spatial-aware / Dense | 48.4 | 33.5 | 46.2 | 34.8 | 45.1 |
| SpaR3D-MoE (Ours) | RGB-only / 32 frames | 63.5 | 44.0 | 70.1 | 48.0 | 69.6 |
On 3D embodied QA benchmarks, SpaR3D-MoE achieves 30.4 EM@1 and 101.5 CIDEr on the ScanQA validation set, outperforming explicit 3D point cloud models like Video-3D LLM (30.1 EM@1) and 3D-LLaVA (92.6 CIDEr). On the SQA3D test set, it achieves a record 58.3 EM@1 among video-input methods.
Ablation Study¶
Ablations on VSI-Bench isolate the specific contributions of each expert and routing component:
| Config | Avg. | Route Plan | Rel. Dir. | Abs. Dist. | Note |
|---|---|---|---|---|---|
| SpaR3D-MoE (Full Model) | 63.5 | 44.0 | 70.1 | 48.0 | 32 frames with ASMS and all 4 experts |
| w/o E0 (Holistic Expert) | 61.8 | 41.3 | 65.6 | 46.3 | Standard tokens dilute higher-order capacity |
| w/o E1 (Cross-Attention Expert) | 62.8 | 43.0 | 69.7 | 47.3 | Degrades fine-grained 2D-3D alignment |
| w/o E2 (Pose Adapter Expert) | 59.3 | 38.8 | 60.4 | 42.2 | Most severe drop (-4.2); breaks sparse view alignment |
| w/o E3 (Structural Expert) | 62.4 | 40.0 | 68.6 | 47.5 | Weakens topological routing and canonical scales |
| Base Router (Vision + Geometry) | 61.2 | 39.0 | 67.5 | 46.0 | Lacks task instruction and pose bias |
| + Pose Bias (\(p\)) | 62.5 | 42.0 | 68.5 | 47.2 | Critical +1.3 gain for viewpoint-sensitive tasks |
| + Query Guidance (\(q\)) | 61.8 | 41.5 | 66.0 | 46.5 | Semantic intent provides +0.6 boost |
In sampling comparisons, ASMS achieves 63.5 at 32 frames, yielding a 10.6% relative gain in Route Plan over uniform sampling (44.0 vs 39.8). At 16 frames, ASMS retains a competitive 61.9 average score, outperforming uniform sampling at the same density (61.2).
Key Findings¶
- Pose Adapter E2 is indispensable under sparse viewpoints: Ablating E2 causes the sharpest overall decline (-4.2 points), with Relative Direction dropping by 9.7 points and Route Plan by 5.2 points, demonstrating that dynamic pose adaptation acts as an essential implicit coordinate transformer across large baselines.
- Retaining baseline expert E0 prevents representational degradation: Removing additive expert E0 leads to a 1.7 point decrease, confirming that providing a low-complexity pathway for generic tokens protects specialized experts from over-processing raw features.
- Camera pose bias is a critical routing driver: Integrating pose features into the routing logits increases overall performance by 1.3 points over the base router, showing that physical viewpoint dynamics provide vital conditioning for expert dispatching.
Highlights & Insights¶
- Pioneering MoE for multimodal 3D spatial fusion: Replaces rigid, monolithic linear projectors with architecturally heterogeneous experts, disentangling metric estimation, coordinate transformation, and visual grounding.
- Topological manifold graph sampling with motion quality gating: Treats frame selection as quality-aware farthest point sampling on a spatiotemporal graph, capturing critical navigation points while actively suppressing blur.
- Outperforming explicit 3D models with sparse RGB views: Surpasses dense point-cloud methods on ScanQA and SQA3D, proving that implicit geometric representations paired with adaptive routing unlock rich spatial awareness without depth sensors.
Limitations & Future Work¶
- Reliance on offline global video trajectories: ASMS requires building the spatiotemporal graph over full candidate sequences offline, limiting immediate deployment in streaming low-latency robotics control.
- Sensitivity to monocular geometric estimation quality: The pipeline relies on VGGT for implicit geometry and camera poses; adverse conditions such as severe specular reflections or dynamic moving objects may induce cascading estimation errors.
- Future directions: Adapting the manifold sampling and routing framework to online, causal video streams, and exploring reinforcement learning with self-reflective reasoning chains to further boost long-horizon multi-room navigation.
Related Work & Insights¶
- vs Explicit 3D MLLMs (e.g., 3D-LLM, Chat-Scene, Video-3D LLM): Explicit approaches demand calibrated RGB-D sensors or dense point clouds, incurring significant acquisition overhead and sacrificing fine textures. SpaR3D-MoE uses only sparse RGB videos, matching or surpassing point-cloud models through implicit geometry and expert routing.
- vs RGB-based Spatial MLLMs (e.g., Spatial-MLLM, VG LLM): Prior RGB-only methods use rigid uniform frame sampling and monolithic static fusion, leading to topological disruption and cross-modal contention. SpaR3D-MoE maintains topological connectivity via ASMS and provides adaptive multi-pathway integration via HGI-MoE.
Rating¶
- Novelty: โญโญโญโญโญ [First to introduce heterogeneous geometry-inductive MoE for 3D spatial reasoning, paired with graph-based manifold sampling]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluations across VSI-Bench, ScanQA, and SQA3D with thorough ablations of all architectural components]
- Writing Quality: โญโญโญโญโญ [Clear structural formulation, rigorous mathematics, and well-designed architectural illustrations]
- Value: โญโญโญโญโญ [Establishes a highly scalable and cost-effective paradigm for embodied 3D reasoning from RGB-only observations]