Learning Generatable Mutual Distance for Scene-Aware Human Motion Generation¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: PDF
Area: 3D Vision
Keywords: Scene-Aware Human Motion Generation, Mutual Distance, Interaction Representation, Diffusion Models, Vision-Language Models
TL;DR¶
Addressing the spatial ambiguity of mutual distance in stochastic generation due to the absence of an absolute spatial coordinate frame, this paper proposes MDNet—a two-stage diffusion framework that introduces vision-language model (VLM) semantic spatial anchors to resolve geometric ambiguity and synthesizes smooth mutual distance trajectories in the frequency domain to guide physically plausible, temporally coherent scene-aware human motions.
Background & Motivation¶
Language-guided scene-aware 3D human motion generation aims to synthesize full-body skeletal dynamics that faithfully follow textual intentions while conforming strictly to the physical geometry of complex 3D environments. However, this task has long been constrained by the dual challenge of abstract, ambiguous natural language descriptions and rigid physical scene constraints. Existing generative approaches suffer from severe representation bottlenecks: implicit feature fusion methods like HUMANISE tightly entangle motion synthesis with scene reasoning, which easily obscures task-relevant spatial regions and leads to semantic misalignment and "floating" artifacts; conversely, explicit grounding methods based on affordance contact maps focus only on sparse human-scene contact points, completely neglecting the spatiotemporal motion evolution of non-contacting limbs and causing extensive body penetrations, severe jitter, and unnatural freezes.
Mutual distance integrates per-joint signed distances to the scene surface (SDF) with per-basis-point Euclidean distances from fixed scene anchors to the human body, providing an expressive representation for full-body spatiotemporal human-scene interaction. Nevertheless, in stochastic generation, mutual distance encodes only relative spatial configurations and inherently lacks an absolute global reference frame. While motion forecasting can naturally disambiguate relative geometry using historical frames, no such reference exists during text-conditioned stochastic generation. Consequently, two completely distinct motions performed in the same scene (e.g., walking toward a chair on the left versus walking toward a table on the right) yield nearly identical mutual distance curves. This fundamental ambiguity prevents mutual distance from independently determining motions that satisfy language-specified targets. Furthermore, standard diffusion models operating in the time domain tend to accumulate high-frequency denoising noise, resulting in temporal discontinuities and jittery interactions.
To overcome this core contradiction, this paper departs from end-to-end direct pose regression by decoupling spatial localization from motion synthesis. It introduces a vision-language model to perform chain-of-thought spatial reasoning over a bird's-eye-view scene representation to establish absolute start and end spatial references, while modeling mutual distance in the frequency domain to suppress high-frequency disturbances. Core idea: Resolve the spatial ambiguity of mutual distance via VLM-predicted semantic anchors to provide an absolute spatial reference frame, and develop a two-stage diffusion framework, MDNet, that models low-frequency interaction trajectories in the DCT frequency domain to guide physically plausible, temporally coherent 3D full-body human motion generation.
Method¶
Overall Architecture¶
MDNet consists of two cascaded stages: the first stage is the Mutual Distance Generator (MDG), and the second stage is the Mutual Distance-based Motion Generator (MDMG). The system first deploys a pretrained 2D vision-language model (Gemini-1.5-Pro) to execute chain-of-thought (CoT) reasoning over a rendered top-down bird's-eye-view (BEV) image of the 3D scene, locating the interaction target object and estimating plausible 2D start and end pelvis positions (semantic anchors), which are subsequently unprojected back into 3D coordinates. In the MDG stage, noisy mutual distance sequences are mapped via discrete cosine transform (DCT) into a truncated low-frequency domain, combined with scene SDF features, CLIP text embeddings, and semantic anchor positional encodings for diffusion denoising, and inverted via inverse DCT (IDCT) to reconstruct temporally smooth mutual distance sequences \(\hat{M}\). In the MDMG stage, guided by predicted \(\hat{M}\), scene geometric features, and semantic conditions, a conditional diffusion model synthesizes full-body SMPL-X pose sequences \(\hat{X}\) under dual differentiable geometric consistency losses derived from scene SDF interpolation.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: 3D Scene SDF & Text Instruction"] --> B["VLM Semantic Anchor Prediction<br/>BEV reasoning for start/end spatial anchors"]
B --> C["Frequency-Domain Interaction Distance Generation<br/>DCT transform & spatial-semantic diffusion"]
C --> D["Distance-Guided Motion Generation & Consistency Constraints<br/>Interaction fusion & dual distance consistency losses"]
D --> E["Output: Physically Plausible Full-Body Motion"]
Key Designs¶
1. VLM Semantic Anchor Prediction: Providing Absolute Spatial Grounding for Relative Interactions Mutual distance inherently lacks global coordinate anchors, directly leading to multiple ambiguous solutions, while fixed-vocabulary 3D object detectors exhibit limited generalization across open-vocabulary descriptions and novel objects. To resolve this bottleneck, this work introduces an efficient zero-shot semantic anchor localization strategy utilizing bird's-eye-view (BEV) rendering and general 2D vision-language models (VLMs). The 3D scene SDF is projected from a virtual orthographic overhead camera into a 2D BEV map that comprehensively captures walkable floor spaces and obstacle layouts. Gemini-1.5-Pro is prompted via chain-of-thought (CoT) reasoning through four explicit steps: (1) recognize all scene objects and depict room context; (2) locate the target object referenced in the text prompt; (3) infer a plausible start pelvis position in open space; and (4) infer a reachable end pelvis position along the edge of the target object. The inferred 2D points are lifted into 3D spatial anchors \(P_{se} \in \mathbb{R}^{2 \times 2}\). This design avoids the computational complexity of 3D-LLM feature reconstruction and supplies the missing spatial reference frame, completely eliminating geometric ambiguity in stochastic mutual distance generation.
2. Frequency-Domain Interaction Distance Generation: Suppressing High-Frequency Jitter and Decoupling Interaction Dynamics The mutual distance representation \(M = [D; B] \in \mathbb{R}^{U \times (J+K)}\) comprises per-joint signed distances \(d_{ju}\) (negative inside obstacles, positive outside) and per-basis-point Euclidean distances \(b_{ku}\) from \(K=150\) fixed scene anchor points to human joints. In time-domain diffusion denoising, random noise injection causes severe temporal oscillations that compromise human-scene physical plausibility. The MDG incorporates the Discrete Cosine Transform (DCT) applied independently along the temporal dimension for each joint and basis channel: $\(Z^{(t)} = C M^{(t)}, \quad \tilde{Z}^{(t)} = C_d M^{(t)}\)$ where \(C_d \in \mathbb{R}^{d \times U}\) retains the first \(d=50\) rows of the orthogonal cosine basis matrix. Because human motion energy is predominantly concentrated in low-frequency bands, preserving only the first 50 coefficients effectively filters out high-frequency noise. Inside the Spatial-semantic Transformer, positional embeddings \(F_{pos}\) from semantic anchors are merged with frequency latents using attention fusion \(Att_{fp}\); 3D-CNN scene features \(F_{sce}\) are integrated via cross-attention \(Att_{es}\); and a transformer decoder fuses text features \(F_{text}\), diffusion timestep embeddings \(F_{step}\), and learnable queries \(Q_m\) to predict denoised DCT coefficients \(\hat{Z}^{(t)}\), which are transformed back via \(\hat{M}^{(t)} = C_d^\top \hat{Z}^{(t)}\) into smooth mutual distance trajectories.
3. Distance-Guided Motion Generation & Consistency Constraints: Enforcing Closed-Loop Human-Scene Geometric Realism The second-stage MDMG network synthesizes SMPL-X human body poses \(X\) using predicted mutual distance trajectories \(\hat{M}\) as a structured geometric scaffold. To prevent unnatural drifts between motion synthesis and scene geometry, MDMG first fuses mutual distance embeddings with scene features \(F_{sce}\) via an attention module \(Att_{ms}\), followed by a transformer encoder to reconstruct motion sequence \(\hat{X}\). Crucially, the model utilizes the pre-computed scene signed distance volume to enforce dual differentiable geometric consistency. During training, predicted joint positions are trilinearly interpolated in the SDF volume to derive instantaneous predicted signed distances \(\hat{d}_{ju}\) and basis point distances \(\hat{b}_{ku}\), which are penalized against ground truth distances using \(\ell_1\) consistency losses: $\(\mathcal{L}_{joint} = \frac{1}{UJ}\sum_{u=1}^{U}\sum_{j=1}^{J} |\hat{d}_{ju} - d_{ju}|, \quad \mathcal{L}_{basis} = \frac{1}{UK}\sum_{u=1}^{U}\sum_{k=1}^{K} |\hat{b}_{ku} - b_{ku}|\)$ This mechanism compels the diffusion backbone to satisfy physical contact constraints and obstacle boundaries explicitly, eliminating artifacts such as foot sliding, joint penetration, and hovering.
Loss & Training¶
MDNet is trained in two decoupled stages: - The first stage optimizes the MDG diffusion network using an \(\ell_1\) loss on the truncated DCT coefficients: $\(\mathcal{L}_{MDG} = \mathbb{E}_{\tilde{Z}^0, t} \left[ \|\tilde{Z}^0 - G_\theta(\tilde{Z}^{(t)}, t, S, F_{text}, P_{se})\|_1 \right]\)$ - The second stage optimizes the MDMG diffusion network with a composite objective combining motion pose reconstruction and geometric consistency: $\(\mathcal{L}_{MDMG} = \mathcal{L}_{motion} + \mathcal{L}_{joint} + \mathcal{L}_{basis}\)$ where \(\mathcal{L}_{motion} = \mathbb{E}[\|X^0 - G_\phi(X^{(t)}, t, \hat{M}, S, F_{text}, P_{se})\|_1]\). Both stages are optimized using AdamW with a learning rate of \(1 \times 10^{-4}\) and a batch size of 32 on a single NVIDIA A100 GPU. During test-time inference, 5 anchor candidate pairs are sampled per text prompt via Gemini-1.5-Pro to support diverse motion generation.
Key Experimental Results¶
Main Results¶
On the HUMANISE indoor benchmark, models are evaluated across semantic alignment (goal distance), physical plausibility (contact score, non-collision score, average penetration penemean, maximum penetration penemax), and motion quality/diversity (Average Pairwise Distance APD, Temporal Motion Variance TMV). Quantitative comparisons are reported in Table 1.
| Method | goal dist. (m) ↓ | contact (%) ↑ | non-collision (%) ↑ | penemean (cm) ↓ | penemax (cm) ↓ | APD ↑ | TMV → |
|---|---|---|---|---|---|---|---|
| Ground Truth (G.T.) | 0.014 | - | - | - | - | 0.000 | 0.1875 |
| Humanise (NeurIPS'22) | 0.422 | 84.06 | 99.77 | 1.283 | 2.656 | 4.094 | 0.2055 |
| Cen et al. (CVPR'24) | 0.384 | 86.36 | 99.60 | 1.144 | 2.551 | 2.073 | 0.1908 |
| Wang et al. (CVPR'24) | 0.156 | 95.86 | 99.69 | 1.367 | 14.768 | 2.597 | 0.1486 |
| MDNet (Ours) | 0.057 | 95.05 | 99.81 | 0.635 | 1.849 | 2.235 | 0.1870 |
Ablation Study¶
Systematic ablations on the HUMANISE test set examine the impact of mutual distance components, semantic position anchors, and frequency-domain DCT modeling, as shown in Table 2.
| Config | goal dist. ↓ | contact ↑ | non-collision ↑ | penemean ↓ | penemax ↓ | APD ↑ | TMV → | Note |
|---|---|---|---|---|---|---|---|---|
| Full model (Ours) | 0.057 | 95.05 | 99.81 | 0.635 | 1.849 | 2.235 | 0.1870 | full two-stage model with DCT |
| w/o mutual distance | 0.298 | 86.54 | 99.77 | 1.921 | 6.680 | 3.903 | 0.3573 | large penetration increase and erratic dynamics |
| w/o per-joint signed dist. | 0.059 | 93.98 | 99.80 | 0.992 | 3.998 | 2.229 | 0.2524 | degraded local surface contact and penetration |
| w/o per-basis point dist. | 0.158 | 94.66 | 99.76 | 0.701 | 2.829 | 2.361 | 0.2083 | degraded global spatial navigation accuracy |
| w/o semantic anchors | 0.365 | 90.37 | 99.79 | 0.767 | 4.010 | 4.433 | 0.2383 | relative distance ambiguity leads to target drift |
| w/o DCT & IDCT | 0.059 | 94.96 | 99.80 | 0.741 | 3.246 | 2.178 | 0.2046 | time-domain diffusion exhibits higher jitter |
| Rand-proj. mutual dist. | 0.147 | 94.42 | 99.70 | 1.343 | 3.928 | 2.369 | 0.2145 | random orthogonal projection destroys structure |
Cross-Dataset Generalization¶
Evaluating the zero-shot transfer capability on the real-world scanned 3D dataset PROX without any fine-tuning produces the results detailed in Table 3.
| Method | goal dist. (m) ↓ | non-collision (%) ↑ | contact (%) ↑ | penemean (cm) ↓ | penemax (cm) ↓ |
|---|---|---|---|---|---|
| Cen et al. (CVPR'24) | 0.281 | 99.65 | 86.62 | 0.963 | 1.921 |
| Wang et al. (CVPR'24) | 0.283 | 99.69 | 95.67 | 0.816 | 2.237 |
| MDNet (Ours) | 0.054 | 99.87 | 95.93 | 0.270 | 0.654 |
Key Findings¶
- Crucial Role of Mutual Distance Representation: Removing mutual distance causes the most severe performance degradation across all physical metrics: average penetration increases from 0.635 cm to 1.921 cm (over 200% increase), and TMV worsens from 0.1870 to 0.3573, demonstrating that explicit full-body human-scene interaction representation is essential for physical plausibility.
- Anchor Disambiguation Effect: Omitting VLM semantic anchors elevates the goal distance error dramatically from 0.057 m to 0.365 m, confirming that relative mutual distance requires absolute spatial grounding to achieve accurate target navigation.
- Frequency-Domain Regularization for Temporal Coherence: Truncating to 50 low-frequency DCT coefficients provides the best trade-off between compact compression and reconstruction fidelity. Removing DCT modeling increases maximum penetration from 1.849 cm to 3.246 cm, accompanied by prominent temporal jitter.
- Superior Zero-Shot Cross-Domain Robustness: On the unseen PROX dataset, fixed-vocabulary 3D detection baselines suffer significant domain degradation, whereas MDNet leverages VLM spatial reasoning and normalized mutual distance constraints to achieve a penetration depth of only 0.270 cm (less than 30% of baseline penetration).
Highlights & Insights¶
- Repurposing Relative Interaction for Generative Tasks: While mutual distance had previously only been applied to motion forecasting with known history, this paper identifies its core bottleneck in stochastic generation—the lack of an absolute reference frame—and successfully unlocks its generative capability through VLM anchor conditioning.
- Harmonious Marriage of DCT and Diffusion Models: Mapping continuous interaction trajectories to the frequency domain natively establishes a low-pass filter against high-frequency diffusion denoising noise, elegantly resolving temporal contact jitter and erratic trajectory variations.
- Cost-Effective BEV-VLM Spatial Reasoning: Instead of training complex and heavy 3D-LLMs, the framework extracts accurate 3D spatial priors by querying general-purpose 2D VLMs over top-down BEV projections using chain-of-thought prompting, offering an effective and reproducible paradigm for embodied AI navigation and interaction.
Limitations & Future Work¶
- Author-Acknowledged Limitations: The overall generation quality depends on the accuracy of the VLM-inferred semantic anchors. Under highly cluttered scenes where severe BEV occlusions occur, or when the VLM hallucinates incorrect coordinates, downstream motion generation may deviate toward infeasible paths.
- Unaddressed Potential Blindspots: The model currently assumes static 3D environments and does not yet handle dynamic object manipulation (such as carrying chairs or opening doors) or multi-person interactions, where dynamic distance field computation would introduce substantial computational overhead.
- Promising Improvement Directions: Expanding VLM inference from terminal endpoints to intermediate topological waypoints could provide continuous navigation guidance. Furthermore, integrating physics simulation engines (e.g., Isaac Gym or MuJoCo) for reinforcement learning post-training could realize closed-loop reactive contact dynamics.
Related Work & Insights¶
- vs Humanise (NeurIPS 2022): Humanise relies on a cVAE that implicitly fuses scene point clouds with text tokens, causing motion synthesis to entangle with spatial reasoning, which leads to foot skating and floating. In contrast, MDNet decouples interaction generation from motion synthesis, reducing goal distance error by 86.5% (0.057 m vs. 0.422 m).
- vs Cen et al. (CVPR 2024): Cen et al. employ a fixed-vocabulary 3D detector to provide bounding box anchors, which struggles to generalize under domain shifts and open-vocabulary descriptions. MDNet leverages open-vocabulary VLM reasoning on BEV images, demonstrating significantly better zero-shot transfer on PROX.
- vs Wang et al. (CVPR 2024): Wang et al. constrain motions via affordance contact maps, which only govern local contact points while leaving the rest of the body unconstrained, resulting in severe intra-body distortion and penetrations (penemax reaches 14.768 cm). MDNet's mutual distance enforces full-body joint constraints, slashing maximum penetration to 1.849 cm.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates the spatial ambiguity of mutual distance in stochastic generation and resolves it elegantly via VLM semantic anchoring and frequency-domain diffusion.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across semantic goal alignment, penetration, and temporal variance, backed by extensive ablations and zero-shot PROX transfer.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulations, clear problem framing, and intuitive visual illustrations.
- Value: ⭐⭐⭐⭐⭐ Establishes an inspiring and robust explicit interaction representation baseline for scene-aware human motion generation and embodied robotics.