Reconstructing 3D Human-Object Interaction via a Unified Triplane Space¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: 3D Vision
Keywords: 3D Human-Object Interaction, Triplane Representation, Implicit Geometric Reconstruction, SMPL-H Pose-fitting, Monocular 3D Reconstruction
TL;DR¶
To tackle the structural disparity between parametric human templates and diverse object geometries in monocular 3D human-object interaction (HOI) reconstruction, this paper introduces a unified triplane feature space, coupling HOIT-VAE implicit priors, camera-aware HOIT-T image-to-triplane prediction, and SMPL-H pose-fitting to deliver balanced, highly accurate interaction meshes.
Background & Motivation¶
Reconstructing 3D human-object interactions (HOI) from a single RGB image is a foundational capability for embodied AI, robotics manipulation, and augmented/virtual reality. However, jointly estimating human and object geometry presents a fundamental structural asymmetry: the human body adheres to a well-constrained statistical topology formalized by parametric body models (e.g., SMPL), whereas interacting objects span boundless categories with highly irregular, non-parametric shapes and varying topological genera. This intrinsic discrepancy creates an acute optimization imbalance for direct vertex regression architectures.
Existing methodologies typically apply local interaction constraints or vertex-level graph formulations. Frameworks like StackFLOW model interactions via vertex-to-vertex offset distributions, while CONTHO introduces explicit contact maps to prevent interpenetration. Although such local contact modeling prevents floating or penetration artifacts, it frequently overfits to local interface points at the expense of global spatial structure. To balance local contact and global coherence, HOI-TG supervises explicit vertex positions via combined graph convolutions and self-attention. Nonetheless, applying uniform positional penalties causes the network to collapse toward the easily optimizable human template, heavily under-constraining diverse objects and yielding severe shape distortion on elongated or small items. Conversely, point-cloud diffusion formulations such as HDM discard object templates but suffer from the inherent disorder of point clouds, requiring multi-step stochastic denoising that proves computationally expensive and unstable.
The core tension is that monocular HOI reconstruction lacks a unified geometric representation that can reconcile topological differences while seamlessly encoding both global spatial arrangement and local surface contact. The paper addresses this gap by decoupling explicit meshes into continuous implicit occupancy fields mapped across orthogonal triplanes, reformulating monocular HOI recovery as a structured image-to-triplane latent prediction task. Core idea: construct a unified triplane feature space that jointly encodes human and object occupancy fields within a shared 3D latent domain, integrating a VAE geometric prior, a camera-aware Transformer, and prior-regularized SMPL-H pose-fitting to eliminate template dependency and structural imbalance.
Method¶
Overall Architecture¶
The reconstruction framework consists of three synergistic stages: latent geometry prior modeling via HOIT-VAE, single-view image-to-triplane prediction via HOIT-T, and geometry-guided SMPL-H pose-fitting. In the first stage, HOIT-VAE ingests sampled surface point clouds with normals, compresses them into a compact triplane latent distribution using cross-attention and self-attention blocks, and decodes them via a convolutional decoder into continuous occupancy grids to extract meshes via Marching Cubes. In the second stage, HOIT-T extracts visual tokens from a frozen DINO backbone, injects closed-form camera embeddings via adaptive Layer Normalization (adaLN), and predicts the triplane tokens aligned with the pre-trained VAE posterior. Finally, to resolve fine-grained hand degradation caused by spatial quantization, an optimization module fits SMPL-H parameters to the decoded human mesh under the regularization of the DPoser-X whole-body diffusion prior.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Single RGB Image + 2D Masks"] --> B["HOIT-T Image Prediction<br/>DINO Tokens + Closed-form Camera Embeddings"]
B --> C["HOIT-VAE Latent Alignment<br/>MSE + Cosine Similarity + KL Divergence"]
C --> D["Triplane Upsampling & Decoding<br/>Convolutional Upsampling + MLP Occupancy"]
D --> E["Dual Mesh Reconstruction<br/>512ยณ Voxel Grid Query + Marching Cubes"]
E --> F["SMPL-H Pose-fitting<br/>Chamfer Distance + DPoser-X Prior"]
F --> G["Output: Spatially Aligned Human & Object 3D Meshes"]
Key Designs¶
1. HOIT-VAE and Semi-continuous Occupancy: Building a Template-free Unified Latent Space Directly predicting high-resolution 3D triplane maps is computationally intractable, while standard binary occupancy modeling causes severe tearing and disconnection on thin object parts. HOIT-VAE addresses this by sampling surface points and normals via Farthest Point Sampling (FPS), augmenting them with Fourier Positional Encoding, and feeding them to an encoder with one cross-attention layer and multiple self-attention layers to produce compact latent parameters. To ensure smooth surface continuity and preserve delicate geometries during rasterization, the occupancy values are modeled with a signed-distance-based semi-continuous formulation: $$ \tilde{o}_i(\mathbf{x}) = \begin{cases} 1, & \text{sdf}_i(\mathbf{x}) < -s \ \frac{1}{2} - \frac{\text{sdf}_i(\mathbf{x})}{2s}, & -s \leq \text{sdf}_i(\mathbf{x}) \leq s \ 0, & \text{sdf}_i(\mathbf{x}) > s \end{cases} $$ with transition threshold \(s = 1/512\). The decoder upsamples the latents through ResNet blocks, allowing an MLP to query projected triplane feature sums and predict dense voxel occupancies, realizing sub-voxel reconstruction for arbitrary shapes without requiring category-specific CAD models.
2. HOIT-T and Camera-aware Modulation: Eliminating Cropping Ambiguities Monocular 3D estimation typically processes tightly cropped images, but cropping alters the principal point and effective field of view, causing monocular depth and scale ambiguities. HOIT-T combines visual tokens from a frozen DINO-ViT-B/14 encoder with an analytical camera embedding. Utilizing the full image intrinsics and bounding boxes, it computes the translation vector \(t_c\) and adjusted intrinsics \(K_\text{crop}\) via an analytical closed-form solver without learnable parameter drift. These geometric parameters are mapped into a 128-dimensional embedding \(\tilde{c}\) to dynamically modulate the transformer tokens via adaptive Layer Normalization (adaLN): $$ \text{adaLN}_c(\mathbf{f}_j) = \text{LN}(\mathbf{f}_j) \cdot (1 + \gamma) + \eta $$ The modulated triplane tokens query visual cues via cross-attention and consolidate spatial relationships through self-attention, anchoring 2D visual cues into metric 3D coordinates.
3. Multi-objective Latent Alignment: Preventing Imbalanced Human Collapse When trained purely on regression objectives, models instinctively bias toward the smooth, low-entropy human distribution, causing object representations to deteriorate. HOIT-T counteracts this by combining mean squared error \(\mathcal{L}_\text{MSE}\), cosine feature similarity \(\mathcal{L}_\text{SIM}\), and KL divergence \(\mathcal{L}_\text{KL}\). The cosine similarity loss strictly aligns the directional vectors between predicted triplane tokens and VAE posterior tokens, penalizing directional feature drift and forcing the network to capture high-frequency object contours and boundary contacts rather than just minimizing bulk Euclidean volume.
4. DPoser-X Guided Pose-fitting: Restoring Articulated Interaction Details While triplane representations successfully capture holistic body mass and irregular object forms, voxel discretization and surface smoothing can blur articulate regions such as open fingers and contact grasp points. To supply standard skeletal parameters, the post-processing stage initializes body and hand parameters using Hand4Whole and optimizes the SMPL-H parameters (pose \(\theta\), shape \(\beta\), orientation \(\psi\), translation \(t_h\), scale \(s\)). The optimization minimizes the Chamfer Distance \(\mathcal{L}_\text{dist}\) against point clouds sampled from the decoded human mesh, coupled with the DPoser-X diffusion body prior \(\mathcal{L}_\text{prior}\) to prevent unnatural joint hyperextension, yielding kinematically valid human rigs anchored to the predicted geometry.
Loss & Training¶
HOIT-VAE is trained on synthetic ProciGen using surface points \(\mathcal{Q}_s\) (20,480 points) and uniform points \(\mathcal{Q}_u\) (20,480 points) with binary cross-entropy and KL divergence: $$ \mathcal{L}\text{VAE} = \sum}, \text{object}}} \mathbb{E{\mathbf{q} \in \mathcal{Q}} \left[ \text{BCE}(\tilde{o}_i(\mathbf{q}), \hat{o}_i(\mathbf{q})) \right] + \lambda\text{KL} \mathcal{L}\text{KL} $$ with \(\lambda_\text{KL} = 10^{-6}\), optimized using AdamW (learning rate \(5\times 10^{-5}\)) for 130K steps on 4 RTX A6000 GPUs. In the second stage, HOIT-T is trained to predict the triplane distribution: $$ \mathcal{L}\text{all} = \lambda_1 \mathcal{L}\text{MSE} + \lambda_2 \mathcal{L}\text{SIM} + \lambda_3 \mathcal{L}_\text{KL} $$ with \(\lambda_1 = 1.0, \lambda_2 = 1.0, \lambda_3 = 10^{-6}\), optimized using AdamW (learning rate \(1\times 10^{-4}\)) for 100K steps.
Key Experimental Results¶
Main Results¶
On the BEHAVE and InterCap benchmark datasets, reconstruction accuracy is evaluated using [email protected] over Chamfer Distance in normalized camera coordinates, alongside contact precision (Contact\(_p\)) and recall (Contact\(_r\)). Baselines include template-based methods (CHORE, CONTHO, HOI-TG) and template-free diffusion (HDM).
| Dataset | Method | Hum. F-score โ | Obj. F-score โ | Comb. F-score โ | Contact\(_p\) โ | Contact\(_r\) โ |
|---|---|---|---|---|---|---|
| BEHAVE | CHORE | 0.3454 | 0.4258 | 0.3966 | 0.587 | 0.472 |
| CONTHO | 0.5690 | 0.3177 | 0.5173 | 0.628 | 0.496 | |
| HOI-TG | 0.5913 | 0.3278 | 0.5354 | 0.662 | 0.554 | |
| HDM | 0.3925 | 0.5049 | 0.4604 | - | - | |
| Ours (w.o. Fit.) | 0.5560 | 0.5391 | 0.5723 | - | - | |
| Ours (w. Fit.) | 0.5199 | 0.5356 | 0.5497 | 0.619 | 0.684 | |
| InterCap | CHORE | 0.4064 | 0.5135 | 0.4687 | 0.339 | 0.253 |
| CONTHO | 0.4261 | 0.2305 | 0.3840 | 0.661 | 0.432 | |
| HOI-TG | 0.4465 | 0.2118 | 0.3953 | 0.700 | 0.473 | |
| HDM | 0.4399 | 0.6072 | 0.5344 | - | - | |
| Ours (w.o. Fit.) | 0.6323 | 0.6192 | 0.6432 | - | - | |
| Ours (w. Fit.) | 0.5882 | 0.6107 | 0.6298 | 0.726 | 0.560 |
Ablation Study¶
Ablation studies on BEHAVE validate grid resolution, semi-continuous formulation, loss components, and camera embeddings.
Table 1: Ablation of HOIT-VAE (Voxel Resolution and Semi-continuous Occupancy) | Resolution | Semi-continuous | Hum. F-score โ | Obj. F-score โ | Comb. F-score โ | |:---:|:---:|:---:|:---:|:---:| | 128 | โ | 0.9505 | 0.9625 | 0.9611 | | 128 | โ | 0.9647 | 0.9746 | 0.9740 | | 256 | โ | 0.9696 | 0.9731 | 0.9738 | | 256 | โ | 0.9848 | 0.9900 | 0.9905 | | 512 | โ | 0.9745 | 0.9814 | 0.9804 | | 512 (Full model) | โ | 0.9858 | 0.9912 | 0.9913 |
Table 2: Ablation of HOIT-T Components (Loss Functions and Camera Embeddings) | Config | \(\mathcal{L}_\text{MSE}\) | \(\mathcal{L}_\text{SIM}\) | \(\mathcal{L}_\text{KL}\) | Cam. Embed. | Hum. F-score โ | Obj. F-score โ | Comb. F-score โ | Note | |:---|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---| | Only MSE | โ | โ | โ | โ | 0.5922 | 0.3885 | 0.5602 | Overfits to human template; poor object score | | MSE + SIM | โ | โ | โ | โ | 0.4838 | 0.4434 | 0.4934 | Lacks distribution prior regularization | | w/o Cam. Embed. | โ | โ | โ | โ | 0.5002 | 0.4794 | 0.5144 | Loses metric scale alignment and focal cues | | Full model | โ | โ | โ | โ | 0.5560 | 0.5391 | 0.5723 | Best balanced performance across targets |
Key Findings¶
- Overcoming Reconstruction Asymmetry: Prior template-based models struggle heavily on objects (HOI-TG scores only 0.3278 on BEHAVE objects), whereas our un-fitted model reaches 0.5391 on objects and achieves 0.5723 on combined score, outperforming all competitors.
- Critical Role of Cosine Loss: Relying solely on MSE drops object accuracy to 0.3885. Introducing cosine similarity enforces angular alignment in feature space, preventing human dominance and stabilizing symmetrical object orientations.
- Thin Surface Protection: The semi-continuous SDF transformation delivers a +1.13% gain in human F-score and +0.98% in object F-score at 512 resolution, preventing hollow mesh artifacts caused by binary threshold truncation.
- Pose-fitting Trade-off: Constraining human surfaces to the lower-dimensional SMPL-H manifold slightly reduces surface F-score (0.5560 vs 0.5199), but surges contact recall on BEHAVE (from 0.554 to 0.684) and eliminates hand blob artifacts.
Highlights & Insights¶
- Unified Spatial Abstraction: Circumvents topological discrepancies between parametric humans and diverse objects by mapping both into orthogonal triplane implicit fields, outperforming discrete point-cloud diffusion in stability and inference speed.
- Analytical Camera Modulation: Uses an exact closed-form geometric solver for camera translations coupled with adaLN modulation, resolving image-crop perspective ambiguities without requiring fragile learned camera regression.
- Extensible Representation: The shared triplane coordinate volume naturally accommodates scene extensions, laying the groundwork for multi-person and multi-object joint spatial reasoning.
Limitations & Future Work¶
- Unseen Object Generalization: HOIT-T image-to-triplane mapping remains bounded by the diversity of its training data; fine geometric recesses on exotic or out-of-distribution tools can experience over-smoothing.
- Severe Occlusion Ambiguity: In cases of heavy mutual occlusion, monocular visual cues lack depth transparency, causing the model to generate plausible yet hallucinated occluded backsides.
- Decoupled Pipeline: SMPL-H fitting functions as a detached post-processing phase; exploring end-to-end differentiable kinematic parameterization represents a promising direction for future work.
Related Work & Insights¶
- vs CONTHO / HOI-TG: Prior vertex-regression methods rely heavily on SMPL topology and fail to generalize across diverse object templates; this paper employs a template-free triplane space that balances deformable human bodies and irregular objects.
- vs HDM / TriDI: Point-cloud diffusion models require iterative sampling steps and complex meshing heuristics; this approach predicts continuous triplanes in a single feed-forward pass, yielding continuous, watertight meshes efficiently.
Rating¶
- Novelty: โญโญโญโญ [Pioneering use of a unified triplane feature space for joint monocular HOI reconstruction, resolving template asymmetry]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive quantitative evaluation across BEHAVE, InterCap, and HODome with detailed ablations]
- Writing Quality: โญโญโญโญโญ [Clear motivation, structured explanations, and mathematically sound formulation]
- Value: โญโญโญโญ [Provides a practical and balanced representation paradigm for embodied interaction and AR/VR reconstruction]