Skip to content

EgoPHI: Estimating 3D Hand-Object Contact and Force from Egocentric Vision

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/eth-siplab/EgoPHI
Project: https://siplab.org/projects/EgoPHI
Area: Robotics
Keywords: Egocentric Vision, Hand-Object Interaction, 3D Contact Estimation, Contact Force Estimation, Physical Interaction Reasoning

TL;DR

EgoPHI is the first vision-based model that jointly predicts dense per-vertex 3D contact regions and interaction forces on both hands and articulated object meshes from a single egocentric RGB image and object template geometry, leveraging physics-based simulation supervision and real-world FTIR sensor validation.

Background & Motivation

Understanding hand-object interaction (HOI) from an egocentric perspective is foundational for modeling human physical engagement with the environment and enabling embodied robots to learn dexterous manipulation via visual imitation. Recent vision-based HOI methods have made substantial strides in hand pose reconstruction, whole-body contact estimation, and fine-grained hand-object contact localization (such as DECO and HACO). However, spatial contact localization alone cannot capture the underlying physical dynamics of an interactionโ€”stable grasping, balancing, and coordinated manipulation are fundamentally governed by contact forces, including their magnitudes and 3D spatial vector orientations across contact interfaces.

Inferring physical forces from a single egocentric RGB frame presents formidable challenges. Pervasive self-occlusion by interacting hands, weak visual cues (such as subtle skin blanching or minute surface deformation), and inherent monocular depth ambiguity severely hinder the accurate estimation of force magnitudes and directional vectors in 3D space. Compounding these hurdles, dense ground-truth force annotations are virtually impossible to capture at scale in natural real-world interactions. Consequently, existing force estimators are either confined to 2D image-space pressure heatmaps (e.g., PressureVision) or restricted to flat rigid contact surfaces (e.g., EgoPressure), falling short of addressing the diverse, non-planar, and articulated objects encountered in daily life.

To resolve the core tension between the lack of scalable ground-truth force data and severe 3D hand-object occlusions, this paper presents a pipeline combining physics simulation supervision with graph-based geometric interaction reasoning. Core idea: build an automated physics simulation pipeline to synthesize dense per-vertex force annotations from existing motion datasets, and introduce an iterative pose refinement and graph-based interaction network that unifies intra-mesh topology and inter-mesh cross-attention to jointly decode 3D contact and forces across hands and objects.

Method

Overall Architecture

EgoPHI takes as input a single egocentric RGB image \(I\), estimated 3D vertex coordinates of the left and right hand meshes \(P_L, P_R \in \mathbb{R}^{V_H \times 3}\), and the CAD template mesh of the target object \(O_{\text{template}}\). The end-to-end inference consists of three stages: visual and geometric feature extraction with cross-modal projection, iterative object pose refinement (IRM) via graph interaction blocks, and joint 3D contact and force estimation on the aligned meshes.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Egocentric RGB Image + Hand Meshes + Object Template"] --> B["Feature Extraction & Cross-Modal Projection<br/>ViT feature sampling + local & relative geometric embeddings"]
    B --> C["Iterative Pose Refinement Module (IRM)<br/>Intra-mesh GAT + Inter-mesh Cross-Attention for cyclic (R, t) update"]
    C --> D["Aligned Coordinate Space Reconstruction<br/>Object reprojection with predicted pose & full feature extraction"]
    D --> E["Joint Contact & Force Decoding<br/>Per-vertex contact probabilities + 3D force magnitude & direction"]
    E --> F["Output: Dense 3D Contact Probabilities & Force Vectors on Hands and Object"]

In the initial stage, a ViT-B/16 backbone extracts patch-level feature maps. Hand and object template vertices are projected onto the image plane to sample visual features via bilinear interpolation, which are then fused with 3D coordinates. In the second stage, an iterative refinement loop optimizes the object's 6D rigid pose in the camera frame. In the third stage, updated hand and object representations pass through dedicated interaction blocks to predict per-vertex contact probabilities and 3D force vectors.

Key Designs

1. Physics-Based Contact and Force Simulation: Dense Supervision without Physical Sensors To overcome the scarcity of physical multi-axis force sensing arrays for dynamic grasping, the authors build a simulation pipeline using the SOFA framework to augment the ARCTIC dataset with dense force annotations. Treating hand and object meshes as quasi-rigid bodies, virtual linear springs are instantiated across interacting surfaces. When mesh vertices enter a predefined contact proximity zone, a penalty-based restoring force is computed proportional to the penetration displacement \(d_i\): $\(\mathbf{f}_i = -k d_i \mathbf{n}_i\)$ where \(k=10\) denotes the uniform stiffness constant and \(\mathbf{n}_i\) is the surface normal at vertex \(i\). To suppress spurious non-contact responses, a binary contact mask \(M\) filters the raw force field to obtain physically valid ground-truth forces \(\tilde{\mathcal{F}} = M \odot \mathcal{F}\). During supervised training, the force vector is decoupled into its scalar magnitude \(|f_i|\) and unit direction vector \(u_i = f_i / |f_i|\), stabilizing optimization against extreme force spikes.

2. Graph-Based Interaction Blocks: Combining Manifold Topology with Inter-Mesh Proximity Physical interaction is simultaneously constrained by local surface continuity and mutual spatial proximity between separate geometries. The authors design modular Graph-Based Interaction Blocks applied across both pose refinement and force estimation. Each block comprises: - Intra-mesh Graph Attention Networks (GAT): propagate features strictly along mesh triangular edges to enforce local structural smoothness; - Inter-mesh Cross-Attention: computes dynamic all-to-all attention between the left hand, right hand, and object vertices to capture cross-entity proximity and spatial configurations under occlusion. This architecture allows the model to propagate geometric evidence from visible hand regions to occluded object boundaries.

3. Iterative Pose Refinement Module: Resolving Geometric Misalignment under Occlusion Directly predicting absolute object pose under severe hand occlusion frequently causes interpenetration or floating artifacts that invalidate subsequent physical contact reasoning. The Iterative Refinement Module (IRM) starts from an initial zero translation and identity rotation, iteratively updating the pose over \(N=3\) steps. In each step, static hand features interact with dynamic object features derived from RoI-aligned visual features and hand-relative 3D distances. The refined vertex features are pooled to predict 6D incremental rotation \(\Delta R_i \in \mathbb{R}^6\) and translation \(\Delta t_i \in \mathbb{R}^3\), updating the camera-frame pose via: $\(R_i = \Delta R_i \cdot R_{i-1}, \quad t_i = \Delta t_i + t_{i-1}\)$ This cyclic geometric alignment reduces the mean per-vertex error (MPV) by over 70%, establishing accurate spatial contact boundaries for downstream force estimation.

4. Contact-Gated Joint Force Prediction: Enforcing Physical Consistency Conditioned on the refined object pose \((R_N, t_N)\), updated vertex embeddings pass through a secondary stack of interaction blocks. The output heads jointly estimate contact probabilities \(C\) alongside force magnitudes and directions. Because physical interaction forces cannot exist without geometric contact, predicted force vectors are explicitly modulated by predicted contact probabilities via element-wise gating: $\(\tilde{\mathbf{f}}_i = C_i \cdot (|f_i| \mathbf{u}_i)\)$ This gating mechanism suppresses spurious floating forces in visually ambiguous or non-contact regions, ensuring strict physical causality between visual contact evidence and predicted force vectors.

Loss & Training

The joint training objective combines six weighted loss terms: $\(\mathcal{L}_{\text{total}} = \lambda_c \mathcal{L}_{\text{contact}} + \lambda_m \mathcal{L}_{\text{force-mag}} + \lambda_v \mathcal{L}_{\text{force-vec}} + \lambda_t \mathcal{L}_{\text{trans}} + \lambda_r \mathcal{L}_{\text{rot}} + \lambda_p \mathcal{L}_{\text{verts}}\)$ Contact estimation \(\mathcal{L}_{\text{contact}}\) uses weighted binary cross-entropy, vertex-level class-balanced (VCB) loss, dataset-level contact regularization, and a local smoothness term. Force magnitude \(\mathcal{L}_{\text{force-mag}}\) employs contact-weighted \(L_2\) loss with mean magnitude regularization, while direction \(\mathcal{L}_{\text{force-vec}}\) minimizes contact-weighted cosine distance. Pose objectives supervise vertex positions, geodesic rotations, and Euclidean translations. Optimization is conducted using Adam with learning rate \(1 \times 10^{-4}\) and batch size 8.

Key Experimental Results

Main Results

EgoPHI is evaluated on the in-distribution ARCTIC test split (s05) and the out-of-distribution H2O dataset (subject4_ego), comparing against the 3D contact model HACO (including an ARCTIC-finetuned 3D force variant) and the 2D image-space force model PressureVision.

Dataset Model Hand Force MAE [N] โ†“ Hand Force RMSE [N] โ†“ Hand Force vIoU โ†‘ Object Force MAE [N] โ†“ Object Contact F1 โ†‘ Hand Contact F1 โ†‘
ARCTIC s05 HACO 6.62 7.29 1.0 not supported not supported 0.196
ARCTIC s05 EgoPHI w/o IRM 4.80 5.94 0.7 5.01 0.038 0.057
ARCTIC s05 EgoPHI (Ours) 4.03 5.06 2.2 4.42 0.060 0.190
H2O s4_ego HACO 6.37 7.13 0.7 not supported not supported 0.198
H2O s4_ego EgoPHI w/o IRM 4.65 5.87 0.8 4.70 0.030 0.026
H2O s4_ego EgoPHI (Ours) 5.16 6.62 0.6 3.88 0.037 0.097

In 2D projected force evaluations, EgoPHI achieves 15.4% Volumetric IoU on ARCTIC, outperforming PressureVision (6.6%) by more than \(2\times\).

Ablation Study & Real-World Validation

The Iterative Refinement Module (IRM) is ablated against the full pipeline. To validate sim-to-real transfer, an acrylic FTIR-based touch-and-force sensing apparatus was constructed with a cube and a cylinder, capturing 2,000 synchronized image-force pairs across eight participants.

Evaluation Setting Configuration Contact Precision โ†‘ Contact Recall โ†‘ Contact IoU โ†‘ Force MAE [N] โ†“ Force RMSE [N] โ†“
ARCTIC (2D projection) w/o IRM 0.186 0.023 0.021 - -
ARCTIC (2D projection) Full model (IRM) 0.245 0.304 0.157 - -
Real-World: Cube Full model 0.115 0.316 0.085 0.48 1.54
Real-World: Cylinder Full model 0.114 0.305 0.067 1.35 3.24

Furthermore, comparing SOFA simulations against real-world FTIR calibrated measurements reveals close physical agreement: per-frame force MAE is \(0.25 \pm 0.16\text{ N}\) on the cube and \(0.36 \pm 0.18\text{ N}\) on the cylinder, yielding Spearman rank correlations of 0.604 and 0.760.

Key Findings

  • Iterative pose alignment is decisive for contact reasoning: removing the IRM drops contact IoU drastically (ARCTIC 2D IoU plunges from 0.157 to 0.021), proving that coarse mesh poses break the spatial correspondence required for force estimation.
  • Error-isolation highlights pose as the main bottleneck: when evaluating with ground-truth hand and object poses, predicted force and contact fields align almost seamlessly with ground truth, showing that force regression capacity is sufficient and errors stem primarily from single-view pose ambiguity under occlusion.
  • Train-test hand pose gap: swapping ground-truth hand poses for HAMER predictions at inference induces a trade-off; force MAE improves slightly under domain adaptation, but contact precision drops due to mesh vertex drift.

Highlights & Insights

  • Joint 3D force and contact estimation across hands and objects: expands egocentric interaction perception from 2D pixel-space pressure and flat planar surfaces to full 3D articulated rigid bodies.
  • Scalable physics simulation for physical ground truth: uses SOFA quasi-rigid contact penalties to generate dense per-vertex force supervision, resolving the data scarcity barrier without cumbersome physical sensor rigs during training.
  • Low-cost FTIR physical validation benchmark: establishes an optical FTIR tracking setup converting light reflection into calibrated physical Newtons, providing a sound methodology for sim-to-real tactile validation.

Limitations & Future Work

  • Frame-independent inference: processes video frames independently, neglecting temporal momentum and Newtonian dynamic constraints across sequential frames.
  • Dependency on known CAD mesh templates: cannot operate on arbitrary unseen object geometries zero-shot; pairing with feed-forward 3D shape reconstruction models (such as SAM3D) is an essential next step.
  • Normal forces only and isotropic stiffness assumption: neglects shear and friction forces while assuming uniform material compliance across objects, limiting precision on soft or deformable items.
  • vs HACO: HACO focuses solely on hand-side contact classification across multiple datasets, lacking object-side contact reasoning and force estimation; EgoPHI recovers contact and forces across both hands and objects.
  • vs PressureVision / PressureVision++: PressureVision estimates 2D fingertip pressure heatmaps in image space and struggles with egocentric hand occlusion; EgoPHI leverages 3D mesh topology and GATs for robust physical grounding.
  • vs EgoPressure: EgoPressure predicts hand pressure on 3D meshes but is constrained to flat planar surfaces (such as tabletops); EgoPHI generalizes to complex 3D rigid and articulated object interactions.

Rating

  • Novelty: โญโญโญโญโญ (First vision-based system for dense 3D contact and force estimation on hands and articulated objects, validated with physics simulation and custom FTIR hardware)
  • Experimental Thoroughness: โญโญโญโญโญ (Comprehensive in-distribution and out-of-distribution benchmarks, error isolation with GT poses, and physical sim-to-real validation)
  • Writing Quality: โญโญโญโญโญ (Clear mathematical formulation, cohesive narrative, and transparent analysis of geometric and domain gap challenges)
  • Value: โญโญโญโญโญ (Highly impactful for robotic manipulation learning, physical simulator grounding, and embodied AI)