Skip to content

SOMA: From Surface Observations to Muscle Anatomy

Conference: ECCV 2026
Paper: ECCV Official Paper
Project Page: MPI SOMA
Area: Human Understanding
Keywords: Muscle deformation, Parametric human models, Biomechanical digital humans, Inverse surface observation, Soft-tissue dynamics

TL;DR

Addressing the fundamental limitation that conventional digital human models are restricted to the outer skin surface while finite element simulations are computationally prohibitive, SOMA presents the first data-driven framework to infer dynamic internal muscle deformations directly from multi-view RGB visual observations; using the newly captured SKIM dataset with a custom ArUco marker suit, decoupled non-linear blendshapes, and multi-physical geometric priors, it enables real-time, penetration-free, anatomically grounded musculoskeletal animation driven solely by skeletal pose.

Background & Motivation

Building highly realistic, controllable digital human avatars is a long-standing aspiration in computer vision and computer graphics, playing a vital role in medical rehabilitation, sports science, and immersive virtual reality. Over the past two decades, parametric body models such as SMPL and GHUM have standardized human representation into low-dimensional pose and shape spaces, enabling scalable surface tracking from monocular or multi-view RGB video. However, virtually all popular models terminate strictly at the outer skin geometry, providing zero insight into the underlying biomechanical structures that generate human movement. As digital avatar applications increasingly advance toward sports biomechanics, ergonomic analysis, and surgical simulation, this reliance on skin-only representations reveals severe physiological discrepancies, underscoring the urgent demand for anatomical digital human models that venture beneath the skin.

Existing methodologies for modeling internal anatomy face steep trade-offs between physical fidelity, dynamic expressiveness, and data acquisition feasibility. Heuristic template-fitting techniques rely on oversimplified geometric assumptions that yield visually plausible avatars but fail to reconstruct authentic muscle origin, insertion, and individual morphology. Data-driven medical imaging approaches leverage MRI or CT scans to register volumetric templates, achieving commendable results for static skeletons or localized body parts. However, high acquisition costs, radiation/privacy constraints, and particularly the fundamental physical inability of magnetic resonance to capture dynamic, full-body rapid movements leave in-vivo dynamic musculature virtually devoid of ground truth. On the other hand, physics-based simulations using Finite Element Methods (FEM) or Projective Dynamics are computationally expensive, non-scalable, and heavily dependent on manually tuned physical parameters, while OpenSim-style rigid-body musculoskeletal tools model only 1D line-actuator muscle forces without 3D volumetric shape deformation. Consequently, learning volumetric muscle behavior directly from non-invasive, accessible visual data remains an unresolved inverse challenge.

This paper breaks away from conventional medical imaging constraints and heavy numerical simulations by establishing a new paradigm: inferring internal muscle deformation directly from surface visual signals. The authors recognize that subtle surface skin bulges, stretch patterns, and local shear during movement are the direct physical manifestations of underlying muscle contractions and soft-tissue volume conservation. Core idea: by introducing the SKIM multi-view dynamic dataset captured with a custom ArUco marker suit, SOMA formulates a decoupled muscle-and-skin non-linear blendshapes pipeline supervised by canonical displacement residuals, regularized by multi-physics geometry and volume preservation priors to achieve the first end-to-end recovery of dynamic anatomical muscle deformations driven purely by skeletal poses.

Method

Overall Architecture

SOMA aims to predict anatomically plausible, penetration-free dynamic muscle meshes \(\hat{\mathcal{M}}_{pred}\) and skin surface meshes \(\hat{\mathcal{S}}_{pred}\) for any input skeletal body pose \(\boldsymbol{\theta}\), while driving high-resolution individual muscle geometries. The method first constructs a canonical (rest T-pose) multi-layer anatomical representation comprising skeleton \(B_{bind}\), aggregated muscle \(M_{bind}\), and skin surface \(S_{bind}\) sharing consistent skeletal rigging. Input pose features are processed through decoupled non-linear networks to sequentially predict primary active muscle deformations and superficial passive skin residuals. The framework is trained under canonical residual supervision and regularized by multi-physics priors—incorporating Laplacian smoothness, biharmonic bending resistance, anisotropic tangential sliding, and signed prismatic volume preservation. Finally, canonical displacements are propagated from the unified muscle boundary to dozens of individual anatomical muscle meshes via precomputed barycentric binding.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Skeletal Pose Features r*(θ)"] --> B["Decoupled Non-Linear Blendshapes Prediction<br/>Cascaded U-Nets for Muscle & Skin Residuals"]
    B --> C["Canonical Residual Field Data Supervision<br/>Non-rigid Deviation Alignment via ArUco Markers"]
    C --> D["Biomechanical Multi-Physical Geometric Priors<br/>Laplacian Smoothness/Bending/Sliding/Prismatic Incompressibility"]
    D --> E["Unified Boundary to Individual Muscle Propagation<br/>Barycentric Binding Driving High-Res Muscle Meshes"]
    E --> F["Output: Dynamic Skin & Anatomical Muscle Meshes"]

Key Designs

1. Decoupled Non-Linear Blendshapes Prediction: Decoupling Active Bulging from Passive Sliding Conventional pose space deformation (PSD) frameworks operate on a unified single-layer surface, inherently failing to represent the disparate mechanical behaviors of subcutaneous tissue and deep musculature. SOMA addresses this by decoupling soft-tissue dynamics into two cascaded, interdependent layers. First, the deep muscle mesh captures primary structural volume changes (e.g., active contraction and bulging of the biceps or quadriceps) driven by the skeletal articulation \(\boldsymbol{\theta}\). SOMA utilizes a non-linear U-Net \(\Psi_M\) that takes relative joint rotation features \(r^*(\boldsymbol{\theta}) \in \mathbb{R}^{9(J-1)}\) to directly regress the canonical 3D displacement field: $\(\hat{\mathcal{M}}_{pred}(\boldsymbol{\theta}) = \mathcal{M}_{bind} + \mathbf{D}_{musc}(\boldsymbol{\theta}), \quad \mathbf{D}_{musc} = \Psi_M(r^*(\boldsymbol{\theta}))\)$ Second, the skin surface must not rigidly follow the underlying muscle, but rather slide tangentially and compress normally to accommodate adipose tissue volume conservation. SOMA introduces a second cascaded U-Net \(\Psi_S\) to predict a superficial residual offset \(\mathbf{D}_{res}\), formulating total skin deformation as: $\(\mathbf{D}_{skin}(\boldsymbol{\theta}) = \mathbf{D}_{musc}(\boldsymbol{\theta}) + \mathbf{D}_{res}(\boldsymbol{\theta}), \quad \mathbf{D}_{res} = \Psi_S(r^*(\boldsymbol{\theta}))\)$ This formulation grants the skin the necessary physical degrees of freedom for tangential slippage and inward compression under flexion, completely avoiding the severe "candy-wrapper" joint collapse and unnatural thinning inherent to single-layer skinning.

2. Canonical Residual Field Data Supervision: Eliminating Rigid Articulation and Viewpoint Bias Directly supervising mesh vertices in global world coordinates is susceptible to global translation drift and articulated joint rotation errors, which obscure subtle non-rigid soft-tissue bulges. SOMA adopts a canonical space residual supervision strategy. Using dense 3D surface trajectories \(\tilde{\mathbf{p}}_{k,t}^*\) tracked from the custom ArUco marker suit across multi-view streams, the method computes the purely kinematic marker positions \(\mathbf{p}_k(\boldsymbol{\theta}_t)\) via forward Linear Blend Skinning (LBS) with pre-assigned skinning weights. Next, the observed world-space non-rigid deviations are mapped back into the canonical rest frame using the inverse blended rotation transformation: $\(\boldsymbol{\delta}_{k,t} = \left( \sum_{j=1}^J w_{k,j} \mathbf{R}_j(\boldsymbol{\theta}_t) \right)^{-1} \left( \mathbf{p}_{k,t}^* - \mathbf{p}_k(\boldsymbol{\theta}_t) \right)\)$ During training, the predicted displacement field \(\mathbf{D}_{skin}\) is interpolated onto the marker locations using barycentric coordinates to obtain \(\hat{\boldsymbol{\delta}}_k(\boldsymbol{\theta})\). A masked Mean Squared Error loss \(E_{data}\) is minimized over all visible markers. This mathematical isolation strips away global motion artifacts, compelling the network to focus exclusively on local, pose-dependent soft-tissue bulging and contraction.

3. Biomechanical Multi-Physical Geometric Priors: Higher-Order Elastic Curvature and Prismatic Incompressibility Because inferring two continuous 3D deformation layers from sparse skin observations is fundamentally under-constrained, unregularized optimization yields high-frequency surface wrinkling and severe layer interpenetrations. SOMA introduces an intricate suite of differential geometric and biomechanical energy terms \(E_{reg}\). At the differential geometry level, the formulation enforces first-order Laplacian smoothness \(E_{smooth}\) weighted by Voronoi mass matrices and second-order biharmonic bending resistance \(E_{bi}\). By assigning a much stronger smoothness weight to the skin than the muscle layer (\(\lambda_{smooth}^\mathcal{S} \gg \lambda_{smooth}^\mathcal{M}\)), the model prevents localized high-frequency skin spikes while permitting large, organic muscle bulging. Edge-length preservation \(E_{stretch}\) maintains membrane structural integrity. Furthermore, an anisotropic tangential constraint \(E_{tang}\) penalizes lateral displacement relative to surface normals, with \(\lambda_{tang}^\mathcal{M} \gg \lambda_{tang}^\mathcal{S}\), anchoring muscles to deform predominantly outward while allowing the skin to slide naturally over the fascia. At the volumetric physics level, recognizing that biological soft tissue is nearly incompressible, SOMA discretizes the subcutaneous fat layer (skin-to-muscle) and deep muscle layer (muscle-to-bone) into volumetric prisms \(\mathcal{P}_l\), penalizing deviations from rest volume: $\(E_{vol} = \sum_{l \in \{\mathcal{M}, \mathcal{S}\}} \lambda_{vol}^l \frac{1}{|\mathcal{P}_l^*|} \sum_{p \in \mathcal{P}_l^*} \frac{\big(\text{Vol}(p) - \text{Vol}(p_0)\big)^2}{|\text{Vol}(p_0)| + \epsilon}\)$ Signed volumes are computed via 2-point Gauss-Legendre Quadrature. Crucially, evaluating skin volume in canonical space and deep muscle volume in posed space forces the network to generate active canonical bulges that counteract LBS-induced joint pinching. Moreover, because signed volume becomes negative upon geometric inversion, any bone-muscle or muscle-skin penetration incurs massive energy penalties, intrinsically preventing self-intersection without expensive collision-detection meshes.

4. Unified Boundary to Individual Muscle Propagation: Driving Anatomical Entities via Barycentric Binding For anatomical education, sports biomechanics, and clinical visualization, viewing an aggregated muscle envelope is insufficient; practitioners need isolated inspections of specific muscles. SOMA establishes a topology-aware deformation propagation pipeline. During initial static setup, rays are cast along the negative surface normal from canonical markers and skin vertices into individual muscle meshes \(\{M_{bind}^m\}_{m=1}^{N_m}\), establishing an unambiguous correspondence map \(\phi(p_k) = m\). At runtime, once the non-linear U-Net predicts vertex displacements for the unified global muscle boundary \(M_{bind}\), these displacements are instantly propagated to the vertices of all underlying individual muscle meshes via precomputed 3D barycentric bindings. This decoupled architecture retains the lightweight inference footprint of a single neural forward pass while enabling full-body interactive visualization and localized single-muscle querying in standard graphics engines.

Loss & Training

The model is trained end-to-end by minimizing the composite loss: $\(E_{total} = \lambda_{data} E_{data} + E_{reg}\)$ where \(E_{reg} = E_{smooth} + E_{bi} + E_{stretch} + E_{tang} + E_{vol}\). Optimization is performed using Adam on the newly curated SKIM dataset (5 subjects, 45 minutes of multi-view RGB capture across 120 synchronized cameras at 25 Hz). Visibility masks ensure that only unoccluded, reliably tracked markers contribute to data gradients at each time step.

Key Experimental Results

Main Results

Quantitative evaluations are conducted on unseen validation motion sequences against two representative baselines: kinematic LBS applied to the static subject scan \(S_{bind}\), and state-of-the-art uncoupled implicit volumetric estimator HIT (based on SMPL topology).

Method Global MPME (mm) ↓ Global MedPME (mm) ↓ P90 Error (mm) ↓ Dynamic DRE >25mm (mm) ↓ Interpenetration S→M (%) ↓ Interpenetration M→B (%) ↓ Volume Change ∆Vol S→M (%) ↓ Volume Change ∆Vol M→B (%) ↓
LBS (Kinematic Lower Bound) 13.11 15.64 26.03 134.28 0.85 0.93 N/A 18.81
HIT [27] (CVPR 2024) 115.34 50.15 354.80 151.65 7.61 N/A 2.99 1.95
SOMA (Ours) 13.43 11.51 19.20 64.52 0.85 0.93 1.37 10.35

Note: LBS applies no surface pose-correctives, so subcutaneous fat volume change cannot be computed (N/A); its low interpenetration ratio is an artifact of skin and deep tissue sharing identical skinning weights and collapsing simultaneously. HIT lacks subject-specific proportions, yielding large global tracking errors and a severe 7.61% layer collision rate due to uncoupled boundary estimation.

Ablation Study

The ablation study systematically isolates the impact of network architectures, vector regularization terms, and geometric/physical energy constraints on dynamic testing sequences.

Config / Ablation Global MPME (mm) ↓ Dynamic DRE >25mm (mm) ↓ Interpenetration S→M (%) ↓ Interpenetration M→B (%) ↓ Volume Change ∆Vol S→M (%) ↓ Volume Change ∆Vol M→B (%) ↓ Note
Full Model (U-Net) 13.43 64.52 0.85 0.93 1.37 10.35 Full model with dual-layer U-Net and all physical priors
Linear Blendshapes 16.52 64.58 0.85 0.93 3.57 11.02 Linear formulation struggles with non-linear bulging
MLP Architecture 14.03 64.93 0.86 1.09 5.74 11.05 Lacks spatial weight sharing, degrading volume stability
w/o Physics 13.71 64.93 0.87 0.95 9.37 13.54 Omitting geometric priors causes severe volumetric distortion
w/o Vector (\(E_{smooth}, E_{tang}\)) 13.46 64.57 0.87 0.94 8.27 12.63 Vertices drift laterally, shearing prismatic cells
w/o \(E_{bi}, E_{stretch}\) 13.54 64.48 0.88 0.94 10.54 13.76 Induces severe high-frequency wrinkling to cheat volume
w/o \(E_{vol}\) 13.67 64.57 0.87 0.97 8.28 11.87 Subcutaneous volume error jumps from 1.37% to 8.28%

Key Findings

  1. Halving Error in Highly Dynamic Zones: In soft-tissue regions undergoing intense non-rigid deformation (DRE \(\tau > 25\) mm), SOMA slashes tracking error from 134.28 mm (LBS) down to 64.52 mm—a dramatic 51.9% reduction—effectively eliminating joint volume loss and candy-wrapper collapse.
  2. Priors Resolve Fundamental Under-determination: Ablating membrane elasticity (w/o \(E_{bi}, E_{stretch}\)) causes the highest subcutaneous volume error (10.54%) because the mesh resorts to unnatural high-frequency wrinkling to satisfy volumetric constraints. Similarly, removing tangential constraints (w/o Vector) allows vertices to slide laterally, shearing volumetric prisms and degrading isochoric behavior.
  3. Plausible Biomechanical Layer Trade-offs: In the full formulation, outer subcutaneous fat volume error drops to 1.37%, while deep muscle-to-bone volume retains a 10.35% deviation. This is an anatomically justified compromise where the optimizer prioritizes strict non-penetration against the rigid skeletal core over unbounded volumetric expansion.

Highlights & Insights

  • Inverse Paradigm from Micro-Residuals to Deep Anatomy: Rather than treating skin surface wrinkling and sliding as deformation artifacts to be smoothed away, SOMA capitalizes on them as informative physical cues to invert dynamic muscle behavior, bridging the gap between computer vision surface capture and internal physiology.
  • Signed Prismatic Volume via Gauss-Legendre Quadrature: Modeling subcutaneous space as volumetric prisms and evaluating signed volumes via 2-point Gaussian integration turns geometric layer inversions into massive negative-volume energy penalties, eliminating multi-layer collisions without expensive spatial collision hierarchies.
  • Lightweight Inference Paired with Modular Muscle Inspection: Decoupling neural displacement inference to a unified muscle boundary before mapping to individual muscles via barycentric coordinates yields interactive real-time performance in Viser, opening practical avenues for clinical and sports animation.

Limitations & Future Work

  • Reliance on Initial Anatomical Template Registration: SOMA relies on baseline volumetric fitting (Komaritzan et al.) from static surface scans; any anatomical misalignment in resting muscle shapes propagates into motion. Joint optimization of baseline anatomy and dynamic displacements is a promising avenue.
  • Dependency on Custom Marker Suit for Training: High-precision training requires subjects to wear custom ArUco marker suits in a multi-camera studio; future work should leverage SKIM to train markerless skin correspondence models to democratize data collection.
  • Unmodeled Isometric Contractions and Muscle Activation: The current model is driven purely by joint kinematics \(\boldsymbol{\theta}\), meaning it cannot represent isometric co-contractions or varied muscle activations when joint angles remain static; incorporating surface electromyography (sEMG) or force sensors represents a compelling next step.
  • Subject-Specific Nature: Models are currently trained per subject; expanding across broader demographic distributions to build a multi-subject parametric anatomical foundation model remains an open challenge.
  • vs Surface Parametric Models (SMPL / GHUM): Conventional surface models rely on kinematic skinning and statistical blendshapes, modeling only the outer dermal layer; SOMA introduces a physically grounded, multi-layered anatomical architecture that reconstructs internal muscle mechanics.
  • vs HIT (CVPR 2024): HIT applies implicit neural fields on generic SMPL templates using static MRI data, resulting in over-smoothed dynamics and severe 7.61% layer collisions during movement; SOMA maintains explicit dual meshes with physical volume constraints, cutting collision rates to 0.85% while capturing sharp dynamic bulges.
  • vs FEM Muscle Simulation & OpenSim: Volumetric FEM simulations require minutes per frame and struggle with in-vivo calibration, whereas OpenSim is restricted to 1D force lines; SOMA provides a data-driven neural alternative that runs at interactive frame rates while maintaining physical plausibility grounded in multi-view capture.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering inverse paradigm recovering dynamic 3D muscle deformations from multi-view RGB visual signals]
  • Experimental Thoroughness: ⭐⭐⭐⭐☆ [Comprehensive dynamic metric evaluations and physical ablations on the new SKIM dataset; slightly constrained by subject count]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous mathematical derivations, crystal-clear motivation, and self-consistent physical explanations]
  • Value: ⭐⭐⭐⭐⭐ [Crucial stepping stone connecting superficial digital humans to deep biomechanical avatars across vision, graphics, and medicine]