Skip to content

title: >- [Paper Note] Path-JEPA: Path Signature Based Predictive Learning for Skeleton Action Recognition description: >- [ECCV 2026][Video Understanding][Skeleton Action Recognition] Path-JEPA pioneers path signatures from rough path theory as predictive targets in a masked JEPA framework, constructing continuous-time geometric targets across multi-scale spatial relations and global-local temporal views with signature-augmented masking to achieve state-of-the-art accuracy and remarkable frame-rate robustness. tags: - ECCV 2026 - Video Understanding - Skeleton Action Recognition - Path Signature - Self-Supervised Learning - JEPA date: 2026-09-19 content_hash: b87f8154b9e6bff5

Path-JEPA: Path Signature Based Predictive Learning for Skeleton Action Recognition

Conference: ECCV 2026
Paper: ECCV Official
Full Text Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-4270.txt
Area: Video Understanding
Keywords: Skeleton Action Recognition, Self-Supervised Learning, Path Signature, JEPA, Continuous-Time Geometry

TL;DR

Path-JEPA is the first framework to adopt path signatures from rough path theory as self-supervised prediction targets for skeleton action recognition, utilizing multi-scale spatial paths (joints, edges, chains) and global-local temporal views within a leak-free JEPA architecture to achieve state-of-the-art performance and intrinsic invariance to frame-rate shifts.

Background & Motivation

Human action recognition from skeleton sequences serves as a cornerstone for human-robot interaction, intelligent surveillance, and automated behavior analysis. Skeleton sequences offer a compact, privacy-preserving modality while faithfully maintaining kinematic biomechanics. Within self-supervised representation learning, the field has transitioned from contrastive frameworks toward masked predictive learning. Early autoencoding methods such as SkeletonMAE and MAMP rely on reconstructing raw 3D joint coordinates; however, this coordinate recovery task is acutely sensitive to sensor jitter and high-frequency noise, often steering the network toward memorizing low-level local artifacts. Subsequent joint embedding predictive architectures (JEPAs), notably S-JEPA and GFP, advanced the paradigm by predicting high-level representations in a latent feature space, substantially boosting semantic understanding.

Despite these notable advances, existing masked prediction objectives suffer from two foundational bottlenecks. First, both coordinate and latent prediction targets are conventionally derived from isolated, masked joint tokens. Action semantics, however, are governed not by isolated coordinate displacements, but by the coordinated, coupled relational geometry across kinematically adjacent and distant limbs; predicting targets devoid of explicit relational structure forces the encoder to infer biomechanical constraints only implicitly. Second, physical human motion is inherently continuous in time, whereas contemporary models operate strictly over discrete, frame-indexed observations. This discrepancy introduces rigid frame-rate dependencies, causing severe representation degradation under temporal resampling, dropouts, or sampling frequency variations.

To resolve these joint challenges, a principled mathematical representation that simultaneously captures multi-body relational geometry and guarantees reparameterization invariance over time intervals is essential. Path signatures, originating from rough path theory, characterize continuous trajectories through iterated integrals. By summarizing paths over intervals via coordinate-free geometric invariants (such as linear displacements and Lévy signed areas), signatures provide intrinsic invariance to time reparameterization. When applied across individual joints, relative bone pairs, and kinematic chains, they capture higher-order spatiotemporal coupling. The core idea is to lift discrete skeleton sequences into continuous trajectories to extract multi-scale relational path signatures (over joints, edges, and kinematic chains) as JEPA prediction targets, paired with signature-augmented masking to prevent relational leakage and establish geometry-grounded, temporally robust representations.

Method

Overall Architecture

The Path-JEPA pipeline comprises continuous path lifting with signature tokenization, signature-augmented masking, and a joint embedding predictive architecture (JEPA). Given a discrete 3D skeleton sequence, the framework lifts joint coordinates into continuous paths via piecewise-linear interpolation and lead-lag transformation. Multi-scale log-signatures are then computed across three spatial hierarchies (joints, edges, chains) and dual temporal horizons (global trajectory and local windows). During self-supervised pretraining, an anatomically coherent mask is sampled and propagated across all dependent signature tokens. The context student encoder processes only visible tokens, while a lightweight predictor cross-attends from learnable queries to reconstruct teacher representations at masked indices, supervised by a cosine distance loss against the EMA teacher encoder.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Skeleton Sequence X ∈ R^(T×J×3)"] --> B["Continuous Path Lifting & Signature Extraction<br/>Linear Interpolation + Lead-Lag Lift"]
    B --> C["Multi-Scale Relational Signature Tokenization<br/>Joints / Edges / Chains × Global-Local Views"]
    C --> D["Signature-Augmented Masking<br/>Motion-Aware Seed + Kinematic Subtree Expansion"]
    D -->|Visible Context Tokens| E["Student Context Encoder f_θ"]
    D -->|Full Unmasked Token Sequence| F["EMA Teacher Encoder f_bar_θ"]
    E --> G["Lightweight Predictor g_ϕ<br/>Learnable Query Cross-Attention"]
    F -->|Target Extraction at Masked Indices| H["Centered Teacher Targets Z*_I"]
    G --> I["Cosine Distance JEPA Loss L_JEPA"]
    H --> I

Key Designs

1. Continuous Path Lifting and Log-Signature Representation: Resolving Temporal Discretization

Physical human motion unfolds along continuous smooth manifolds, while discrete frame sampling introduces sensitivity to camera frame rates. To construct a continuous-time representation, each joint trajectory \(X_j \in \mathbb{R}^{T \times 3}\) is lifted via piecewise-linear interpolation into a continuous path \(\tilde{X}_j: [0, \tau] \to \mathbb{R}^3\). To capture instantaneous velocity and motion progression, the trajectory is augmented with a time-lagged copy via a lead-lag transformation: $\(\tilde{X}_j^{\text{LL}}(t) = [\tilde{X}_j(t), \tilde{X}_j(t-\delta)]^\top \in \mathbb{R}^6\)$ The differential between lead and lag streams encodes directional velocity dynamics. Over any temporal interval \([s, t]\), the truncated path signature \(S(\mathbf{x})_{s,t}^{\le n}\) is evaluated. The level-1 terms capture net coordinate displacement \(S^{(1)}_i = x_i(t) - x_i(s)\), while level-2 terms capture iterated double integrals \(S^{(2)}_{ij} = \int_s^t \int_s^{u_1} dx_i dx_j\), where cross-terms \(i \neq j\) represent the Lévy signed area (quantifying planar curvature and directional rotation). To avoid combinatorial dimensionality growth, the model maps the signature to a log-signature using the Baker-Campbell-Hausdorff (BCH) basis. At dimension \(d'=6\) and truncation level \(n=2\), this compresses the raw 43 signature terms into 22 non-redundant coefficients, ensuring computational efficiency and mathematical completeness.

2. Multi-Scale Relational Signature Tokenization: Bridging Kinematic Hierarchy and Dual Horizons

Rather than treating joints in isolation, Path-JEPA establishes a unified token vocabulary spanning spatial kinematic granularities and temporal observation horizons: - Joint Paths (J-PSF): Computed over individual joint paths, capturing localized intrinsic dynamics; - Pairwise/Edge Paths (P-PSF): Formed from relative displacement paths \(Y_{u,v}(t) = \tilde{X}_v(t) - \tilde{X}_u(t)\) between kinematically connected joints \((u, v)\), making the representation inherently invariant to global subject translations; - Chain Paths (C-PSF): Formed by concatenating displacements along multi-joint anatomical limbs (e.g., shoulder \(\to\) elbow \(\to\) wrist), directly embedding coordinated, multi-joint actions into higher-order geometric terms.

Across the temporal dimension, the sequence is divided into multi-scale windows \(\omega_{\ell, k} = (t_{\ell, k-1}, t_{\ell, k}]\). To unify coarse action progression with high-frequency dynamics, each token concatenates a sequence-wide global log-signature with a window-specific local log-signature: $\(m_{p,\ell,k} = \Big[ \log S(\tilde{X}_p^{\text{LL}})_{0, T} \,,\, \log S(\tilde{X}_p^{\text{LL}})_{t_{\ell,k-1}, t_{\ell,k}} \Big] \in \mathbb{R}^{2 D_{\log}}\)$ Tokens are subsequently projected to the transformer hidden dimension via modality-specific linear layers, augmented with type embeddings and a structural spatiotemporal embedding encoding temporal scale \(\ell\) and kinematic root distance \(d_G(p)\).

3. Signature-Augmented Masking: Preventing Cross-Path Information Leakage

In a multi-scale relational token space, naive random token masking causes severe target leakage: because edge and chain signatures are derived from multiple joints, an unmasked edge or chain token can trivially reveal the trajectory of an allegedly masked joint. To preserve the integrity of the self-supervised pretext task, Path-JEPA introduces signature-augmented masking. Seed joints are sampled with probability proportional to their motion magnitude and expanded to encompass the entire anatomical subtree, yielding a kinematically coherent masked joint set \(\mathcal{M}\). The mask is then propagated across the entire signature hierarchy: $\(\mathcal{I}(\mathcal{M}) = \{(p, \ell, k) : p \cap \mathcal{M} \neq \emptyset\}\)$ Every token whose underlying path traverses any masked joint \(j \in \mathcal{M}\)—including the joint token itself, connected edge tokens, and traversing chain tokens—is strictly pruned from the student's context input. This forces the student encoder to infer missing limb dynamics purely from unmasked anatomical context, eliminating short-circuit geometric leakage.

Loss & Training

The framework is instantiated with an asymmetric JEPA design. The student context encoder \(f_\theta\) (6-layer vanilla Transformer, hidden dimension 512, 8 heads) processes only visible context tokens \(\mathcal{C}\). The lightweight predictor \(g_\phi\) (4-layer Transformer) takes learnable query embeddings \(Q_\mathcal{I}\) indexed by masked token metadata and cross-attends to the student context, generating predictions \(\hat{Z}_\mathcal{I}\). The EMA teacher encoder \(\bar{f}_\theta\) processes the full, unmasked sequence, and its features at masked positions are centered via running mean subtraction \(c\) (decay 0.9) to yield targets \(Z^*_\mathcal{I}\). Optimization is driven by a cosine distance loss over the masked indices: $\(\mathcal{L}_{\text{JEPA}} = \frac{1}{|\mathcal{I}|} \sum_{i \in \mathcal{I}} \left( 1 - \frac{\hat{\mathbf{z}}_i^\top \mathbf{z}^*_i}{\|\hat{\mathbf{z}}_i\| \, \|\mathbf{z}^*_i\|} \right)\)$ The teacher weights are updated via an EMA schedule whose momentum warms up from 0.996 to 1.0. Pretraining is executed for 300 epochs on NTU benchmarks. For downstream tasks, the frozen teacher encoder serves as a feature extractor for linear probing, or is fine-tuned end-to-end.

Key Experimental Results

Main Results

Evaluations are conducted on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD under linear evaluation and fine-tuning protocols, using joint-position inputs only (Joint-only).

Table 1: Linear Evaluation Accuracy Comparison (%) (Source: Paper Table 1)

Paradigm / Method Input Modality NTU60 X-Sub NTU60 X-View NTU120 X-Sub NTU120 X-Set PKU-MMD Ph-I PKU-MMD Ph-II
Contrastive Learning
CrosSCLR [25] J+B+M 77.8 83.4 67.9 66.7 84.9 21.2
AimCLR [18] J+B+M 78.9 83.8 68.2 68.8 87.4 39.5
ActCLR [28] J+B+M 84.3 88.8 74.3 75.7
HiCo [13] J 81.1 88.6 72.8 74.1 89.3 49.4
Coordinate Reconstruction
SkeletonMAE [52] J 74.8 77.7 72.5 73.5 82.8 36.1
MAMP [37] J 84.9 89.1 78.6 79.1 92.2 53.8
Feature Prediction
GFP [48] J 85.9 92.0 79.1 80.3 56.2
S-JEPA [1] J 85.3 89.8 79.6 79.9 92.2 53.5
Path-JEPA (Ours) J 87.6 92.9 81.5 81.8 93.4 56.7

Table 2: Supervised vs. Self-Supervised Fine-Tuning Accuracy (%) (Source: Paper Table 2)

Paradigm Method Input Modality NTU60 X-Sub NTU60 X-View NTU120 X-Sub NTU120 X-Set
Supervised ST-GCN [53] J 81.5 88.3 70.7 73.2
2s-AGCN [44] J 88.5 95.1 82.9 84.9
MS-G3D [32] J 91.5 96.2 86.9 88.4
CTR-GCN [9] J 92.6 96.9 88.9 90.8
Self-Supervised SkeletonMAE [52] J 86.6 92.9 76.8 79.1
MotionBERT [58] J 93.0 97.2 84.8 86.4
MAMP [37] J 93.1 97.5 90.0 91.3
MaskCLR [22] J 93.5 97.5 90.5 91.9
GFP [48] J 93.9 97.0 89.6 91.2
S-JEPA [1] J 93.1 97.6 90.3 91.2
Path-JEPA (Ours) J 94.4 98.2 90.8 91.9

Ablation Study

Ablations on frame-rate robustness, signature tokenization components, and masking designs on NTU60 X-Sub isolate the core mechanisms.

Table 3: Robustness to Frame-Rate Variations on NTU60 X-Sub (%) (Source: Paper Table 4)

Method 15fps 20fps 30fps (Default) 45fps 60fps Range (Max - Min)
SkeletonMAE [52] 69.9 73.2 74.8 71.1 70.3 4.9%
MAMP [37] 81.8 83.5 84.9 82.7 81.2 3.7%
S-JEPA [1] 82.4 82.9 85.3 84.7 84.5 2.9%
Path-JEPA (Ours) 87.1 87.5 87.7 87.8 88.1 1.0%

Table 4: Ablation Study on Tokenization Components (Source: Paper Table 6)

Global View Local View Joint Edge Chain Accuracy (%) Note
\(\checkmark\) \(\checkmark\) 84.4 Global joint trajectory baseline
\(\checkmark\) \(\checkmark\) 83.3 Local-only suffers from missing global context (-2.5%)
\(\checkmark\) \(\checkmark\) \(\checkmark\) 85.8 Adding local window dynamic details (+1.4%)
\(\checkmark\) \(\checkmark\) \(\checkmark\) \(\checkmark\) 86.4 Incorporating pairwise relative motion (+0.6%)
\(\checkmark\) \(\checkmark\) \(\checkmark\) \(\checkmark\) \(\checkmark\) 87.6 Full model: multi-joint coordinated chains (+1.2%)

Key Findings

  • Progressive Gains from Relational Hierarchy: Beginning from a global joint-only signature baseline of 84.4%, incorporating local window views (+1.4%), edge relative motions (+0.6%), and multi-joint kinematic chains (+1.2%) drives accuracy to 87.6% on NTU60 X-Sub (+3.2% total gain), validating the benefit of multi-scale geometric abstraction.
  • Data Efficiency in Extreme Low-Data Regimes: Under semi-supervised fine-tuning with only 1% labels (approx. 400 training samples), Path-JEPA attains 72.2% on NTU60 X-Sub, outperforming S-JEPA (67.5%) by 4.7 percentage points. The geometric interpretability of path signatures provides a superior inductive bias for linear separability under sparse supervision.
  • Empirical Validation of Continuous Invariance: Across a wide resampling range from 15fps to 60fps, Path-JEPA exhibits an accuracy variation of merely 1.0% (87.1% to 88.1%), whereas S-JEPA and SkeletonMAE degrade by 2.9% and 4.9% respectively, substantiating the theoretical reparameterization invariance of iterated path integrals.

Highlights & Insights

  • Rough Path Theory Meets Masked Predictive Learning: By pioneering path signatures as JEPA targets, Path-JEPA replaces heuristic coordinate recovery or unconstrained latent prediction with mathematically grounded, continuous-time geometric invariants.
  • Leak-Free Signature-Augmented Masking: Correctly identifying that relational spatial tokens enable geometric shortcutting when partially masked, the framework prunes entire kinematic subtrees and their multi-scale dependencies, enforcing genuine anatomical inference.
  • Surpassing Specialized GCNs with Standard Transformers: Utilizing a standard 6-layer Transformer without manual graph adjacency engineering, the model achieves 94.4% (X-Sub) and 98.2% (X-View) under fine-tuning, demonstrating that relational signature tokenization inherently embeds spatial connectivity.

Limitations & Future Work

  • Heuristic Choice of Signature Hyperparameters: Truncation order (\(n=2\)), lead-lag delay parameter \(\delta\), and window durations are manually preselected rather than learned dynamically to adapt to varying action tempos.
  • Pre-extraction Computation: Calculating log-signatures requires continuous interpolation and tensor contraction during preprocessing; implementing hardware-accelerated signature kernels directly in PyTorch/CUDA could streamline end-to-end efficiency.
  • Single Modality Scope: The current pipeline operates strictly on 3D joint coordinates without joint pretraining across RGB video frames or linguistic action descriptions.
  • vs S-JEPA [1]: S-JEPA predicts latent teacher embeddings derived from isolated joints, remaining vulnerable to frame-rate shifts and lacking explicit relational geometry. Path-JEPA incorporates multi-scale path signatures, achieving a +2.3% gain on NTU60 linear evaluation alongside superior temporal stability.
  • vs SkeletonMAE [52] & MAMP [37]: Coordinate reconstruction models suffer from noise sensitivity and overfit to high-frequency trajectory details. Path-JEPA shifts to interval-level geometric invariants, leading MAMP by +5.3% at 15fps downsampled frame rates.
  • vs Supervised GCNs (CTR-GCN [9], MS-G3D [32]): GCNs rely on explicit graph topology and message passing, whereas Path-JEPA packages pairwise and chain kinematics into signature tokens, allowing standard self-attention to effortlessly capture cross-body coordination.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Groundbreaking integration of rough path signatures with masked JEPA pretraining for continuous, relational motion modeling.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks spanning linear probing, fine-tuning, semi-supervised low-data splits, cross-dataset transfer, and frame-rate sensitivity.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulation, coherent motivation, and precise alignment between theory and empirical results.
  • Value: ⭐⭐⭐⭐⭐ Sets a compelling precedent for applying continuous-time path signature representations to spatio-temporal self-supervised learning.