Skip to content

SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos

Conference: ECCV 2026
Paper: CVF Open Access
Code: https://github.com/eth-siplab/SkillSpotter
Area: Video Understanding
Keywords: Egocentric and Exocentric Video, Skill Assessment, Temporal Action Detection, Pose Fusion, Adaptive Temporal Suppression

TL;DR

Addressing the joint timestamp localization and execution quality grading (good vs. tip) in untrimmed ego-exo videos, SkillSpotter integrates bidirectional cross-view attention, gated 3D body pose fusion, and adaptive temporal suppression, boosting class-specific mAP from 12.40 to 21.82 and achieving 94% of the empirical human inter-annotator agreement ceiling in balanced accuracy.

Background & Motivation

Enabling real-time personalized coaching through Augmented Reality (AR) glasses or fixed camera rigs across domains such as sports, culinary arts, and musical performance requires a system to understand not merely what an actor is doing, but critically how well each action is executed. For years, temporal action detection (TAD) has focused exclusively on localizing temporal segments of action categories in untrimmed videos without assessing execution quality; conversely, action quality assessment (AQA) typically presumes pre-trimmed single-action video clips and directly regresses a quality score, failing to support continuous, moment-by-moment feedback. Ego-Exo4D formalized this challenge in its proficiency demonstration benchmark: detecting timestamps of skilled actions in continuous synchronized egocentric and exocentric (ego-exo) recordings while simultaneously classifying each detected moment as either good execution or a tip for improvement.

However, directly transferring existing state-of-the-art TAD architectures to this task reveals severe limitations. In this paper, seven modern TAD architectures are adapted and systematically re-evaluated using an extended protocol that decouples localization from quality grading. This audit uncovers that current architectures grade execution quality near-randomly (balanced accuracy between 49.51% and 55.99%, close to the 50.9% random baseline). This breakdown stems from three structural tensions: first, skilled action density varies drastically across scenariosโ€”over 80% of events occur within 0.5 s of a neighbor in Basketball, whereas fewer than 20% do in Musicโ€”causing fixed-radius NMS to either erase dense co-occurring events or retain duplicate false positives; second, visual appearance features alone are easily degraded by egocentric camera motion and severe occlusions, struggling to capture subtle biomechanical differences; third, naively concatenating ego and exo visual features induces cross-view representation collapse, degrading grading accuracy.

This work attacks these challenges by coupling skeletal biomechanical priors with scenario-adaptive density and view-specific coordination. Core idea: build SkillSpotter, a pose-aware multi-view framework that coordinates ego and exo streams via bidirectional cross-view attention to eliminate feature concatenation collapse, injects gated 3D body kinematics to complement visual appearance, and learns an adaptive temporal suppression module with continuous exponential decay to dynamically adapt to varying event densities.

Method

Overall Architecture

SkillSpotter takes untrimmed synchronized egocentric and exocentric video streams alongside estimated 3D body pose sequences as input, outputting a set of timestamped skilled action detections with quality classification logits (good execution or tip for improvement). The video inputs are processed using frozen Omnivore representations for ego and exo streams, while 3D skeletal joints and derived kinematic features are reconstructed via Aria camera trajectory estimation or multi-view triangulation. An ActionFormer backbone processes these multi-modal streams across an \(L=8\) level temporal feature pyramid (\(D=512\)), where cross-view attention and gated pose fusion are successively applied at each pyramid level before passing to classification, temporal offset regression, and suppression radius prediction heads.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Synchronized Ego/Exo Video & 3D Pose"] --> B["Omnivore & Pose Conv Temporal Feature Extraction"]
    B --> C["Bidirectional Cross-View Attention<br/>Full-span multi-head attention & adaptive gated fusion"]
    C --> D["Gated 3D Body Pose Fusion<br/>Kinematic feature pyramid & residual gated injection"]
    D --> E["Classification & Temporal Regression Heads<br/>Predict candidate timestamps, confidence & quality logits"]
    E --> F["Adaptive Temporal Suppression<br/>Continuous exponential decay & class-aware suppression"]
    F --> G["Output: Timestamped Skilled Actions & Quality Grades"]

Key Designs

1. Bidirectional Cross-View Attention: Preventing Classification Collapse from Naive Feature Concatenation

When incorporating synchronized ego-exo inputs, naive channel-wise concatenation of ego and exo feature streams severely harms classification discriminability, as the shared temporal backbone struggles to decouple viewpoint-specific geometric distortions and occlusions (e.g., ActionFormer balanced accuracy drops from 55.99% to 50.34%). This design maintains separate temporal feature pyramids \(\{f_{ego,l}\}\) and \(\{f_{exo,l}\}\) and introduces multi-head cross-attention (\(N_h=4\)) operating over the entire temporal span at each level \(l\), allowing egocentric tokens to attend across exocentric context and vice versa: $\(f_{ego,l}' = f_{ego,l} + \sigma(g_e^l) \cdot \text{MHA}(f_{ego,l}, f_{exo,l}), \quad f_{exo,l}' = f_{exo,l} + \sigma(g_x^l) \cdot \text{MHA}(f_{exo,l}, f_{ego,l})\)$ The enhanced streams are merged via a gated projection anchored to the egocentric residual path: \(f_l = f_{ego,l}' + \sigma(g_f^l) \cdot \text{Proj}([f_{ego,l}'; f_{exo,l}'])\), where scalar fusion gate \(g_f^l\) is initialized to -1.0. This preserves the primacy of the first-person perspective while enabling smooth cross-view information exchange that completely prevents representation collapse.

2. Gated 3D Body Pose Fusion: Supplying Skeletal Kinematics to Overcome Visual Appearance Blind Spots

Action execution quality is fundamentally grounded in body joint kinematics, which pure RGB appearance representations fail to capture robustly under variable lighting, background clutter, and viewpoints. This design constructs a 118-dimensional skeletal input vector \(P \in \mathbb{R}^{D_p \times T}\) containing 51 raw 3D keypoint coordinates (17 COCO joints \(\times 3\)) and 67 kinematic features (joint angles, inter-joint pairwise Euclidean distances, and inter-frame velocities). At inference, ego poses are estimated from Aria trajectories while exo poses are estimated via ViTPose and multi-view RANSAC-DLT triangulation, ensuring zero ground-truth leakage. A 1D convolutional backbone encodes the pose sequence into a temporal feature pyramid \(p_l \in \mathbb{R}^{D \times T_l}\) (\(D=512\)), merged at each pyramid level via a gated residual connection: $\(f_l' = f_l + \sigma(g_p^l) \cdot \text{Proj}([f_l; p_l])\)$ The learnable scalar gate \(g_p^l\) is initialized to -2.0, providing stable training warm-up and allowing the model to adaptively balance skeletal kinematics against visual appearance.

3. Adaptive Temporal Suppression: Matching Dynamic Scenario Event Densities via Continuous Exponential Decay

Standard TAD pipelines rely on fixed-radius Soft-NMS, which fails under extreme event-density variations across skilled activities (over 80% of events have a neighbor within 0.5 s in Basketball, compared to fewer than 20% in Music). Furthermore, hard suppression eliminates valid co-occurring good and tip feedback at the same timestamp. This module uses a two-layer MLP conditioned on detection features \(h_k \in \mathbb{R}^D\) and learnable scenario embeddings \(e_s\) to predict continuous per-candidate suppression radii: $\(r_k = \text{Softplus}(\text{MLP}(h_k)) + e_s\)$ For candidate detections sorted by confidence scores, higher-scoring candidate \(i\) suppresses lower-scoring candidate \(j\) via a continuous, class-aware exponential decay: $\(w_{ij} = \exp\left(-\frac{|t_i - t_j|}{r_i}\right) \cdot \mathbb{I}[s_i > s_j] \cdot \mathbb{I}[y_i = y_j]\)$ The class indicator \(\mathbb{I}[y_i = y_j]\) guarantees that detections of different quality classes never suppress each other, allowing co-occurring good/tip events to coexist. The score is updated via \(\tilde{s}_j = s_j \cdot \exp(-\sum_i w_{ij} / \tau)\). By supervising adjusted scores against ground-truth Gaussian temporal targets using binary cross-entropy \(\mathcal{L}_{sup} = \text{BCE}(\text{logit}(\tilde{s}_k), \hat{y}_k)\), the suppression module is trained end-to-end with smooth gradients.

Loss & Training

The overall training objective combines action classification focal loss, cumulative temporal score alignment loss, and the adaptive suppression auxiliary loss: $\(\mathcal{L} = \mathcal{L}_{cls} + \lambda_{reg} \mathcal{L}_{reg} + \lambda_{sup} \mathcal{L}_{sup}\)$ with \(\lambda_{reg} = \lambda_{sup} = 1.0\). The temporal alignment term \(\mathcal{L}_{reg} = \frac{1}{N_{pos}} \sum_{c=1}^{N_c} \sum_t (\text{CDF}_c^{pred}(t) - \text{CDF}_c^{gt}(t))^2\) aligns cumulative distribution functions across time for each class. Models are trained with AdamW (learning rate \(1 \times 10^{-3}\), weight decay 0.05) with 5 warmup epochs and 15 total epochs, taking under 15 minutes on a single NVIDIA H200 GPU. The complete Ego+Exos model has 67.37M parameters and operates at 19.70 FPS.

Key Experimental Results

Main Results

Evaluated on the Ego-Exo4D proficiency demonstration benchmark across three view settings (Ego, Exos, Ego+Exos), metrics include class-specific mAP (\(mAP_S\)), class-agnostic detection mAP (\(mAP_A\)), balanced accuracy (BA), and macro-F1 (F1), averaged over matching radii \(\{0.25, 0.5, 1.0\}\text{s}\):

Model Ego \(mAP_S\) Ego \(mAP_A\) Ego BA Exos \(mAP_S\) Exos BA Ego+Exos \(mAP_S\) Ego+Exos BA
Random Baseline 0.73 1.49 50.90 0.70 50.44 0.70 50.15
Ego-Exo4D Baseline 3.27 โ€“ โ€“ 3.84 โ€“ 3.57 โ€“
VideoMambaSuite 7.63 8.65 49.51 7.06 46.03 3.69 52.70
TadTR 7.79 10.79 49.68 6.37 52.31 4.12 52.81
DyFADet 10.09 12.53 48.89 3.57 49.52 3.18 47.63
TriDet 10.35 14.79 49.17 8.99 50.06 8.23 48.92
CausalTAD 11.42 16.07 52.86 11.82 50.78 13.16 54.98
TemporalMaxer 12.34 16.69 54.34 10.38 52.18 11.27 50.09
ActionFormer 12.40 17.11 55.99 13.18 55.03 13.82 50.34
SkillSpotter (Ours) 21.82 27.89 60.40 21.12 60.59 21.34 60.39
Gain over best baseline +9.42 +10.78 +4.41 +7.94 +5.56 +7.52 +5.41

Ablation Study

Progressive component ablation based on the ActionFormer backbone across viewpoints:

Configuration Ego \(mAP_S\) Ego \(mAP_A\) Ego BA Ego+Exos \(mAP_S\) Ego+Exos BA Note
ActionFormer (Soft NMS) 12.12 16.85 55.34 13.71 50.24 Default baseline
ActionFormer (No NMS) 13.96 22.50 54.81 17.30 51.79 Removing NMS improves detection recall but degrades grading
+ Adaptive Supp. 18.82 25.87 59.74 21.18 43.06 Large single-view gains; concatenation causes multi-view BA collapse
+ Pose Fusion 21.82 27.89 60.40 21.99 45.77 3D kinematics peak single-view performance
+ Cross-View Attn (Ours) โ€“ โ€“ โ€“ 21.34 60.39 Bidirectional cross-view interaction restores multi-view BA (+14.62)

Key Findings

  • Major gains across core modules: Adaptive temporal suppression provides the largest single boost (\(mAP_S\) +4.86, BA +4.93 on Ego). The learned median suppression radii strongly correlate with ground-truth temporal spacing across scenarios (Spearman \(\rho = 0.83, p = 0.010\)), autonomously assigning small radii to Basketball (0.23 grid units) and large radii to Music (0.32 grid units) without density supervision.
  • Cross-view attention eliminates classification collapse: Naive concatenation collapses multi-view BA to 45.77 when coupled with pose and adaptive suppression; bidirectional cross-view attention restores BA to 60.39 (+14.62 gain), matching single-view levels.
  • Scenario performance mirrors motion scale: Substantial improvements occur in full-body motion activities (Basketball \(mAP_S\) 40.76, Rock Climbing 28.87), whereas fine-grained hand-object manipulation tasks (Cooking \(mAP_S\) 3.08, Music 1.57) remain challenging.
  • Reaching 94% of human agreement ceiling: The empirical inter-annotator agreement ceiling between expert annotators on Ego-Exo4D is 64.6 BA (Cohen's \(\kappa = 0.29\)). SkillSpotter's 60.40 BA achieves 94% of this empirical upper bound, indicating that further progress in balanced accuracy is largely constrained by annotation subjectivity.

Highlights & Insights

  • Decoupled evaluation protocol and empirical ceiling: By decomposing localization (\(mAP_A\)) from quality classification (BA), this work reveals that standard TAD architectures grade execution quality near-randomly, while establishing the first empirical inter-annotator ceiling (64.6 BA) for this benchmark.
  • Differentiable adaptive temporal suppression: Rather than applying static post-processing heuristics, the proposed formulation employs continuous exponential decay with class-aware coexistence masks, supervised via Gaussian heatmaps and optimized end-to-end via gradient descent.
  • Universal plug-and-play backbone transferability: The adaptive suppression and pose fusion modules consistently improve performance across other modern TAD architectures (TriDet, CausalTAD, TemporalMaxer) by 8โ€“10 mAP points, and generalize effectively to mistake detection on HoloAssist.

Limitations & Future Work

  • Lack of fine-grained hand-object interaction features: The global 3D body skeleton does not adequately capture subtle finger manipulation in cooking and music; integrating POTTER hand pose yielded a slight performance drop, highlighting the need for specialized hand-object spatial interaction models.
  • Multi-view fusion does not surpass single-view upper bound: While bidirectional attention rescues multi-view training from classification collapse, multi-view performance (\(mAP_S\) 21.34) does not exceed single egocentric performance (21.82), leaving effective complementary geometric fusion an open research problem.
  • Future directions: Integrating egocentric physiological cues (such as heart-rate variability via EgoHRV, skin conductance, or gaze tracking) alongside multi-modal LLMs to extend discrete timestamp classification into descriptive, actionable natural language coaching feedback.
  • vs. Temporal Action Detection (ActionFormer / TriDet / CausalTAD): Standard TAD detects un-graded action intervals and relies on fixed-radius Soft-NMS. SkillSpotter handles timestamp-level joint localization and quality grading using learned adaptive suppression with class coexistence.
  • vs. Action Quality Assessment (AQA): Conventional AQA models require pre-trimmed video clips and regress continuous quality scores. SkillSpotter executes joint localization and grading over continuous untrimmed multi-view video streams.
  • vs. HoloAssist / DR-MoE: Prior mistake detection benchmarks evaluate classification on pre-cut segments. SkillSpotter extends this to untrimmed temporal localization, doubling ActionFormer's detection performance (\(mAP_S\) 3.33 to 7.24).

Rating

  • Novelty: โญโญโญโญ [Well-motivated combination of differentiable temporal suppression, cross-view attention, and 3D pose fusion for joint skill localization and grading]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive benchmark of 7 TAD baselines, ablation of all components, cross-backbone transfer, and human agreement ceiling analysis]
  • Writing Quality: โญโญโญโญโญ [Clear problem formulation, insightful diagnostic analysis of classification collapse, and well-structured technical narrative]
  • Value: โญโญโญโญโญ [Sets a strong new standard and establishes foundational diagnostic metrics for ego-exo skill assessment and AR coaching systems]