Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/yaa1haa1/PartialVisGraph
Area: Video Understanding
Keywords: skeleton-based action recognition, constrained field-of-view, hypergraph neural network, visibility prior, embodied AI
TL;DR¶
To tackle substantial joint visibility dropout caused by restricted camera views in mobile robotics, wearable platforms, and crowded environments, this paper introduces PartialVisGraph—the first hypergraph framework dedicated to skeleton action recognition under constrained Field-of-View (FoV)—combining learnable virtual hyperedges, a soft incidence matrix, and visibility prior gating to achieve up to 68.8% gains over strong baselines under severe view restrictions.
Background & Motivation¶
Skeleton-based human action recognition has achieved prominent success in human-machine collaboration, interactive virtual reality, health monitoring, and surveillance systems. By extracting joint coordinates and their kinematic topologies, skeleton sequences provide a lightweight, geometry-aware representation of human movement that is inherently robust against background clutter, illumination shifts, and camera viewpoint variations. In the emerging era of embodied intelligence and edge robotics, action recognition modules must increasingly be deployed on mobile platforms, wearable headsets, and service robots. However, virtually all prevailing models operate under the idealized premise that clean, complete skeleton sequences are consistently available throughout the video. In practical embodied settings, physical sensor placement, dynamic agent movement, and constrained lens angles frequently leave large portions of the human body outside the camera frame.
The fundamental limitation of existing Graph Convolutional Networks (GCNs) lies in their reliance on standard graphs with pairwise physical edges. Information propagation is strictly constrained along anatomical connections; when a central torso or supporting limb moves out of the field-of-view, the underlying physical graph fragments, disrupting feature propagation and collapsing the topological representation. Furthermore, complex human actions are naturally defined by high-order, coordinated movements across multiple distant body parts—for example, kicking requires simultaneous counter-balancing arm elevation, and standing up demands synchronized dynamics across legs, pelvis, and spine. Under constrained FoV conditions where direct visual evidence is sparse and fragmented across non-adjacent joints, conventional binary graphs cannot aggregate these weak, spatially distributed cues into coherent semantic representations.
To resolve this limitation, this work establishes the problem of skeleton-based action recognition under constrained field-of-view. The authors argue that overcoming truncated observations requires breaking the rigid boundary of pairwise connections through data-driven hypergraphs, coupled with an explicit visibility gate to prevent missing joint noise from corrupting reliable joint representations. Core idea: adaptively construct dynamic soft hypergraphs via a bank of learnable virtual hyperedge tokens, aggregate joint features with a Single-Head Sample-Adaptive Transformer that incorporates logarithmic visibility priors, and regularize the structure via hyperedge orthogonality and class clustering to achieve robust high-order action reasoning under truncated fields of view.
Method¶
Overall Architecture¶
PartialVisGraph processes partially observed skeleton sequences by combining an anatomical graph inductive bias with flexible higher-order hypergraph aggregation. The input skeleton sequence is first projected into a high-dimensional feature space and augmented with learnable spatial-temporal positional embeddings. The resulting representation is processed through a sequence of Hybrid Graph-Hypergraph (HGH) blocks interleaved with Multi-Scale Temporal Convolutional Networks (MS-TCN) to capture spatial and temporal dynamics, concluding with temporal average pooling and a linear classifier.
Within each HGH block, input features are partitioned along the channel dimension into a GCN branch (retaining physical topology inductive bias) and a hypergraph branch. The hypergraph branch employs the Incidence-Aware Hypergraph Attention (IAHA) module: temporally pooled joint features interact with a bank of learnable virtual hyperedges to dynamically compute a soft incidence matrix; a Single-Head Sample-Adaptive Transformer (SHSAT) then aggregates visible joint information onto hyperedge tokens under logarithmic visibility masking, before redistributing high-order features back to individual joints.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Constrained FoV skeleton sequence<br/>X ∈ R^(T×V×Cr) & visibility map vis"] --> B["Linear projection & positional embedding<br/>Feature tensor Xe ∈ R^(T×V×Cf)"]
B --> C["Channel partition<br/>Xgcn (Cf/4) & Xhyper (3Cf/4)"]
C --> D["Dual-branch spatial modeling<br/>Physical GCN branch ∥ IAHA hypergraph branch"]
D --> E["1. Soft incidence matrix construction<br/>Virtual hyperedge bank paired with joint summaries via sparsemax"]
E --> F["2. Single-Head Sample-Adaptive Transformer (SHSAT)<br/>Joint-to-hyperedge aggregation with log-visibility prior Bvis"]
F --> G["Hyperedge feature redistribution<br/>Fj = Hsoft × Fe projected to joints & concatenated with GCN"]
G --> H["Multi-scale temporal modeling & classification<br/>MS-TCN layers → temporal pooling → linear classifier ŷ"]
Key Designs¶
1. Soft incidence matrix construction: adaptive hyperedges beyond pairwise physical graphs
Prior hypergraph formulations in action recognition impose rigid structural constraints—such as enforcing disjoint hyperedges where each joint belongs to exactly one hyperedge, or restricting the number of hyperedges to equal the number of joints. To support flexible body grouping under arbitrary view restrictions, IAHA introduces a learnable virtual hyperedge bank \(E_v \in \mathbb{R}^{K_v \times C_v}\) (configured with \(K_v = 10\)). Given joint features \(F_h \in \mathbb{R}^{T \times V \times C_v}\), temporal average pooling first summarizes global motion into a sequence-level joint descriptor \(F_s = \text{Pool}_t(F_h) \in \mathbb{R}^{V \times C_v}\).
The model then calculates the cosine similarity matrix \(S \in \mathbb{R}^{V \times K_v}\) between joint summaries and virtual hyperedge tokens, followed by row-wise sparsemax normalization to yield the continuous soft incidence matrix \(H_{\text{soft}}\):
The sparse activation eliminates spurious associations and naturally permits inactive (all-zero) hyperedge columns when certain body synergies are absent. Simultaneously, non-zero entries capture soft, multi-membership affinities, allowing a joint to contribute dynamically across multiple semantic action groups without requiring rigid manual hypergraph definitions.
2. Single-Head Sample-Adaptive Transformer (SHSAT): visibility prior gating
Aggregating joint features onto hyperedges under partial visibility presents a critical vulnerability: out-of-view or occluded joints carry zeroed or distorted coordinates that corrupt hyperedge representations if processed naively. To resolve this, virtual tokens are replicated across time to form \(E_{vt} \in \mathbb{R}^{T \times K_v \times C_v}\) and concatenated with joint features \(F_h\). An attention mask \(M \in \{0, 1\}^{(K_v+V) \times (K_v+V)}\) derived from \(H_{\text{soft}}\) enforces that joints only transmit signals to their assigned hyperedges.
Critically, the architecture incorporates the sequence visibility map \(vis \in [0, 1]^{T \times V}\) (with 1 indicating visibility and 0 indicating occlusion/out-of-frame). A logarithmic visibility bias \(B_{\text{vis}} = \log(vis + \varepsilon)\) is broadcast into the attention logits:
Because out-of-view joints have near-zero visibility, their attention logits are heavily penalized by large negative values, driving their softmax weights to zero and completely preventing unobserved joints from corrupting the hyperedge representations. After updating the \(K_v\) hyperedge tokens \(F_e\), high-order features are redistributed back to joints via matrix multiplication \(F^{(t)}_j = H_{\text{soft}} F^{(t)}_e\), creating a resilient cycle of aggregation and holistic feature completion.
3. Multi-objective hypergraph regularization: preventing assignment collapse
End-to-end data-driven hypergraph discovery is susceptible to structural collapse, such as all joints congregating onto a single hyperedge or hyperedge tokens collapsing to identical vectors. To ensure stable and diverse topology discovery, the authors introduce three targeted auxiliary objectives alongside standard classification loss \(L_{ce}\): - Hyperedge diversity loss (\(L_{\text{pool}}\)): Evaluates the Gram matrix \(G = \tilde{P}\tilde{P}^\top\) of \(L_2\)-normalized virtual hyperedge tokens to penalize off-diagonal squared entries, driving virtual hyperedges toward mutual orthogonality: $\(L_{\text{pool}} = \frac{\|G - I\|_F^2}{K_v(K_v - 1)}\)$ - Assignment regularization loss (\(L_{\text{assign}}\)): Measures column sums of \(H_{\text{soft}}\) to produce hyperedge utilization distribution \(p\). It combines a batch-level balance loss \(L_{\text{balance}}\) (penalizing divergence from uniform distribution) and a per-sample hinge loss \(L_{\text{hinge}}\) (capping maximum assignment probability at threshold \(\tau\)), formulated as \(L_{\text{assign}} = w_{\text{balance}}L_{\text{balance}} + w_{\text{hinge}}L_{\text{hinge}}\) to eliminate degenerate winner-take-all hyperedges. - Class-centre clustering loss (\(L_{\text{cluster}}\)): Maps flattened incidence matrices into low-dimensional embeddings \(z_n\) and minimizes their distance to learnable class centers \(c_k\) via \(L_{\text{pull}}\), while repelling distinct class centers via a margin-based hinge loss \(L_{\text{repel}}\) to ensure action-consistent topological configurations.
Loss & Training¶
The overall training objective combines classification and structural regularizers:
Training is guided by two tailored data protocols: - Curriculum learning: In the initial epoch, the visible joint ratio is maintained high (95%–100%). Over 70 epochs, the FoV window progressively contracts down to a challenging 20%–80% range, stabilizing initial feature convergence before exposing the network to extreme partial visibility. - Temporal CutMix: Applied with 0.3 probability, segments from two skeleton sequences \(a\) and \(b\) are concatenated across the temporal axis, with label losses weighted by segment duration \(L_{\text{mix}} = \frac{t}{T}L_a + \frac{T-t}{T}L_b\), compelling the model to generalize across interrupted temporal dynamics.
Key Experimental Results¶
Main Results¶
Evaluation is conducted on NTU RGB+D 60 and NTU RGB+D 120 under simulated constrained FoV benchmarks centered on the Head (joint 3), Spine (joint 20), and Base of Spine (joint 0), across Easy (~75% visible), Medium (~50% visible), and Hard (~25% visible) splits using a standard 4-stream ensemble.
The table below reports action recognition accuracy on NTU RGB+D 60 across various FoV centers and difficulty levels:
| FoV Center | Protocol | Difficulty | PartialVisGraph (Ours) | Hyper-GCN (ICCV 2025) | Hyperformer (2022) | FR-Head (CVPR 2023) | Gain |
|---|---|---|---|---|---|---|---|
| Head (3) | X-Sub | Easy (%) | 90.9 | 85.4 | 84.5 | 83.2 | +5.5% |
| X-Sub | Medium (%) | 88.5 | 65.9 | 64.3 | 57.3 | +22.6% | |
| X-Sub | Hard (%) | 79.2 | 31.6 | 29.4 | 26.5 | +47.6% | |
| X-View | Hard (%) | 84.5 | 34.4 | 32.9 | 25.7 | +50.1% | |
| Spine (20) | X-Sub | Easy (%) | 90.6 | 82.1 | 82.1 | 80.6 | +8.5% |
| X-Sub | Medium (%) | 87.1 | 58.5 | 56.7 | 51.3 | +28.6% | |
| X-Sub | Hard (%) | 77.7 | 24.1 | 23.7 | 21.7 | +53.6% | |
| X-View | Hard (%) | 83.2 | 27.8 | 27.2 | 20.5 | +55.4% | |
| Base of Spine (0) | X-Sub | Easy (%) | 89.2 | 65.1 | 68.8 | 61.9 | +20.4% |
| X-Sub | Medium (%) | 83.8 | 34.1 | 40.0 | 33.8 | +43.8% | |
| X-Sub | Hard (%) | 71.4 | 7.3 | 10.8 | 9.0 | +60.6% | |
| X-View | Hard (%) | 76.5 | 7.7 | 9.5 | 7.2 | +67.0% |
On NTU RGB+D 120, under the most severe restriction (Base of Spine, Hard split), baselines suffer catastrophic breakdown (1.6%–4.7% accuracy), whereas PartialVisGraph retains 66.7% (X-Sub) and 67.4% (X-Set), producing remarkable advantages exceeding +62.0% (with max relative gains of 68.8%).
Furthermore, on standard Full FoV benchmarks, PartialVisGraph also sets a new state-of-the-art:
| Method | Publication | Params (M) | GFLOPs | NTU 60 X-Sub (%) | NTU 60 X-View (%) | NTU 120 X-Sub (%) | NTU 120 X-Set (%) |
|---|---|---|---|---|---|---|---|
| CTR-GCN | ICCV 2021 | 1.5 | 1.97 | 92.4 | 96.4 | 88.9 | 90.6 |
| FR-Head | CVPR 2023 | 2.0 | – | 92.8 | 96.8 | 89.5 | 90.9 |
| HD-GCN | ICCV 2023 | 1.7 | 1.77 | 93.0 | 97.0 | 89.8 | 91.2 |
| SkateFormer | ECCV 2024 | 2.0 | 3.62 | 93.5 | 97.8 | 89.8 | 91.4 |
| ProtoGCN | CVPR 2025 | – | – | 93.5 | 97.5 | 90.4 | 91.9 |
| Hyper-GCN | ICCV 2025 | 2.3 | 2.88 | 93.7 | 97.8 | 90.9 | 92.0 |
| PartialVisGraph (Ours) | ECCV 2026 | 4.0 | 2.90 | 93.8 | 97.6 | 91.0 | 92.3 |
Ablation Study¶
All ablation models are evaluated on the single joint modality. The impact of individual modules across Full FoV and Constrained FoV (averaged across the three difficulty splits) is summarized below:
| Configuration | Full FoV Accuracy (%) | Constrained FoV Avg Accuracy (%) | Note |
|---|---|---|---|
| PartialVisGraph (Full Model) | 92.0 | 78.8 | Complete proposed framework |
| w/o \(L_{\text{pool}}\) | 91.8 (-0.2) | – | Removing orthogonality causes redundant hyperedges |
| w/o \(L_{\text{assign}}\) & \(L_{\text{cluster}}\) | 91.2 (-0.8) | – | Removing regularization collapses hypergraph into a single edge |
| w/o Temporal CutMix | 91.0 (-1.0) | – | Removing temporal segment mixing reduces generalization |
| w/o \(B_{\text{vis}}\) (Visibility Prior) | – | 78.4 (-0.4) | Removing log-visibility allows unobserved noise to leak |
| w/o Curriculum Learning | – | 78.4 (-0.4) | Training directly on mixed difficulties impairs early stability |
Key Findings¶
- Conventional architectures suffer catastrophic failure under constrained FoV: When visibility drops to ~25% in the Hard split, GCN and hypergraph baselines drop below 10%–35% accuracy. Without an explicit visibility gate, missing joint zero-fillings are treated as valid structural nodes, corrupting feature message passing throughout the network.
- Regularization is essential to prevent hypergraph collapse: Visualizations demonstrate that omitting \(L_{\text{assign}}\) results in a collapsed soft incidence matrix where one single hyperedge encompasses all joints. The proposed regularizer maintains clean semantic decomposition (e.g., separating leg-pelvis locomotion from arm-torso balancing).
- Visibility prior enforces clean message propagation: Injecting the logarithmic visibility prior \(B_{\text{vis}}\) into attention logits ensures a clean 0.4% performance boost by mathematically nullifying out-of-view joints before feature aggregation occurs.
Highlights & Insights¶
- Formulates a realistic, high-impact embodied vision benchmark: Bridges the gap between sanitized academic skeleton datasets and messy, view-constrained real-world deployments in robotics and wearable computing.
- Elegant closed-loop hypergraph feature redistribution: Pairs a virtual hyperedge bank with sparsemax to induce dynamic soft incidence matrices, routing features through attention and redistributing them via simple matrix multiplication without combinatorial search.
- Logarithmic visibility gating: The simple injection of \(\log(vis + \varepsilon)\) into transformer attention logits provides a mathematically sound, plug-and-play mechanism to suppress missing observations across modalities.
Limitations & Future Work¶
- Reliance on external visibility annotations: The formulation assumes binary visibility indicators \(vis\) are supplied upstream; inaccurate confidence scores from 3D pose estimators in cluttered scenes may degrade gating efficacy.
- Parameter overhead: Introducing Transformer attention layers increases the parameter count to 4.0M (compared to 1.5M–2.0M for compact GCNs), requiring pruning or distillation for ultra-low-power microcontrollers.
- Generative completion: Future extensions could incorporate generative priors (e.g., diffusion trajectory models) to explicitly hallucinate missing limb dynamics alongside discriminative classification.
Related Work & Insights¶
- vs Hyperformer [Zhou et al., 2022] & Hyper-GCN [Zhou et al., ICCV 2025]: Prior hypergraph architectures enforce restrictive assumptions (disjoint node-to-hyperedge sets or fixed hyperedge counts matching joint counts) and evaluate solely on complete skeletons. PartialVisGraph enables unconstrained multi-membership hypergraphs and achieves stable recognition under extreme visibility dropout (~80% vs ~30% for baselines).
- vs Occluded Skeleton Action Recognition [Chen et al., ACM MM 2023]: Previous occlusion-handling techniques focus on random joint dropout (simulating sensor noise), whereas constrained field-of-view creates contiguous, regional spatial cutoffs (e.g., lower body fully off-screen). The proposed FoV benchmark provides a standardized foundation for evaluating regional truncation.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneers skeleton action recognition under constrained FoV with a learnable virtual hypergraph and visibility-gated Transformer]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Comprehensive evaluations across multiple FoV centers, three difficulty levels, and full-FoV settings on NTU 60 and 120]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous loss formulations, and compelling ablation analyses]
- Value: ⭐⭐⭐⭐⭐ [Directly addresses the operational requirements of embodied AI, service robotics, and wearable devices in real-world environments]