PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition¶
Conference: NeurIPS2026 (task-list assignment; this note uses arXiv v2)
arXiv: 2605.11497
Area: Human Understanding
Keywords: zero-shot action recognition, skeletonization gap, pose-anchored semantics, cross-attention, semantic prototype adaptation
TL;DR¶
Instead of adding an RGB action-recognition branch, PoseBridge preserves visual semantics inside pose estimation before they are compressed into joint coordinates, then transfers them through skeleton-conditioned bridging and semantic prototype adaptation, improving over the strongest compared baseline by 13.3–17.4 percentage points across eight Kinetics-200/400 splits.
Background & Motivation¶
Zero-shot skeleton-based action recognition usually converts video into joint trajectories before aligning skeleton representations with action text. This reduces sensitivity to backgrounds, clothing, and illumination and supports transfer of motion patterns to unseen classes; methods such as PURLS and TDSM consequently focus on skeleton–language alignment. However, joint coordinates are not complete action evidence: putting on a hat, headphones, or glasses can all involve hands moving toward the head, while the manipulated object and its relations to the hands and head are not explicitly retained in the coordinate sequence.
The paper calls this the skeletonization gap: text prototypes require richer semantics than skeleton observations preserve, and the information loss occurs before alignment. Further optimizing skeleton–text distances does not guarantee recovery of discarded visual evidence. Language-generated object or scene descriptions provide class priors but cannot directly establish what appears in this particular video. An additional RGB backbone supplies evidence but changes the input protocol and adds computational cost and opportunities for appearance shortcuts.
The pose estimator has already observed the video. Shallow features retain local visual details and deep features represent body configuration, but naive fusion can still absorb background information rather than action semantics. Core idea: train intermediate features from the same human pose estimation pass into body-organized semantic evidence, preserve information upstream of skeletonization, and correct both recognition queries and language prototypes rather than improving alignment only after coordinate extraction.
Method¶
Overall Architecture¶
The input is a video and the output is an unseen action label; generalized zero-shot recognition also allows seen labels. PoseBridge has two training stages: train the pose estimator and pose-anchored semantic extractor on MS COCO, then freeze them to extract 2D skeletons and frame-level semantic cues from action videos. Subsequent recognition training uses only seen-class action samples to learn bridging and alignment.
The four key designs follow the data flow: hierarchical pose refinement, body-aware pooling, skeleton-conditioned bridging, and semantic prototype adaptation. The first two turn HPE intermediate features into cacheable semantic cues; the latter two enrich the test query and correct class prototypes, respectively. Prototype adaptation uses a separate temporal-pooling path over the cues, not the test query to update class centroids dynamically.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["COCO pose and<br/>image–text supervision"] -.-> B["Hierarchical pose refinement"]
V["Video through HPE<br/>frozen after training"] --> B
B --> C["Body-aware pooling"]
A -.-> C
V --> S["2D skeleton encoding"]
C --> P["Frame-level semantic cues"]
S --> D["Skeleton-conditioned bridging"]
P --> D
D --> Q["Recognition query"]
P --> E["Semantic prototype adaptation"]
T["Action text prototypes"] --> E
Y["Seen-class training<br/>samples only"] -.-> E
E --> R["Seen centroids and<br/>unseen neighbor-transfer prototypes"]
Q --> O["ZSL / GZSL matching"]
R --> O
Dashed edges indicate training supervision or the source of seen-class statistics; solid edges indicate extraction and matching data flow. Unseen classes supply text only and cannot use unseen action videos to construct centroids. “No additional RGB branch” does not mean “no RGB access”: video still enters the pose estimator, and recognition retains its intermediate visual representations.
Key Designs¶
1. Hierarchical pose refinement: transfer shallow details into deep body structure
Using only the deepest HPE features may preserve body pose while discarding fine-grained evidence such as objects near the hands. PoseBridge selects three shallow-to-deep feature levels by default. It projects the current shallow feature into the next level's channel space, resizes it to the next resolution, applies convolutional spatial refinement, and injects it residually with coefficient 0.5. The updated feature then continues to the deeper level.
This is not simple concatenation of every feature level. Progressive injection brings details into deep representations already organized by the keypoint task, producing a spatial feature map that retains pose structure. The residual path preserves the deep backbone rather than replacing body structure with shallow appearance, although the paper does not establish that all background information is eliminated.
2. Body-aware pooling: organize visual semantics around the predicted body
The deep map covers the whole crop, so ordinary global average pooling mixes background responses with evidence near the body. The model converts predicted joint probabilities into heatmaps, averages them across joints, and normalizes the result into a body attention map. This map weights a spatial average over the refined feature map, which is projected into one frame-level semantic vector; a small denominator constant provides numerical stability.
The prior favors responses near the body and action-relevant body parts. Joint coordinates cannot express which object is beside the hand, but HPE features at that location can still contain object and interaction evidence. Pooling therefore operates on visual features, not reweighted joint coordinates. Its body-centered support also means that distant objects or very small interaction targets may remain unobserved.
To prevent the pooled vector from merely describing clothing or appearance, HPE training includes image–text semantic supervision. COCO image-level captions are encoded with a frozen CLIP text encoder and aligned with pose-anchored vectors through symmetric contrastive learning. Captions provide weak image-level supervision, not per-person action labels or annotated hand–object relations. Joint optimization with the original pose loss constrains both localization and semantic preservation.
3. Skeleton-conditioned bridging: let motion queries select useful temporal semantics
After HPE is frozen, each action video supplies a skeleton sequence and a semantic cue sequence. Shift-GCN encodes the skeleton, and its projected representation queries the frame-level pose semantics as keys and values through multi-head cross-attention. The default uses 16 temporal cues per video and four attention heads, rather than sending each RGB frame through another action backbone.
A learnable element-wise sigmoid gate controls the attention output before residual addition to the skeleton representation. Layer normalization, a feed-forward network, and a second residual normalization produce the final bridge representation. Motion remains the starting point of the query, while temporal visual evidence is selected conditionally. Compared with direct concatenation, this lets the skeleton trajectory select cues that resolve its semantic ambiguity.
The same cue sequence also undergoes attention-based temporal pooling to produce a video-level pose-semantic representation. This representation supports seen-class centroid statistics and the pose-semantic training anchor; the skeleton representation is the motion anchor, while the bridge representation is the final zero-shot matching query. These roles should not be conflated with a single closed-set classification output.
4. Semantic prototype adaptation: correct unseen text with seen-class displacements
Enriching the query does not automatically remove the language–visual mismatch on the target side. The model averages video-level pose-semantic representations from training samples of each seen class, then moves that class's text prototype toward its centroid with default coefficient 0.2. Keeping text dominant prevents prototypes from becoming visual templates usable only for seen classes.
An unseen class has no visual centroid, so averaging its videos is not allowed. The model selects its five nearest seen classes in text space, assigns temperature-softmax weights from text similarity, and transfers each neighbor's “pose centroid minus text prototype” residual instead of replacing the unseen prototype with a neighboring centroid. The essential relation is:
Here, \(\boldsymbol\mu_j\) uses only seen-class training samples, and \(\omega_{cj}\) is a normalized similarity weight over the selected text neighbors. Centroids and adapted prototypes are L2-normalized. Transfer assumes that semantically related classes have transferable text–pose biases; correction may fail when neighboring classes have different object-interaction structures.
A Worked Example¶
Consider a video showing hands moving toward the head. Its skeleton can be compatible with both putting on a hat and putting on glasses, making the trajectory alone ambiguous. Three HPE feature levels are progressively refined, and body-aware pooling preserves visual responses near the head and hands as temporal cues. This is an explanatory example of the method, not an additional measured case.
At recognition time, the skeleton query aggregates evidence through cross-attention over the default 16 cues. If the target action is unseen, its text prototype has already been corrected using displacement transfer from five seen text neighbors. The bridge query is then compared with the adapted prototypes. No training videos of that unseen action or separate hat/glasses detector are required, but preservation of object details still depends on HPE features and the crop boundaries.
Loss & Training¶
The HPE stage uses RTMPose-s with \(256\times192\) inputs and 420 training epochs on COCO. Its objective combines the original SimCC pose loss with a symmetric image–text contrastive loss weighted by 0.1, using contrastive temperature 0.07. Semantic supervision uses the frozen CLIP ViT-B/32 text encoder and a 512-dimensional semantic space. The top-down pipeline uses YOLOv11 for human detection, so “no external object detector” must not be broadened into “no human-detection step.”
The ZSSAR stage freezes HPE and the CLIP ViT-L/14 text encoder and caches encoded action-level descriptions. Skeleton features have dimension 256, the shared embedding has dimension 512, and dropout is 0.1. Action descriptions follow SMIE-style ChatGPT expansion without additional body-part or temporal-phase prompts. The primary change is therefore visual evidence and semantic transfer rather than stronger description engineering.
Both the skeleton anchor and pose-semantic anchor use seen-class classification cross-entropy, cross-entropy for matching seen text prototypes, and supervised contrastive loss. The bridge representation uses semantic matching and supervised contrastive loss only, without a seen-class closed-set classifier, to avoid constraining the final query to seen-class decision boundaries.
Cosine consistency additionally aligns skeleton and pose-semantic representations. The pose branch's probability distribution over seen text prototypes is distilled into the skeleton branch, with stop-gradient on the teacher and distillation temperature 4.0. This teaches semantic relations to the skeleton representation but does not establish that HPE cues can be removed at test time without losing the reported performance.
Recognition training uses AdamW for 30 epochs, batch size 128, learning rate \(10^{-3}\), and weight decay \(2\times10^{-3}\). Five warm-up epochs precede cosine annealing, with gradient clipping and EMA. At test time, normalized bridge representations and adapted prototypes are scored by dot product. ZSL considers unseen classes only; GZSL considers both sets and subtracts calibration coefficient \(\kappa\) from seen-class scores to reduce seen bias. ZSL and GZSL accuracies are not interchangeable.
Key Experimental Results¶
Main Results¶
All results below are percentages, and splits mean “seen/unseen.” NTU uses Xsub; random-split results average three class partitions. ZSL measures unseen-class top-1 accuracy. GZSL-H is the harmonic mean of seen accuracy \(S\) and unseen accuracy \(U\):
| Dataset and protocol | Metric | PoseBridge | Comparison | Comparison result | Gain (percentage points) |
|---|---|---|---|---|---|
| NTU60 standard 55/5 | ZSL | 88.8 | Neuron / FS-VAE | 86.9 | 1.9 |
| NTU60 standard 48/12 | ZSL | 73.2 | Neuron | 62.7 | 10.5 |
| NTU120 standard 110/10 | ZSL | 80.3 | BSZSL (additional RGB) | 77.7 | 2.6 |
| NTU120 standard 96/24 | ZSL | 70.9 | TDSM | 65.1 | 5.8 |
| NTU60 standard 48/12 | GZSL-H | 68.0 | Neuron | 59.1 | 8.9 |
| NTU120 random 110/10 | GZSL-H | 66.0 | SCoPLe | 54.1 | 11.9 |
| PKU-MMD random 46/5 | GZSL-H | 71.9 | FS-VAE | 59.0 | 12.9 |
| Kinetics-200 180/20 | ZSL | 55.6 | TDSM | 38.2 | 17.4 |
| Kinetics-200 160/40 | ZSL | 38.3 | TDSM | 24.4 | 13.9 |
| Kinetics-200 140/60 | ZSL | 28.6 | TDSM | 15.3 | 13.3 |
| Kinetics-200 120/80 | ZSL | 26.4 | TDSM | 13.1 | 13.3 |
| Kinetics-400 360/40 | ZSL | 52.3 | TDSM | 38.9 | 13.4 |
| Kinetics-400 320/80 | ZSL | 39.6 | TDSM | 26.2 | 13.4 |
| Kinetics-400 300/100 | ZSL | 32.9 | TDSM | 18.5 | 14.4 |
| Kinetics-400 280/120 | ZSL | 32.1 | TDSM | 16.1 | 16.0 |
Sources: Tables 1–3. PoseBridge's standard NTU120 110/10 GZSL-H is 68.0, below the additional-RGB method BSZSL at 75.8. It is therefore not best across every modality and metric.
Most previous methods in the main table use Kinect 25-joint 3D skeleton protocols, whereas PoseBridge uses RTMPose COCO-17 2D skeletons. Appendix Table 8 re-evaluates Neuron, TDSM, and FS-VAE in the same 2D skeleton format: their NTU60 48/12 ZSL accuracies, followed by PoseBridge, are 58.8, 64.7, 48.8, and 73.2. This controls skeleton-format differences but does not give the baselines identical HPE semantic cues, so it is not an equal-information comparison under a strictly coordinate-only protocol.
Ablation Study¶
Table 5 fixes NTU60 Xsub 48/12. HR, BP, SB, and PA denote hierarchical pose refinement, body-aware pooling, skeleton-conditioned bridging, and semantic prototype adaptation.
| HR | BP | SB | PA | ZSL | GZSL-H |
|---|---|---|---|---|---|
| No | No | No | No | 54.7 | 52.0 |
| No | No | No | Yes | 56.6 | 57.3 |
| No | No | Yes | No | 60.4 | 60.7 |
| No | No | Yes | Yes | 61.4 | 61.7 |
| No | Yes | Yes | Yes | 64.0 | 63.2 |
| Yes | No | Yes | Yes | 64.8 | 63.8 |
| Yes | Yes | No | Yes | 66.9 | 63.9 |
| Yes | Yes | Yes | No | 71.0 | 66.5 |
| Yes | Yes | Yes | Yes | 73.2 | 68.0 |
Key Findings¶
- The full model improves over the skeleton baseline by 18.5 ZSL points and 16.0 GZSL-H points. Removing SB, HR, BP, or PA reduces ZSL to 66.9, 64.0, 64.8, or 71.0, respectively. The designs are complementary; gains from adding one component alone do not substitute for its contribution in the full configuration.
- Figure 4 contrasts folding paper with typing on a keyboard to illustrate disambiguation of similar body trajectories. Grad-CAM and t-SNE provide qualitative evidence. The text cache contains neither readable image pixels nor curve values, so no localization accuracy or clustering metric is inferred.
- Efficiency must distinguish recognition with cached cues from online extraction. The following table preserves the measurement scopes of Tables 4, 9, and 10; 1398.0 FPS is not treated as complete video-to-label throughput.
| Measurement scope and configuration | Parameters (M) | GFLOPs | FPS | Additional information |
|---|---|---|---|---|
| PoseBridge recognition stage (Table 4) | 128.2 total; 3.8 trainable | 7.3 | 1398.0 | Includes skeleton and text encoders |
| Original RTMPose-s (Table 9) | 5.47 | 0.68 | 4740.71 | COCO AP 70.86 |
| Enhanced RTMPose-s (Table 9) | 8.23 | 0.83 | 2534.7 | COCO AP 70.24 |
| Skeleton baseline online pipeline (Table 10) | 131.0 | 8.009 | 1101.9 | Paper-defined cumulative pipeline |
| Full PoseBridge online pipeline (Table 10) | 136.4 | 8.168 | 934.6 | Paper-defined cumulative pipeline |
Semantic enhancement reduces HPE AP by 0.62, and cumulative online FPS is also below the skeleton baseline. Better action recognition is therefore not directly explained by better pose localization. The frame/clip measurement convention and complete inclusion of human detection and decoding still require implementation-level verification.
Highlights & Insights¶
- Identify where information is lost before choosing where to align representations. If the input has already compressed away discriminative evidence, a more elaborate downstream mapping may still fit an insufficient observation.
- Turn pose-estimation intermediates into task-relevant reusable resources rather than generic RGB features. Body priors, semantic supervision, and skeleton queries jointly constrain their use, suggesting applications that rely on keypoints but also need interaction evidence.
- Adapt both the query and prototype instead of assigning the entire modality gap to one network. Unseen prototype adaptation transfers neighbors' cross-modal displacements while retaining the target's own text semantics rather than copying a seen visual centroid.
Limitations & Future Work¶
- The authors explicitly target an HPE-aware protocol, not a strictly coordinate-only one. Applications with existing skeleton files but no original video or HPE intermediates cannot directly use the full method.
- Severe occlusion, small objects, objects outside body-attended regions, motion blur, and crowds can degrade both skeletons and semantic cues. More robust body-region modeling and inexpensive multiscale evidence preservation are concrete directions.
- COCO captions provide weak image-level supervision and can describe entities other than the current person; text neighbors may also have different interaction structures. Multi-person caption association, neighbor reliability, and negative transfer to unseen classes need further analysis.
- Numerical reporting contains unresolved issues: Table 1 lists ReViSE 55/5 with \(S=74.2\), \(U=34.7\), and \(H=29.2\), which does not follow the harmonic-mean definition. This note does not replace the authors' numbers with recomputed values. Table 11 lists 766.9 M parameters for ViTPose; that anomalous value is likewise not silently corrected.
- Random splits report only three-run means without variance; the text cache does not support reconstruction of precise hyperparameter curves in Figures 5–6. Reproduction also requires verification of neighborhood temperature, GZSL calibration, and implementation details. This version promises code release but supplies no verifiable repository link.
- Omitting an RGB backbone does not remove visual privacy concerns or bias: HPE still accesses video and inherits biases from COCO, captions, and pretrained models. Deployment requires scenario-specific validation, data minimization, and human oversight.
Related Work & Insights¶
- vs PURLS / TDSM: PURLS strengthens part-level skeleton–text relations, while TDSM improves alignment through a diffusion-based mechanism. PoseBridge intervenes earlier to recover evidence lost before skeleton output; these directions need not be mutually exclusive.
- vs Neuron / SkeletonContext: Language context enriches class descriptions, whereas PoseBridge obtains sample-specific evidence from the current video's HPE stream. The distinction is the source of semantic information, not merely the text encoder.
- vs BSZSL / SKI-VLM: Additional RGB pathways provide richer observations but have different costs and modality protocols. PoseBridge reuses HPE and conditions on 2D skeletons, which does not imply an information budget identical to pure-skeleton methods.
- vs PFMESR: Both consider pose-estimation features. PoseBridge adds pose-structured refinement, body pooling, semantic supervision, and unseen prototype adaptation for zero-shot semantic transfer rather than generic feature fusion.
Rating¶
- Novelty: 4/5. Moves the zero-shot bottleneck upstream to skeletonization and bridges semantics on both query and prototype sides.
- Experimental Thoroughness: 4/5. Covers standard, random, in-the-wild, same-skeleton-format, and component-ablation settings, with remaining protocol and statistical caveats.
- Writing Quality: 4/5. Connects motivation to the method and clarifies training and cost scopes in the appendix, although several table values need verification.
- Value: 4/5. Useful for action recognition with HPE-stream access, but not directly applicable to coordinate-only deployments.