Skip to content

PoseImageNet: Pose Estimation for Extensive Classes Based on Rich Structure Prototypes

Conference: ECCV 2026
Paper: ECCV Official
Area: Human Understanding / Object Pose Estimation
Keywords: Pose Estimation, Structure Prototype, Thin Plate Spline (TPS), Cross-Class Query Matching, Multi-Class Benchmark

TL;DR

Addressing the long-standing restriction of pose estimation to single classes and rigid topologies, this work introduces PoseImageNet on ImageNet spanning 720 semantic classes and 3,500+ structure prototypes grounded on TPS deformation alignment, alongside a momentum cost-matched query decoder baseline achieving 57.0% [email protected] across extensive classes.

Background & Motivation

Perceiving an object's spatial location, segmentation mask, and structural skeleton constitutes foundational capabilities of human visual intelligence, respectively corresponding to object detection, semantic segmentation, and pose estimation. While modern detectors and segmentors exhibit open-world parsing capability across thousands of diverse object categories on large-scale datasets such as ImageNet and LVIS, pose estimation has remained largely confined to a handful of specific single-class domains, primarily human bodies, faces, vehicles, and common animals. Although complex real-world scenes abound with heterogeneous physical objects whose articulated topological configurations are pivotal for high-level tasks like action analysis, human-object interaction, and robotic manipulation, a standardized fully-supervised benchmark for extensive multi-class pose estimation has remained conspicuously missing.

Recent explorations attempt to bridge this gap by concatenating prior single-class datasets into multi-class benchmarks (e.g., MP-100, UniKPT) or shifting towards few-shot and zero-shot transfer settings. However, naive multi-dataset concatenation remains severely deficient in semantic richness and structural variabilityโ€”for instance, UniKPT encompasses only 4 coarse super-classes, falling far short of the diversity seen in ImageNet classification benchmarks over a decade ago. Crucially, objects within the same semantic class frequently exhibit radical topological divergences: a telephone category inherently contains both foldable flip-phones and rigid bar smartphones; teapots or musical instruments present variations exceeding cross-class differences. Forcing uniform keypoint topologies across a semantic label induces intractable ambiguity and invalidates geometric correspondence.

To resolve this bottleneck, this work conceptualizes the visual pose of generic objects as a sparse topological abstraction that both captures the salient structure of an individual instance and characterizes non-rigid geometric deformation across structurally isomorphic instances. Core idea: Grounded on Thin Plate Spline (TPS) non-rigid deformation as a geometric alignment test, decouple pose annotation into prototype clustering followed by canonical keypoint labeling across 720 classes and 3,500+ prototypes, and propose a query-based baseline supervised by momentum cost bipartite matching to dynamically adapt shared queries to heterogeneous keypoint topologies.

Method

Overall Architecture

The proposed framework comprises two synergistic components: a rigorous dataset curation protocol based on Thin Plate Spline (TPS) warpability, and a unified query-based network tailored for multi-class, multi-structure pose estimation. At inference time, the model executes a two-stage forward pass: a ResNet-50 backbone extracts multi-scale feature maps, followed by global average pooling to classify the input into one of \(N\) structure prototypes. Concurrently, a bank of cross-class shared keypoint queries interacts with the multi-scale image features via a deformable Transformer decoder to predict candidate 2D coordinates and visibility confidences. Finally, the historical momentum cost matrix associated with the predicted prototype is retrieved to solve a bipartite Hungarian matching, adaptively organizing the candidate predictions into the ordered target keypoints of that prototype.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image x"] --> B["ResNet-50 Backbone<br/>Multi-scale Features {F_r}"]
    B --> C["Structure Prototype Classification<br/>Predict Prototype Logits y"]
    B --> D["Deformable Transformer Decoder<br/>M Shared Queries E Interaction"]
    D --> E["Candidate Keypoint Prediction<br/>Coordinates P and Visibilities V"]
    C --> F["Prototype Momentum Cost Matrix Retrieval<br/>Fetch Stored Cost C_bar"]
    E --> G["Bipartite Hungarian Matching<br/>Solve Assignment delta(k)"]
    F --> G
    G --> H["Final Output<br/>K Aligned Keypoints and Visibilities"]

Key Designs

1. Structure Prototype Decoupled Annotation Based on TPS Warpability

Defining keypoints across diverse objects hits an intrinsic roadblock when topological geometries diverge within identical semantic labels. Forcing unified keypoint definitions across an entire semantic class causes annotator confusion and geometric distortion. To establish an objective criterion, this paper formalizes two fundamental geometric axioms: first, keypoints must faithfully sketch the primary structural skeleton of the object; second, instances sharing the same structural prototype must be smoothly warpable into one another via Thin Plate Spline (TPS) transformations using their annotated keypoints as geometric control points. If two instances of the same class cannot be cleanly aligned through keypoint-driven TPS warping, they strictly belong to distinct structural prototypes.

Guided by this formulation, dataset annotation proceeds in two sequential phases: annotators first cluster ImageNet images into discrete structure prototypes \(P = \{x_i\}\) based on visual deformability, discarding ambiguous or severely occluded instances; then, within each identified prototype, a canonical reference instance is chosen to label the primary control keypoints that facilitate smooth mutual warping across all associated samples. Two complementary quality-control inspection viewsโ€”grid-based visual review and TPS canonical warping validationโ€”ensure cross-annotator rigor, ultimately yielding 3,500+ distinct prototypes across 720 semantic classes with up to 20 instances per prototype (complemented by balanced UniKPT subsets).

2. Cross-Class Shared Learnable Queries and Deformable Attention Decoding

Different structure prototypes exhibit wide variances in keypoint counts and semantics (ranging from 4 corner points for a notebook to 26 contour points for a suit). Conventional pose estimators assign dedicated channels or fixed spatial indices to specific keypoints (e.g., the \(i\)-th channel detects the \(i\)-th anatomical joint). Under thousands of heterogeneous prototypes, this fixed mapping causes severe index misalignmentโ€”for example, the "eye" keypoint corresponds to index 3 in avian classes but index 7 in canines.

To enable cross-category knowledge sharing, the network maintains a global pool of \(M\) learnable keypoint queries \(E \in \mathbb{R}^{M \times D}\) (configured with \(M = 300 \ge K = 100\)). These queries interact with multi-scale backbone feature maps \(\{F_r\}_{r=1}^3\) and attend to each other via a deformable Transformer decoder: $\([\hat{P}; E'] = \mathcal{F}_{dec}(E, \{F_r\}_{r=1}^3)\)$ A lightweight linear classification head on updated embeddings \(E'\) predicts binary visibilities \(\hat{V} \in \mathbb{R}^M\), while coordinate heads output unconstrained coordinates \(\hat{P} \in \mathbb{R}^{M \times 2}\). This architectural decouplement allows the model to learn localized structural primitives and geometric attributes that generalize across semantic boundaries.

3. Momentum Cost Matrix Bipartite Matching for Adaptive Output Assignment

To dynamically bind the \(M\) unordered candidate query outputs to the \(K\) padded target points of a specific structure prototype, the model maintains a persistent historical matching cost matrix \(\bar{C} \in \mathbb{R}^{M \times K}\) for each individual prototype. During training iterations, an immediate cost matrix \(C[m, k]\) is evaluated against ground-truth targets by combining coordinate distance and visibility binary cross-entropy: $\(C[m, k] = V^*[k] \cdot \|\hat{P}[m] - P^*[k]\|_1 + \alpha \cdot \text{BCE}(\hat{V}[m], V^*[k])\)$ where \(\alpha\) is a balancing coefficient. To prevent oscillatory query assignment across training steps and encourage individual queries to specialize in consistent geometric features, the cost matrix is updated via exponential moving average: $\(\bar{C} \leftarrow (1 - \lambda)\bar{C} + \lambda C\)$ Applying the Hungarian algorithm on \(\bar{C}\) produces an optimal injective mapping \(\delta(k)\), routing the \(k\)-th ground-truth keypoint to the \(\delta(k)\)-th query. Unmatched queries (\(M - K\)) are explicitly suppressed. During inference, the converged prototype-specific cost matrix \(\bar{C}\) is fetched to resolve optimal matching deterministically.

Loss & Training

The overall multi-task training loss balances prototype classification and matched keypoint regression: $\(\mathcal{L}_{full} = \mathcal{L}_{ce}(y, y^*) + \beta \sum_{k=1}^K \mathcal{L}_{kp}[k]\)$ where \(\mathcal{L}_{ce}\) is cross-entropy loss over \(N\) structure prototypes, and \(\beta = 10\) is the task weighting parameter. For each assigned target keypoint \(k\), the supervised regression loss is formulated as: $\(\mathcal{L}_{kp}[k] = V^*[k] \cdot \|\hat{P}[\delta(k)] - P^*[k]\|_1 + \text{BCE}(\hat{V}[\delta(k)], V^*[k])\)$ The \(M - K\) unassigned queries are supervised by an auxiliary visibility suppression loss \(\mathcal{L}_{un}\). Hyperparameters are set to \(K=100, M=300, \lambda=0.8\), and \(\alpha=0.1\), with samples split into 70% training and 30% testing partitions per prototype.

Key Experimental Results

Main Results

The authors benchmarked 8 representative pose estimation paradigms adapted for multi-class visibility prediction on PoseImageNet across 13 super-classes. The primary evaluation metric is PCK-based [email protected] (\(T_p = 0.2\), averaged over 9 visibility thresholds from 0.1 to 0.9).

Method Paradigm animal vehicle struc. tool wear. Overall [email protected]
SimBase ResNet / Heatmap 42.3 27.7 47.7 35.2 37.8 39.9
QueryPose Transformer / Query 45.4 32.5 49.3 37.5 41.1 42.8
ViTPose ViT-B / Heatmap 46.6 32.6 50.4 39.2 42.2 43.8
PCT Transformer 47.2 32.8 51.2 39.7 43.1 44.6
ED-Pose Explicit Query 48.1 33.2 52.1 41.1 43.9 45.4
NerPE Neural Prior 49.0 34.2 53.7 42.7 45.2 46.8
GroupPose Group Query 49.2 35.9 55.1 43.2 46.2 47.7
RTMO One-Stage SOTA 49.9 36.8 55.9 43.9 46.4 48.2
Ours Deformable + Match 52.0 54.5 70.4 55.3 62.3 57.0

On the global benchmark, the proposed method achieves 57.0% [email protected], outperforming the strongest baseline RTMO (48.2%) by a significant margin of +8.8%. The performance gains are particularly pronounced on rigid and semi-rigid objects (+17.7% on vehicle, +14.5% on structure, and +15.9% on wear items).

Ablation Study

To dissect the impact of adaptive bipartite matching and prototype classification errors, the study evaluates fine-grained metrics from [email protected] to [email protected] and their mean average (AVG).

Config [email protected] [email protected] [email protected] [email protected] AVG Note
SimBase Baseline 2.6 12.8 26.4 39.9 20.4 Traditional fixed-channel heatmap baseline
RTMO Baseline 17.5 36.7 43.5 48.2 36.5 Advanced competitive one-stage detector
Full Model 27.2 43.3 51.8 57.0 44.8 Adaptive momentum bipartite matching
Fix Matching 8.0 17.8 26.8 35.3 22.0 Rigid 1-to-1 mapping drops AVG mAP by -22.8%
Using GT Class 38.6 61.9 73.3 79.5 63.3 Upper bound reaches 79.5% (+22.5%) with perfect prototype labels

Key Findings

  • Adaptive matching is indispensable for multi-structure generalization: Enforcing a fixed 1-to-1 index matching (Fix Matching) causes the AVG mAP to collapse from 44.8% to 22.0% (-22.8%), confirming that fixed channels fail catastrophically under diverse topologies and that dynamic bipartite matching is essential to unlock query sharing.
  • Prototype classification accuracy is the primary performance bottleneck: When using ground-truth prototype IDs during evaluation (Using GT Class), [email protected] jumps to 79.5% (AVG reaching 63.3%, a +18.5% gain). This reveals that spatial localization itself is highly accurate, but misclassified prototypes cause mismatched cost matrices and subsequent index permutation errors.
  • Robust inter-annotator consistency: Quality verification demonstrates high consistency with Fleiss' Kappa of 82.7 for prototype assignment and PCK-Match of 98.2 for keypoint placement, proving the operational objectivity of TPS-grounded pose annotation.

Highlights & Insights

  • Formulating visual pose via TPS control-point warpability: The work transforms an ambiguous, subjective notion of "pose" into an objective mathematical condition: whether two objects can smoothly align under non-rigid thin-plate-spline warping using keypoints as control vertices.
  • Momentum-accumulated matching stabilizes query specialization: In an end-to-end Transformer setup, raw batch Hungarian matching frequently oscillates, causing query competition and mode collapse; maintaining an EMA cost matrix (\(\lambda = 0.8\)) per prototype fosters smooth query specialization.
  • High transferability to articulated robotics and 3D CAD alignment: Sparse keypoints across thousands of daily objects naturally align with functional affordances (handles, hinges, grasp contact regions), serving as an off-the-shelf foundation for robotic manipulation and open-vocabulary object pose estimation.

Limitations & Future Work

  • Admitted limitations: Structure prototypes remain strictly partitioned within individual semantic categories rather than hierarchically unified across semantic boundaries (e.g., baseballs and basketballs could share an identical rigid sphere prototype); additionally, errors in top-level prototype classification directly corrupt output keypoint ordering.
  • Future directions: Integrating vision-language foundation models (e.g., CLIP / VLMs) to inject hierarchical structural taxonomy, or exploring permutation-invariant keypoint set prediction to diminish coupling with discrete prototype classifiers.
  • vs UniKPT / MP-100: UniKPT and MP-100 naively merge existing single-class benchmarks, covering only 4 super-classes or 100 categories and failing to address multi-topology intra-class variation; PoseImageNet provides 720 classes with 3,500+ TPS-grounded prototypes directly from ImageNet.
  • vs Conventional Fixed-Channel Decoders (SimBase / RTMO): Standard pose architectures rely on rigid channel-to-keypoint binding, which breaks down entirely when scaled to thousands of heterogeneous structures; the proposed cross-class query sharing with momentum cost matching provides an elegant unified alternative.

Rating

  • Novelty: โญโญโญโญโญ [Pioneering dataset definition anchoring visual pose to TPS deformation alignment across extensive ImageNet classes]
  • Experimental Thoroughness: โญโญโญโญโ˜† [Comprehensive 3,500+ prototype benchmark with revealing ablations on matching dynamics and classification bounds]
  • Writing Quality: โญโญโญโญโญ [Clear mathematical grounding, elegant narrative structure, and faithful correspondence between architecture and designs]
  • Value: โญโญโญโญโญ [Fills an essential gap in fully-supervised multi-class pose estimation, providing high-value priors for generic object understanding and robotics]