Skip to content

Occlusion-Resilient Category-Agnostic Pose Estimation with Conditional Flow Matching

Conference: ECCV 2026
Paper: CVF Open Access
Area: Human Understanding / 3D Vision / Image Generation
Keywords: Category-Agnostic Pose Estimation, Flow Matching, Occlusion Robustness, Graph Flow Encoder, Riemannian Manifold

TL;DR

FlowCape reformulates Category-Agnostic Pose Estimation (CAPE) as a structured pose transport process driven by conditional flow matching, integrating heatmap-guided initialization to anchor candidate modes and a topology-constrained Riemannian Pose Head with Graph Flow Encoders to achieve state-of-the-art pose recovery under severe occlusion.

Background & Motivation

Category-Agnostic Pose Estimation (CAPE) aims to localize category-defined keypoints on an arbitrary query image given only a few annotated support examples or textual point descriptions. In sharp contrast to category-specific pose estimation, which relies on fixed skeletons and closed-set supervised appearance priors, CAPE demands that models transfer structural and semantic keypoint knowledge to previously unseen categories whose physical shapes, keypoint semantics, and topological configurations vary drastically. This generalization challenge becomes particularly severe under heavy occlusion or self-occlusion, where local visual evidence is completely absent.

Most existing CAPE frameworks operate under a discriminative regression or local feature matching paradigm in Euclidean space. While effective when joints are clearly visible, these approaches suffer from two fundamental bottlenecks under severe occlusion. First, localized correlation matching provides weak global structural control, causing invisible or heavily occluded keypoints to drift toward spurious background clutter or anatomically implausible positions. Second, deterministic regression models struggle with the "generative fallacy": when faced with multimodal uncertainty arising from missing evidence, they tend to output over-confident, arbitrary average predictions. Although diffusion-based pose estimators can capture uncertainty, their extensive multi-step reverse denoising steps incur prohibitive computational latency.

Conditional Flow Matching (FM) provides a continuous, simulation-free generative alternative by learning a deterministic velocity field that transports an initial distribution to the target pose along straight probability paths. However, naively applying flow matching to CAPE causes catastrophic drift because sampling from an uninformed prior leaves the model struggling with excessive transport distance and multimodal ambiguity under missing visual cues. Core idea: model category-agnostic pose estimation under severe occlusion as structured conditional flow transport over a skeleton-regularized manifold, anchoring the initial state via cross-modal heatmap-guided initialization and directing probability-flow ODE rollout through a topology-aware Riemannian Pose Head.

Method

Overall Architecture

FlowCape is structured around three primary stages: Heatmap-Guided Initialization, Conditional Flow Matching velocity prediction, and Probability-Flow ODE rollout inference. Given an input query image \(I_q\), a category skeleton graph \(\mathcal{G}\), and textual point-description embeddings \(T_s = \{d_i\}_{i=1}^N\), the pipeline first extracts spatial visual features \(F_q = \Phi_q(I_q)\) via an image encoder and text embeddings via a frozen text model. The Heatmap-Guided Initialization module computes cross-modal similarity between visual feature tokens and text descriptors, softly decoding them into continuous coordinates and blending them with a centered geometric anchor to obtain a stable initial pose \(x_0\). A shared Riemannian Pose Head containing stacked Graph Flow Encoders (GFE) predicts the conditional velocity field \(v_\theta(x_t, t \mid c)\) by jointly reasoning over query image tokens, textual semantics, and the skeleton graph topology. At test time, the final keypoint coordinates are generated by numerically integrating the learned probability flow ODE using explicit Euler steps augmented with heatmap potential guidance.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Query Image Iq + Text Descriptions Ts + Skeleton Graph G"] --> B["Feature Extraction<br/>SwinV2 Image Features Fq and GTE Text Semantics di"]
    B --> C["Heatmap-Guided Initialization<br/>Cross-modal similarity mapping + confidence-adaptive blending"]
    C --> D["Graph Flow Encoder Velocity Field Prediction<br/>Self-attention + Cross-attention + Skeleton graph smoothing"]
    D --> E["Skeleton-Aware Riemannian Flow Supervision<br/>Laplacian metric regularization + rollout consistency losses"]
    E --> F["Probability Flow ODE Rollout<br/>Euler numerical integration + heatmap potential guidance"]

Key Designs

1. Heatmap-Guided Initialization: Anchoring Local Neighborhoods against Transport Drift Standard generative flow matching samples the initial state from an isotropic Gaussian or uniform prior. In CAPE, this wide prior distribution forces the ODE trajectory to traverse extensive ambiguous regions, which easily destabilizes the transport under severe occlusion where pose posteriors are highly multimodal. To resolve this, FlowCape reuses the heatmap branch of the Riemannian Pose Head at \(t=0\) to construct an instance-grounded initial pose. It calculates the dot-product similarity between \(\ell_2\)-normalized image tokens and keypoint text embeddings: \(H_{iuv} = \langle \hat{d}_i, \hat{f}_{uv} \rangle / \tau_h\). Differentiable soft-argmax converts the similarity heatmaps into continuous 2D coordinates \(x_{0,i}^{\mathrm{hm}}\). To prevent spurious peak activations in heavily occluded keypoints from poisoning the initialization, a centered geometric template \(q_0 = 0.5 \cdot \mathbf{1}^{N \times 2}\) is blended with the heatmap coordinates based on a confidence weighting factor \(\alpha_i = \sigma((\sigma(p_i) - \delta)\gamma)\), where \(p_i = \max_{u,v} H_{iuv}\): $\(x_{0,i} = \alpha_i x_{0,i}^{\mathrm{hm}} + (1 - \alpha_i) q_{0,i}\)$ During training, the raw initialization \(x_0^{\mathrm{raw}}\) is supervised directly, but its gradient is severed via stop-gradient \(x_0 = \text{sg}(x_0^{\mathrm{raw}})\) before feeding into the flow matching pipeline, isolating initialization representation learning from trajectory velocity gradients.

2. Graph Flow Encoder: Fusing Structural Graph Topology and Multimodal Context To preserve structural consistency throughout continuous transport, the intermediate state \(x_t\) must continuously condition on local image evidence and global topological constraints. The shared Riemannian Pose Head contains a specialized Graph Flow Encoder (GFE). The \(i\)-th keypoint token is initialized by projecting both the support semantic description and the current spatial state: \(z_i^{(0)} = W_d d_i + W_x x_{t,i}\), followed by injecting sinusoidal time embeddings and global text modulation \(\bar{d}\) via Adaptive Layer Normalization (AdaLN). Inside each GFE block, the network alternates across three operations: self-attention across keypoint tokens to model cross-joint contextual dependencies, cross-attention into projected query image features \(M\) to retrieve local visual cues, and graph convolution smoothing guided by the skeleton adjacency matrix \(\mathcal{G}\). This design ensures that velocity predictions for occluded joints are dynamically regularized by visible structural neighbors through topological graph message passing, yielding a physically plausible conditional velocity field \(v_\theta(x_t, t \mid c)\).

3. Skeleton-Aware Riemannian Flow Supervision and Trajectory Consistency To penalize structural tearing and non-rigid distortion during generative flow, FlowCape enforces flow-matching loss under a non-Euclidean Riemannian metric derived from the graph Laplacian \(L_{\mathcal{G}} = D - A\). Given the ground-truth target pose \(x_1\), the linear reference probability path is \(x_t = (1-t)x_0 + tx_1\) with constant tangent velocity \(u = x_1 - x_0\). Defining the velocity residual as \(r = v_\theta(x_t, t \mid c) - u\), the Riemannian flow matching loss is formulated as: $\(\mathcal{L}_{\mathrm{rfm}} = \frac{1}{\sum_{i=1}^N m_i} \mathbf{r}^T \mathbf{L}_{\mathcal{G}} \mathbf{r}\)$ This metric explicitly penalizes relative velocity discrepancies between topologically connected joints, complementing the standard Euclidean velocity regression loss \(\mathcal{L}_{\mathrm{flow}}\). Furthermore, FlowCape adopts a non-uniform time sampling \(t \sim \mathcal{U}(0,1)^{1/2}\) to concentrate supervision on early trajectory steps, and introduces an initial velocity alignment loss \(\mathcal{L}_{\mathrm{v0}}\) along with a multi-step unrolled ODE rollout loss \(\mathcal{L}_{\mathrm{ode}}\) over \(K_{\mathrm{tr}}\) steps to prevent error accumulation across inference integration.

4. Heatmap Potential-Guided Probability Flow ODE Rollout During test-time inference, the model initializes particles at \(x^{(0)} = x_0\) and solves the probability-flow ODE via explicit Euler numerical integration across \(K_{\mathrm{te}}\) steps. Because the shared Riemannian Pose Head simultaneously outputs refined heatmap predictions \(H^{(k)}\) at each step \(k\), FlowCape constructs a test-time potential function \(U(x; H^{(k)})\) by bilinearly sampling the heatmap responses at the current coordinate estimate. The spatial gradient of this potential \(g^{(k)} = \lambda_{\mathrm{kg}} \nabla_x U(x^{(k)}; H^{(k)})\) acts as an attractive force toward high-confidence visual evidence, yielding a modified velocity: \(\tilde{v}^{(k)} = v^{(k)} + g^{(k)}\). The state updates sequentially as: $\(x^{(k+1)} = x^{(k)} + \frac{1}{K_{\mathrm{te}}} \tilde{v}^{(k)}\)$ This guidance mechanism drives the particles along smooth, skeleton-consistent paths while drawing them into fine-grained local visual extrema, achieving coarse-to-fine localization.

Loss & Training

The overall training objective combines velocity matching, geometric regularization, and trajectory consistency: $\(\mathcal{L} = \lambda_{\mathrm{flow}}\mathcal{L}_{\mathrm{flow}} + \lambda_{\mathrm{rfm}}\mathcal{L}_{\mathrm{rfm}} + \lambda_{\mathrm{init}}\mathcal{L}_{\mathrm{init}} + \lambda_{\mathrm{v0}}\mathcal{L}_{\mathrm{v0}} + \lambda_{\mathrm{ode}}\mathcal{L}_{\mathrm{ode}} + \lambda_{\mathrm{hm}}\mathcal{L}_{\mathrm{hm}}\)$ Hyper-parameter ablations establish optimal weights at \(\lambda_{\mathrm{init}}=0.1\) and \(\lambda_{\mathrm{traj}}=1.0\). The framework is implemented in MMPose and trained with Adam for 200 epochs using a batch size of 16. The initial learning rate is \(1 \times 10^{-5}\), decayed by a factor of 10 at epochs 160 and 180. All images are resized to \(256 \times 256\) pixels. Training is conducted on an NVIDIA A100 GPU server.

Key Experimental Results

Main Results

To rigorously evaluate pose estimation under severe occlusion, the authors introduced the Occ80 benchmark, which features 80 diverse categories and 15,562 annotated instances characterized by heavy occlusion, divided across three non-overlapping category splits (S1โ€“S3). The table below details [email protected] comparisons on Occ80:

Backbone Method Source S1 S2 S3 Avg.
ResNet-50 CapeFormer CVPR 2023 77.63 77.67 71.13 75.48
DINOv2-ViT-B/14 EdgeCape ICLR 2026 82.23 81.21 76.35 79.93
SwinV2-Tiny GraphCape ECCV 2024 84.78 84.68 79.35 82.94
SwinV2-Tiny CapeX ICLR 2025 85.68 85.79 79.69 83.72
SwinV2-Tiny GenCape ICLR 2026 78.59 79.60 73.63 77.27
SwinV2-Tiny FlowCape ECCV 2026 86.72 86.59 80.23 84.51
SwinV2-Small GraphCape ECCV 2024 84.97 85.23 79.92 83.37
SwinV2-Small CapeX ICLR 2025 89.35 89.99 84.88 88.07
SwinV2-Small GenCape ICLR 2026 87.61 84.35 80.66 84.21
SwinV2-Small FlowCape ECCV 2026 91.56 90.49 85.32 89.12

On the standard MP-100 benchmark (1-shot setting, average of 5 splits), FlowCape-T achieves an average PCK of 88.40% (matching CapeX-T and outperforming it on Splits 3 and 4), while FlowCape-S reaches 91.65% PCK, surpassing CapeX-S (91.50%). This confirms that FlowCape maintains top-tier discriminative accuracy on generic, unoccluded pose estimation while delivering substantial gains on occluded targets.

Ablation Study

Ablations on Occ80 Split-1 using the SwinV2-Tiny backbone isolate the contributions of core modules and supervision objectives:

Configuration \(\mathcal{L}_{\mathrm{ODE}}\) H-Init \(\mathcal{L}_{\mathrm{init}}\) \(\mathcal{L}_{\mathrm{rfm}}\) \(\mathcal{L}_{\mathrm{flow}}\) \(\mathcal{L}_{v_0}\) [email protected] \(\Delta\)
Baseline (CapeX) - - - - - - 85.68 0.00
ODE trajectory only โœ“ - - - - - 72.81 -13.31
+ Heatmap Initialization (H-Init) โœ“ โœ“ โœ“ - - - 85.56 -0.12
+ Riemannian Graph Loss โœ“ โœ“ โœ“ โœ“ - - 86.22 +0.54
+ Euclidean Flow Loss โœ“ โœ“ โœ“ โœ“ โœ“ - 86.58 +0.90
+ Initial Velocity Regularization โœ“ โœ“ โœ“ - - โœ“ 86.41 +0.73
Full FlowCape โœ“ โœ“ โœ“ โœ“ โœ“ โœ“ 86.72 +1.04

Comparison across initialization modes on Occ80 Split-1 further highlights trajectory stability: - 20-step Query Centered (20S-Q): 83.13% - 30-step Query Centered (Query): 84.77% - Heatmap-Guided (HM): 86.72% - Confidence-Guided (Conf.): 86.71%

Key Findings

  • Initialization quality dictates flow viability: Operating ODE rollout without heatmap guidance causes accuracy to collapse by -13.31% (down to 72.81%). Because occluded pose posteriors are wide and multimodal, unconstrained integration from an uninformed center anchor accumulates drastic drift. Incorporating Heatmap-Guided Initialization bounds the integration trajectory within a plausible local neighborhood, recovering PCK to 85.56%.
  • Topological Riemannian regularization prevents skeleton tearing: Adding \(\mathcal{L}_{\mathrm{rfm}}\) yields a solid +0.54% gain over pure initialization. Qualitative error visualizations show that the Laplacian metric suppresses unnatural stretching and joint inversion in occluded animal limbs and human torsos.
  • Architectural depth trade-offs: The Riemannian Pose Head peaks at depth \(N_h=3\) (86.72% PCK). Increasing depth to \(N_h=4\) degrades accuracy to 84.96%, indicating that excessive head depth introduces over-parameterization and drift into the vector field. Hidden dimension \(D_h=256\) provides optimal trade-offs between representation capacity and computation.

Highlights & Insights

  • Continuous flow matching resolves the generative fallacy of CAPE: Instead of forcing a neural network to make brittle deterministic guesses under severe occlusion, FlowCape formulates keypoint localization as continuous transport along an optimal probability path, decoupling semantic matching from progressive geometric refinement.
  • Skeleton Laplacian embedded directly as a Riemannian metric: Rather than treating graph networks merely as feature extractors, FlowCape embeds the graph Laplacian into the velocity loss function, penalizing topological distortions mathematically in the velocity tangent space.
  • Unified dual-purpose head architecture: A single Riemannian Pose Head serves as both a cross-modal correlation matcher at \(t=0\) and a conditional vector field predictor during \(t>0\), maximizing parameter efficiency and ensuring shared feature representations across stages.

Limitations & Future Work

  • Inference latency from ODE integration: Solving the probability flow ODE requires 20โ€“30 Euler steps. Although substantially faster than 1000-step diffusion models, multi-step rollout still incurs noticeable latency compared to single-forward discriminative regression models, limiting high-frequency real-time deployment on edge devices.
  • Sensitivity to fine-grained textual prompt fidelity: Keypoint conditioning relies heavily on the expressiveness of text descriptions (e.g., from GTE embeddings). Ambiguous descriptions for exotic or anatomically complex animal joints may produce weak initial heatmaps.
  • Future directions: Integrating trajectory distillation techniques such as Rectified Flow or Consistency Models could compress the 30-step rollout into 1โ€“2 evaluation steps while preserving occlusion robustness.
  • vs CapeX / CapeLLM: CapeX introduced textual point explanations to eliminate visual support dependencies, but still relies on direct DETR-style coordinate regression that fails under heavy occlusion. FlowCape adopts textual descriptions while introducing continuous generative flow transport, achieving vastly superior topological resilience.
  • vs GraphCape / GenCape: While GraphCape and GenCape pioneered graph reasoning in CAPE via GNN layers and adaptive adjacency matrices, FlowCape uses graph structure not just for feature passing, but as a Riemannian metric operator that constrains velocity fields and trajectory dynamics.
  • vs DiffPose / FMPose3D: DiffPose introduces generative modeling via diffusion but requires expensive reverse sampling. FlowCape adopts straight-line probability paths via Flow Matching with task-specific heatmap initialization, achieving faster convergence and higher structural fidelity.

Rating

  • Novelty: โญโญโญโญโญ Pioneering application of conditional flow matching and Riemannian skeleton metrics to category-agnostic pose estimation.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive validation across MP-100 and the newly curated 80-category Occ80 severe-occlusion benchmark.
  • Writing Quality: โญโญโญโญโญ Rigorous mathematical derivations, coherent system narrative, and clean visual representations.
  • Value: โญโญโญโญโญ Establishes a robust, physically plausible paradigm for general keypoint localization under adverse visual conditions.