Skip to content

UF0-6D: Unified Flow-based Zero-Shot 6D Object Pose Estimation without Refinement

Conference: ECCV 2026
Paper: ECCV Paper Page
Code: Available soon
Area: 3D Vision / Robotics & Embodied AI
Keywords: zero-shot 6D pose estimation / conditional flow matching / SE(3) manifold / refinement-free / Riemannian optimal transport

TL;DR

UF0-6D reformulates zero-shot 6D object pose estimation as conditional Riemannian flow matching on the \(SE(3)\) manifold, learning an instance-conditioned posterior via geodesic-consistent bridge velocities and symmetry-aware optimal transport to replace coarse-to-refine cascades and iterative render-and-compare loops with a single ODE integration pass.

Background & Motivation

Estimating the rigid 6D transformation—comprising 3D rotation \(R \in SO(3)\) and 3D translation \(t \in \mathbb{R}^3\)—is a fundamental perception capability underlying robotic manipulation, autonomous grasping, and augmented reality. However, in open-world zero-shot scenarios involving unseen target objects, heavy occlusion, background clutter, textureless surfaces, and inherent pose ambiguities caused by geometric symmetries pose formidable challenges. Existing CAD-based zero-shot pipelines overwhelmingly follow a multi-stage paradigm: retrieving an initial template or coarse hypothesis, followed by repeated render-and-compare iterations to correct discretization errors. Meanwhile, model-free few-shot alternatives rely on fragile sparse feature correspondences that degrade severely under occlusion and textureless conditions, introducing substantial latency and engineering complexity.

The fundamental reason prior methods depend so heavily on iterative refinement is their inability to explicitly represent the multi-modal conditional posterior distribution of poses under geometric ambiguity. While generative paradigms such as score-based diffusion models (e.g., GenPose) attempt to model pose distributions, estimating score functions on high-dimensional non-Euclidean manifolds like \(SE(3)\) is notoriously unstable, especially when rotational symmetries induce complex multi-modal probability landscapes. Recently developed flow matching techniques offer a deterministic probability-flow alternative without score estimation, yet existing Riemannian formulations have been largely restricted to category-level generation without accommodating instance-specific query observations and heterogeneous object representations.

To resolve this tension, this work argues that multi-stage hypothesis filtering and rendering loops can be eliminated by directly learning an instance-conditioned Riemannian probability flow on the rigid Lie group \(SE(3)\). Core idea: reformulate zero-shot 6D pose estimation as conditional Riemannian flow matching (CRFM) on \(SE(3)\), mapping query crops and object representations into a shared point-token geometry condition, supervising the vector field with geodesic-consistent bridge velocities and symmetry-aware Riemannian optimal transport, and performing inference via a single ODE integration pass.

Method

Overall Architecture

UF0-6D presents an end-to-end, single-stage generative framework. Given a detected target crop (RGB or RGB-D) and an object geometry representation (CAD renders or sparse reference-view images), the system first maps both inputs into a unified point-token set and encodes them into an instance condition vector via cross-attention. Conditioned on this representation, the network models a deterministic probability flow on the Lie group \(SE(3)\), transporting an isotropic base distribution on the Lie algebra to the target pose posterior along geodesic paths. During inference, without any coarse-to-fine cascades or iterative rendering, the target 6D pose is sampled directly via a single numerical ODE integration pass.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Observed Input<br/>Query crop image I_crop (+ depth D_crop)"] --> Cond["Unified Geometry Conditioning<br/>Point-token extraction and cross-set attention"]
    Obj["Object Prior<br/>CAD rendered views / 3D-GS reference views"] --> Cond
    Cond --> Prior["Manifold Base Distribution<br/>g_0 ~ rho_0 on SE(3) initialization"]
    Prior --> Geodesic["Geodesic-Consistent Flow Matching<br/>Decoupled SO(3) and R^3 bridge velocity fitting"]
    Geodesic --> SymOT["Symmetry-Aware Riemannian OT<br/>Shortest geodesic target selection on se(3)"]
    SymOT --> ODE["Refinement-Free Single-Stage Inference<br/>Single ODE integration and likelihood readout"]
    ODE --> Out["Final 6D Object Pose<br/>g = [R, t] in SE(3)"]

Key Designs

1. Unified Geometry Conditioning: Bridging Model-Based and Model-Free Settings via Point Tokens To eliminate the architectural divergence between CAD-based and reference-based pose estimation, UF0-6D introduces an instance-conditioned representation \(c = \Psi(\mathcal{P}^q, F^q, \mathcal{P}^o, F^o)\). For the detected query crop \(I^{\text{crop}}\), a Vision Transformer (ViT) backbone extracts dense patch tokens reshaped into a feature map \(E^q\), from which \(N_q\) 2D points (or 2.5D points when depth is provided) \(\mathcal{P}^q\) and their indexed point features \(F^q\) are sampled. For the object template, if a CAD mesh is available, \(V\) views and depths are rendered; if only sparse reference images are available, a compact 3D proxy is rapidly reconstructed via 3D Gaussian Splatting (3D-GS) to re-render consistent views. The visible pixels across views are processed by the same ViT to construct the object point-token set \((\mathcal{P}^o, F^o)\). A cross-attention encoder \(\Psi\) fuses query and template tokens into a shared conditioning vector \(c\), maintaining full compatibility across CAD-based and reference-based regimes.

2. Geodesic-Consistent Flow Matching: Manifold-Preserving Probability Flows on \(SE(3)\) To bypass the topological distortion and score divergence of Euclidean diffusion on geometric manifolds, the continuous normalizing flow is established directly on the Lie group \(SE(3) \simeq SO(3) \times \mathbb{R}^3\). For an initial sample \(g_0 = [R_0, t_0]\) drawn from an isotropic Lie-algebra base distribution and a target ground-truth pose \(g_1 = [R_1, t_1]\), a time-dependent geodesic bridge \(u(t)\) is constructed for \(t \in [0, 1]\). The translation component follows Euclidean linear interpolation \(t_t = (1-t)t_0 + t t_1\), while 3D rotation follows the Lie-algebra geodesic on \(SO(3)\): $\(R_t = R_0 \exp\Big(t \log(R_0^\top R_1)\Big)\)$ The target bridge velocity in body-frame coordinates yields a constant Lie-algebra angular velocity \(\omega = \log(R_0^\top R_1) \in \mathfrak{so}(3)\), with \(\dot{R}_t = R_t \omega\) and linear velocity \(\dot{t}_t = t_1 - t_0\). The model trains a velocity network \(v_\theta(t, g_t, c)\) to match this geodesic vector field under a product Riemannian metric, ensuring that the probability trajectory remains strictly on the manifold throughout continuous time.

3. Symmetry-Aware Riemannian Optimal Transport: Resolving Multi-Modal Ambiguities Symmetric objects possess an equivalence set of ground-truth poses under their symmetry group \(\mathcal{S} \subset SE(3)\), given by \(\mathcal{Y}(g_1) = \{g_1 \circ s \mid s \in \mathcal{S}\}\). Assigning supervision targets arbitrarily causes the velocity field to average across mutually conflicting modes, leading to unphysical off-manifold predictions. UF0-6D integrates a Riemannian optimal transport (ROT) objective. Defining the transport cost via the \(\mathfrak{se}(3)\) matrix logarithm \(c(g_a, g_b) = \|\log(g_a^{-1} g_b)\|_{\mathfrak{se}(3)}\), the training protocol selects the symmetry-equivalent pose that minimizes the geodesic distance from the initial source sample \(g_0\): $\(\tilde{g}_1 = \arg\min_{g \in \mathcal{Y}(g_1)} \|\log(g_0^{-1} g)\|_{\mathfrak{se}(3)}\)$ By guiding each trajectory toward its nearest symmetric target \(\tilde{g}_1\), this formulation prevents mode interference from first geometric principles without requiring ad-hoc symmetry classification branches.

4. Refinement-Free Single-Stage Inference: Single-Pass ODE Integration and Likelihood Readout During test-time inference, UF0-6D abandons multi-stage cascades and render-and-compare refinement loops, directly integrating the probability-flow ODE \(dg_t/dt = v_\theta(t, g_t, c)\). In standard model-based evaluation, a single trajectory (\(M=1\)) yields the final prediction with minimal computation. In model-free settings where multiple hypotheses (\(M > 1\)) are sampled, candidate poses are ranked without training an auxiliary energy network by tracking the log-density evolution \(\frac{\partial}{\partial t} \log p_t(g_t) = -\nabla_g \cdot v_\theta(t, g_t, c)\). The Riemannian divergence is estimated efficiently using Hutchinson's unbiased trace estimator with random vectors \(\epsilon \sim \mathcal{N}(0, I)\). Candidates are aggregated into the final estimate \(\hat{g} = [\hat{R}, \hat{t}]\) via Euclidean weighted mean for translation and a likelihood-weighted Fréchet mean on \(SO(3)\) for rotation.

Loss & Training

UF0-6D is trained purely on synthetic datasets and tested zero-shot on real benchmarks. The conditional Riemannian flow matching loss minimizes velocity regression errors under the bi-invariant product metric: $\(\mathcal{L}_{\text{CRFM}}(\theta) = \mathbb{E}_{(c, g_1)\sim \mathcal{D}, g_0 \sim \rho_0, t \sim \mathcal{U}[0, 1]} \left[ \| v_\theta(t, g_t, c) - u(t; g_0, \tilde{g}_1) \|_{SE(3)}^2 \right]\)$ where \(\|\cdot\|_{SE(3)}^2\) couples the Killing-form metric on \(SO(3)\) with the Euclidean norm on \(\mathbb{R}^3\), and \(g_t\) is the geodesic interpolated state between \(g_0\) and the optimal symmetric target \(\tilde{g}_1\).

Key Experimental Results

Main Results

On the seven core datasets of the BOP Benchmark (LM-O, T-LESS, TUD-L, IC-BIN, ITODD, HB, YCB-V) using CNOS detections, UF0-6D is evaluated on Average Recall (AR, the mean across \(AR_{\text{VSD}}, AR_{\text{MSSD}}, AR_{\text{MSPD}}\)) and per-image runtime:

Paradigm Method Input Modality Pose Refinement Mean AR ↑ Runtime (s) ↓
With Refinement (Single Hyp.) MegaPose (CoRL 2022) RGB MegaPose Refiner 50.9 31.724
With Refinement (Single Hyp.) GenFlow (CVPR 2024) RGB GenFlow Refiner 55.7 10.553
With Refinement (Single Hyp.) Co-op (CVPR 2025) RGB Co-op Refiner 64.0 1.852
With Refinement (Single Hyp.) SAM-6D (CVPR 2024) RGB-D SAM-6D Refiner 70.4 4.367
With Refinement (Single Hyp.) Co-op (CVPR 2025) RGB-D Co-op Refiner 73.6 2.331
With Refinement (Multi Hyp.) FoundationPose (CVPR 2024) RGB-D FoundationPose Refiner 73.4 29.317
With Refinement (Multi Hyp.) Co-op (CVPR 2025) RGB-D Co-op Refiner 75.2 7.162
Without Refinement (Coarse) MegaPose (CoRL 2022) RGB-D None (Coarse Only) 20.8 15.465
Without Refinement (Coarse) GenFlow (CVPR 2024) RGB-D None (Coarse Only) 23.5 3.839
Without Refinement (Coarse) FreeZe (ECCV 2024) RGB-D None (Coarse Only) 69.3 13.474
Without Refinement (Coarse) Co-op (CVPR 2025) RGB None (Coarse Only) 58.4 0.979
Ours (Refinement-Free) UF0-6D (RGB) RGB None (Single-Stage Flow) 70.6 0.379
Ours (Refinement-Free) UF0-6D (RGB-D) RGB-D None (Single-Stage Flow) 81.2 0.832

Per-dataset AR breakdown (RGB-D): LM-O: 75.9, T-LESS: 73.3, TUD-L: 95.9, IC-BIN: 71.2, ITODD: 73.2, HB: 90.0, YCB-V: 88.7.

In the model-free zero-shot setup on YCB-Video without fine-tuning, performance is measured by AUC of ADD and ADD-S:

Method Ref. Views Finetune-free ADD-S (AUC) ↑ ADD (AUC) ↑
PREDATOR (CVPR 2021) 16 ✓ 71.0 24.3
LoFTR (CVPR 2021) 16 ✓ 52.5 26.2
FS6D-DPM (CVPR 2022) 16 ✗ 88.4 42.1
FoundationPose (CVPR 2024) 16 ✓ 97.4 91.5
UF0-6D (Ours) 3 ✓ 93.8 86.1
UF0-6D (Ours) 6 ✓ 96.8 91.0
UF0-6D (Ours) 12 ✓ 97.6 91.8

Ablation Study

Ablation on geometry conditioning components across the seven core BOP datasets, and the impact of numerical ODE integration steps on LM-O:

Modality Setting \(P^q\) (Query Pts) \(F^q\) (Query Feat) \(P^o\) (Obj Pts) \(F^o\) (Obj Feat) Mean AR ↑ Note
RGB w/o both feats ✓ ✗ ✓ ✗ 42.7 Geometry-only tokens collapse
RGB w/o query feat ✓ ✗ ✓ ✓ 55.9 Missing visual evidence (-14.7)
RGB w/o object feat ✓ ✓ ✓ ✗ 60.8 Missing template identity (-9.8)
RGB Full UF0-6D ✓ ✓ ✓ ✓ 70.6 Full conditioning baseline
RGB-D w/o both feats ✓ ✗ ✓ ✗ 55.4 Severe degradation without semantics
RGB-D w/o query feat ✓ ✗ ✓ ✓ 69.1 Query features dominate (-12.1)
RGB-D w/o object feat ✓ ✓ ✓ ✗ 72.9 Object features matter (-8.3)
RGB-D Full UF0-6D ✓ ✓ ✓ ✓ 81.2 Optimal multimodal synergy
Modality ODE Steps LM-O AR ↑ Runtime (s) ↓ Note
RGB 1 57.8 0.154 Large truncation error
RGB 5 62.9 0.226 Fast coarse estimation
RGB 10 65.4 0.342 Near saturation
RGB 20 (Default) 66.2 0.379 Optimal accuracy-speed trade-off
RGB 40 66.3 0.637 +0.1 gain with 1.7x latency
RGB-D 1 66.4 0.357 Single-step baseline
RGB-D 5 72.8 0.469 Balanced fast mode
RGB-D 10 75.1 0.641 High accuracy
RGB-D 20 (Default) 75.9 0.832 Default benchmark setting
RGB-D 40 76.0 1.342 Diminishing returns (+0.1)

Key Findings

  • Dominance of Query Semantic Features: Ablating appearance features drops RGB performance from 70.6 to 42.7, demonstrating that the flow model does not simply memorize structural priors; query features \(F^q\) contribute more heavily (-14.7 AR) than object features \(F^o\) (-9.8 AR) because observation-level visual cues dictate which pose modes remain plausible.
  • Trajectory Smoothness and ODE Saturation: Performance quickly saturates at 20 integration steps. Increasing steps from 20 to 40 yields only a marginal +0.1 AR improvement while substantially increasing execution time, confirming that the geodesic bridge constructs a smooth, low-curvature vector field.
  • Unprecedented Data Efficiency in Model-Free Setups: With only 12 uncalibrated reference images, UF0-6D outperforms FoundationPose configured with 16 views (97.6 vs. 97.4 ADD-S), and retains 93.8 ADD-S with as few as 3 views, underscoring the resilience of 3D Gaussian Splatting proxy conditioning.

Highlights & Insights

  • Paradigm Shift in 6D Pose Estimation: Proves that iterative render-and-compare refinement is an artifact of suboptimal posterior modeling, and that continuous normalizing flows on \(SE(3)\) can achieve superior precision in a single ODE pass.
  • Principled Handling of Symmetry: Formulating symmetry selection as Riemannian optimal transport on Lie algebras solves multi-modal target assignment geometrically, preventing off-manifold trajectory collapse.
  • Unified Representation Architecture: The point-token conditioning interface seamlessly bridges CAD meshes and sparse-view radiance fields, enabling a single trained model to excel across both model-based and model-free regimes.

Limitations & Future Work

  • Vulnerability to Extreme Initialization Offsets: When extreme occlusions or severe out-of-distribution camera distortions occur, trajectories starting from distant base samples may converge into suboptimal local symmetry modes.
  • Preprocessing Overhead for Model-Free Scenarios: While flow inference takes under a second, reconstructing 3D Gaussian Splatting proxies from reference views introduces an upfront offline computation cost.
  • Adaptive ODE Solvers: The current pipeline uses fixed-step integration; developing adaptive Riemannian solvers (e.g., adaptive Runge-Kutta on Lie groups) could further reduce inference steps to under 5 for easy instances.
  • vs FoundationPose (CVPR 2024): FoundationPose relies on coarse candidate ranking followed by iterative render-and-compare refinement, incurring tens of seconds per image. UF0-6D replaces this multi-stage pipeline with single-stage flow matching, outperforming FoundationPose by +7.8 AR on RGB-D (81.2 vs. 73.4) while executing over 35x faster.
  • vs GenPose (NeurIPS 2023) & RFMPose (NeurIPS 2026): GenPose relies on Euclidean score matching and struggles with manifold topology and multi-modal rotational symmetry. RFMPose focuses on category-level generation. UF0-6D is the first to achieve instance-level zero-shot pose estimation on \(SE(3)\) with unified geometry conditioning and unsupervised likelihood evaluation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates zero-shot 6D pose estimation as conditional Riemannian flow matching on \(SE(3)\) with symmetry-aware optimal transport, fundamentally eliminating iterative refinement loops.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across the seven core BOP datasets and YCB-Video spanning both model-based and model-free regimes, accompanied by extensive ablations and trajectory visualizations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulation of Lie algebra dynamics, cohesive narrative flow, and well-structured design justifications.
  • Value: ⭐⭐⭐⭐⭐ Establishes a new state-of-the-art on the BOP benchmark while reshaping the accuracy-speed Pareto frontier, offering immediate practical utility for real-time robotic manipulation.