Skip to content

Geometry-Preserving Image Generation for 6D Object Pose Estimation

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/JiafengZhang-1117/GenerationPose
Area: Image Generation
Keywords: 6D object pose estimation, geometry-preserving generation, Diffusion Transformer (DiT), synthetic-to-real, ControlNet

TL;DR

To tackle the acute bottleneck of costly real annotations and the synthetic-to-real domain gap in 6D object pose estimation, this paper introduces GenerationPose—a Diffusion Transformer (DiT) framework integrating ControlNetGeom and mask-aware self-attention blocking to convert synthetic 3D CAD/mesh/point-cloud renders into photorealistic pseudo-real images that strictly preserve 6D object pose labels.

Background & Motivation

6D object pose estimation aims to estimate the 3D rotation and 3D translation of rigid objects from one or more images, serving as a critical cornerstone for robotic manipulation, augmented reality, and industrial automation. However, scaling up deep pose estimators is severely hindered by a real-world data bottleneck. Annotating precise 6D ground truth requires dedicated multi-sensor setups, calibration targets, and laborious alignment routines. Consequently, existing real datasets remain limited in object diversity, material properties, illumination variations, and complex scene clutter, causing pose models to struggle significantly when deployed to novel objects or unseen real-world environments.

Synthetic rendering from 3D CAD models or point clouds offers an infinitely scalable alternative with flawless, cost-free ground-truth poses. Nonetheless, synthetic images inevitably suffer from a substantial reality gap in surface textures, specular highlights, and contextual backgrounds compared to physical cameras. Conventional domain randomization perturbs lighting, colors, and backgrounds to force models to learn domain-invariant features, yet the resulting heuristic distributions fail to capture the photorealistic intricacies of natural scenes. Conversely, generic image-to-image translation techniques based on GANs or unconstrained diffusion models frequently cause subtle geometric warping, boundary drift, or perspective distortion, which directly invalidates the rendered 6D ground-truth pose annotations.

To bridge this sim-to-real gap without compromising geometric integrity, a generative pipeline must enforce explicit, decoupled architectural constraints over object geometry and spatial boundaries. Core Idea: Formulate synthetic-to-real appearance transfer as a geometry- and mask-conditioned latent diffusion framework (GenerationPose) built upon a Diffusion Transformer (DiT), which introduces a lightweight ControlNetGeom branch to inject rigid geometric priors and a mask-aware self-attention mechanism with foreground-to-background (FG→BG) blocking to eliminate background semantic contamination, producing diverse, photorealistic images while strictly preserving the underlying 6D object pose.

Method

Overall Architecture

The overall pipeline of GenerationPose takes a synthetic geometry render \(x_{\text{geom}}\) and its corresponding binary object mask \(m\) as structural and spatial conditions. Starting from latent Gaussian noise, the model iteratively denoises within the latent space of a pre-trained VAE, yielding a photorealistic pseudo-real image \(\tilde{x}\) that honors realistic lighting and texture distributions while maintaining sub-pixel pose alignment with the input geometry.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Synthetic render x_geom and mask m"] --> B["ControlNetGeom: Dual-Pathway Geometry Injection"]
    B --> C["Mask-Aware Self-Attention: FG→BG Blocking"]
    C --> D["Dual-Mask Objectives & Strength-Controlled Sampling"]
    E["Output: Pose-Consistent Pseudo-Real Image x_out"]
    D --> E

During training, the framework constructs pose-aligned triplets \((x_{\text{real}}, x_{\text{geom}}, m)\) under identical camera intrinsics and object poses. An object-centric square crop expanded from the bounding box of \(m\) is applied consistently to both \(x_{\text{real}}\) and \(x_{\text{geom}}\) and resized to \(256 \times 256\). The real crop is encoded by the VAE encoder into \(z_{\text{real}}\), and corrupted with Gaussian noise at timestep \(t\) to form \(z_t\). The DiT-B/2 backbone processes the patchified latent tokens modulated by adaptive layer normalization (AdaLN) embeddings. Three orthogonal streams drive the generation: a texture stream for latent denoising, a geometric stream via ControlNetGeom, and a spatial stream via mask-derived attention biasing.

Key Designs

1. ControlNetGeom: Dual-Pathway Geometry Injection Standard text or global image conditioning cannot provide the sub-pixel structural anchoring required to preserve fine pose cues, and relying solely on cross-attention in early, high-noise diffusion steps often permits geometric drift. To overcome this, the authors design ControlNetGeom, a lightweight convolutional branch running parallel to the DiT-B/2 backbone. It encodes the geometry render deterministically via the VAE posterior mean \(z_{\text{geom}} = E(x_{\text{geom}}).\text{mean}\), extracts features through compact convolutional layers, and interpolates them onto the DiT token grid to produce spatially aligned geometry tokens \(u \in \mathbb{R}^{B \times N \times D}\). ControlNetGeom injects geometric priors via two distinct pathways: first, projecting \(u\) into global conditioning tokens \(c_{\text{cond}} = W_{\text{cond}} u\) injected into cross-attention across all DiT blocks; second, generating early-layer residual hints \(h_i = W_i u\) for the first \(K\) shallow DiT blocks, added directly to the latent tokens: $\(x \leftarrow x + h_i, \quad i = 0, \dots, K-1\)$ With zero-initialized weights \(W_i\), this early-layer injection firmly locks in the rigid 3D object pose and wireframe boundaries when noise is high, while leaving deeper layers free to synthesize specular reflections, realistic materials, and contextual shading.

2. Mask-Aware Self-Attention: FG→BG Blocking and Spatial Decoupling In unconstrained self-attention, all foreground and background tokens attend to each other across the entire image grid. This global receptive field causes background clutter, distracting edges, and noise to leak into the object region, eroding silhouette boundaries that are vital for 6D pose estimators. To establish strict spatial isolation, the binary object mask is downsampled to the token grid \(\hat{m} \in \{0, 1\}^{B \times N}\), and an additive attention bias matrix \(M\) is injected into the self-attention layer of each DiT block: $\(M_{ij} = \begin{cases} \beta, & \text{if } \hat{m}_i = 1 \text{ and } \hat{m}_j = 0 \\ 0, & \text{otherwise} \end{cases}\)$ where \(\beta \ll 0\) (a large negative value driving attention weights to zero after Softmax). This explicitly blocks foreground queries from attending to background keys, preventing external contamination. Meanwhile, background queries can still attend to foreground tokens to generate physically consistent contact shadows and ambient occlusion at object boundaries, while cross-attention to geometry tokens remains unconstrained.

3. Dual-Mask Objectives & Strength-Controlled Sampling: Structure vs. Appearance Trade-off Because spatial attention constraints and loss optimization pose conflicting demands on mask boundaries—attention blocking benefits from a dilated mask to comfortably enclose soft object boundaries, while loss supervision requires a tight mask to exclude background noise—the method introduces a dual-mask strategy. A dilated, randomly shifted mask \(m_{\text{full}}\) serves as the robust attention bias, whereas an eroded, tightened mask \(m_{\text{loss}}\) provides clean foreground supervision. At test time, GenerationPose adopts an anchor-based DDIM sampling strategy governed by a strength hyperparameter and random style seeds. By mapping strength to an initial denoising timestep \(t_{\text{start}}\), users can smoothly modulate the degree of appearance freedom: lower strength retains strict geometric adherence, while higher strength allows vivid texture exploration and varied ambient illumination, all while maintaining boundary deviations well under one pixel.

Loss & Training

The overall training objective combines standard diffusion noise prediction with periodic multi-scale perceptual supervision: $\(\mathcal{L} = \mathcal{L}_{\text{diff}} + \mathbb{I}[s \bmod K_{\text{perc}} = 0] \cdot \left(\lambda_{\text{l1}} \mathcal{L}_{\text{L1-fg}} + \lambda_{\text{fg}} \mathcal{L}_{\text{perc-fg}} + \lambda_{\text{bg}} \mathcal{L}_{\text{perc-bg}}\right)\)$ Here, \(\mathcal{L}_{\text{diff}}\) optimizes noise prediction alongside a variance term. Every \(K_{\text{perc}}\) training steps, the predicted clean latent \(\hat{z}_0\) is decoded into \(\hat{x} = D(\hat{z}_0)\) to compute the foreground L1 loss \(\mathcal{L}_{\text{L1-fg}} = \|(\hat{x} - x_{\text{real}}) \odot m_{\text{loss}}\|_1\) and multi-layer VGG-based foreground perceptual loss \(\mathcal{L}_{\text{perc-fg}}\), supplemented by a weak background perceptual regularizer \(\mathcal{L}_{\text{perc-bg}}\). To ensure robust cross-domain generalization, synthetic geometry inputs \(x_{\text{geom}}\) undergo synthetic degradation (simulating point-cloud sparsity) while \(x_{\text{real}}\) remains pristine.

Key Experimental Results

Main Results

Evaluation is conducted under a rigorous cross-domain, zero-shot protocol: models are trained on uCO3D point-cloud renders and evaluated zero-shot on official CAD/mesh renders across BOP benchmarks (TYOL, LM, YCBV, ICMI). Geometric boundary fidelity is quantified via Normalized Contour Deviation (NCD, lower is better), Boundary F1 score (BF1@\(\tau\), higher is better), and Inlier ratio (Inlier@\(\tau\), higher is better) at \(\tau \in \{2, 5\}\) pixels. Distribution-level photorealism is evaluated using FID and KID.

Dataset NCD (%) ↓ BF1@2 ↑ BF1@5 ↑ Inlier@2 ↑ Inlier@5 ↑
TYOL 0.481 0.950 0.998 0.951 0.998
YCBV 1.469 0.838 0.945 0.846 0.951
ICMI 1.090 0.844 0.969 0.846 0.970
LM (LineMOD) 0.428 0.960 0.998 0.961 0.998

In terms of appearance quality, macro-averaged across LM, ICMI, TYOL, and YCBV against real image crops, GenerationPose substantially closes the reality gap:

Input FID ↓ KID ↓
Raw Render 132.84 0.0986
GenerationPose (Ours) 114.57 0.0794

For downstream 6D pose utility, models trained with GenerationPose pseudo-real data were evaluated using Wide-Depth-Range Pose (WDR) on LineMOD under five data regimes (ADI accuracy at 0.10d and 0.20d diameter thresholds):

Training Setting (LineMOD on WDR) ADI.10d (%) ↑ ADI.20d (%) ↑
PBR (Physically Based Rendering) 40.96 60.98
Ours (GenerationPose alone) 79.87 94.09
Real (Real annotated data) 83.01 96.43
Real + PBR 83.54 96.84
Real + Ours 85.55 97.27

Evaluating across representative downstream pose estimators demonstrates consistent gains across multiple BOP datasets (ADI.10d metric):

Pose Estimator LM LM-O (Occlusion) YCB-V
GDR-Net (Baseline) 91.00 52.90 60.10
GDR-Net + Ours 92.71 56.20 78.73
WDR (Baseline) 83.01 37.90 27.50
WDR + Ours 85.55 44.76 41.40

Ablation Study

To isolate individual components without train-inference mismatch, separate models were trained and tested from scratch for each ablated variant (macro-averaged over TYOL, YCBV, ICMI, and LM):

Config NCD (%) ↓ BF1@2 ↑ Inlier@2 ↑ Note
Full Model 0.867 0.898 0.901 full model (default strength 0.6)
w/o ControlNetGeom 1.318 0.881 0.887 Sharp geometric degradation (+52.0% NCD error)
w/o mask-aware attention 0.873 0.892 0.896 Boundary agreement drops due to BG interference
w/o dilation (\(r=0\)) 0.886 0.895 0.898 Lack of boundary tolerance impairs contour stability

Sampling strength trade-off between geometric stability and visual diversity:

Strength NCD (%) ↓ BF1@2 ↑ Inlier@2 ↑ Note
0.3 1.085 0.889 0.898 Insufficient denoising; retains raw synthetic artifacts
0.6 (Default) 0.867 0.898 0.901 Optimal sweet spot balancing realism and geometry
1.0 1.045 0.896 0.900 High freedom introduces slight contour drift

Key Findings

  • ControlNetGeom is essential for pose rigidity: Removing ControlNetGeom leads to the largest degradation across all metrics (NCD jumps from 0.867% to 1.318%), proving that layer-wise hints and cross-attention tokens are vital to prevent geometric warping.
  • Pseudo-real data rival real training images: Training WDR purely on generated pseudo-real data reaches 79.87% ADI.10d on LineMOD—nearly doubling the 40.96% achieved by conventional PBR data and approaching the 83.01% of real data without any manual pose labeling.
  • Significant improvements on complex, cluttered datasets: On the challenging YCB-V benchmark, augmenting the state-of-the-art GDR-Net with GenerationPose images boosts ADI.10d from 60.10% to 78.73% (+18.63%), demonstrating that diverse pseudo-real textures and lighting effectively enhance deep visual representations.

Highlights & Insights

  • Decoupled Structural Anchoring and Photorealistic Synthesis: By locking early DiT layers via residual hints from ControlNetGeom while keeping later layers flexible for cross-attention and denoising, the model successfully disentangles rigid geometry constraints from complex photometric synthesis.
  • Topological Boundary Protection via Attention Masking: The FG→BG attention blocking elegantly eliminates background semantic contamination without breaking the parallel compute graph of Transformers or requiring expensive post-hoc composition steps.
  • Robust Zero-Shot Cross-Modal Transfer: Training solely on noisy, incomplete point-cloud renders from uCO3D allows seamless zero-shot transfer to clean CAD mesh renders in BOP, demonstrating strong robustness against rendering domain shifts.

Limitations & Future Work

  • Absence of Explicit Occluder Modeling: The current framework assumes the input render and mask depict the entire visible surface of the target object, and does not explicitly simulate complex inter-object occlusions or dynamic shadows from external manipulators.
  • Challenging Physical Optics: Highly specular, mirror-like reflections or semi-transparent glassware still exhibit minor discrepancies when compared to rigorous physical path tracing.
  • Closed-Loop Task Feedback: Future directions could integrate feedback gradients from downstream pose estimators directly into the diffusion training loop to enable task-driven active data generation.
  • vs CycleGAN / SimGAN: GAN-based refinement lacks explicit 3D geometric inductive biases, frequently suffering from structural collapse, feature hallucination, or boundary drift that invalidates pose annotations. GenerationPose leverages a latent DiT with ControlNet guidance to guarantee strict 6D pose preservation.
  • vs PBR Domain Randomization: Physically-based rendering with domain randomization relies on heuristic asset libraries that still look synthetic and unnatural. GenerationPose harnesses the rich priors of foundation diffusion models to synthesize naturalistic textures, subtle shading, and diverse backgrounds.
  • vs Category-Level Diffusion Augmentation: Unlike category-level methods focusing on coarse shape synthesis, this work addresses high-precision instance-level 6D pose estimation, keeping contour deviations under 1% to reliably train downstream 6D pose estimators.

Rating

  • Novelty: ⭐⭐⭐⭐ [Clever integration of a lightweight geometry ControlNet and FG→BG mask attention blocking into a Diffusion Transformer for pose-preserved generation]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive cross-domain zero-shot evaluation across multiple BOP benchmarks, comprehensive component ablations, and solid downstream pose estimation gains]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation addressing the sim-to-real pose bottleneck, crisp architecture description, and well-structured experimental analysis]
  • Value: ⭐⭐⭐⭐⭐ [Provides a practical, low-cost, and robust pipeline for scaling up high-quality training data in robotics and 6D vision]