Steering 3D Generations: Preference Alignment via Direct Reward and Preference Optimization¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://media.eventhosts.cc/Conferences/ECCV2026/pdfs/12161.pdf
Area: 3D Vision
Keywords: 3D generation, preference alignment, direct reward optimization, direct preference optimization, rectified flow
TL;DR¶
Addressing severe geometric artifacts and intractable human preference formalization in single-image-to-3D generation, this paper proposes a preference alignment framework combining Gaussian-curvature neural SDF data curation with Direct Reward Optimization (DRO) and Direct Preference Optimization (DPO) on rectified flow, achieving high fidelity and controllable geometry via parameter-efficient LoRA.
Background & Motivation¶
Driven by the rapid rise of generative AI across gaming, industrial design, and robotics, synthesizing high-fidelity, controllable 3D assets from a single input image has become a pivotal research pursuit. While recent paradigms have transitioned from 2D lifting methods (such as SDS/VSD distillation and multi-view diffusion) toward native 3D latent diffusion models (e.g., CLAY, Trellis, and Hunyuan3D-2), achieving notable advances in speed and structural diversity, a pronounced gap in quality and utility still lingers between 3D generative models and their 2D counterparts. This disparity primarily stems from the extreme topological complexity of 3D representations (meshes, point clouds, neural fields) and the severe noise—including non-manifold topology, holes, floating artifacts, and inconsistent surface normals—inherent to web-scale raw 3D datasets.
Contemporary 3D generative frameworks trained purely on statistical reconstruction losses (such as latent \(L_2\) regression or Chamfer Distance) treat generation as a black-box mapping without mechanisms to align with human-centric geometric preferences. In practical artistic creation and engineering workflows, users exhibit nuanced, context-dependent quality expectations—such as seeking smooth vehicle body surfaces, sharp mechanical feature edges, and watertight manifolds. Such subtle criteria are notoriously difficult to encode into differentiable point-wise or pixel-wise objective functions. Consequently, unaligned diffusion models often settle into degenerate solutions with jagged noise, over-smoothing, or non-physical surface depressions.
To overcome these dual roadblocks of dirty data and unaligned generation, this work introduces a two-pronged paradigm: establishing a high-quality watertight training foundation followed by explicit preference-aware optimization. Core Idea: Establish a watertight mesh foundation using neural SDF reconstruction with a double-trough Gaussian curvature prior, and fine-tune a rectified flow diffusion backbone using Direct Reward Optimization (DRO) for absolute feedback and Direct Preference Optimization (DPO) for paired data via parameter-efficient LoRA.
Method¶
Overall Architecture¶
Given a single image, the proposed framework synthesizes watertight, high-fidelity 3D meshes adhering to user-centric geometric preferences. The overall architecture is built upon two core pillars: first, a robust data curation pipeline based on neural SDF reconstruction with a double-trough Gaussian curvature prior that extracts clean, watertight meshes from unoriented noisy point clouds and systematically synthesizes curvature-perturbed preference pairs; second, a latent set rectified flow diffusion model where the pre-trained weights are frozen and low-rank adaptation (LoRA) modules are injected into attention projection layers, fine-tuned via an automated rule-based oracle driving Direct Reward Optimization (DRO) alongside paired Direct Preference Optimization (DPO).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Single Image Input & Raw 3D Assets"] --> B["Gaussian Curvature Prior Neural SDF Curation<br/>Double-trough energy separating developable patches & ridges"]
B --> C["Automated Geometric Oracle & Pair Construction<br/>Evaluating watertightness / complexity / curvature variance"]
C --> D["Rectified Flow Backbone & LoRA Fine-Tuning<br/>Preserving pre-trained prior with 1% trainable parameters"]
D --> E["Dual-Track Preference Alignment<br/>Absolute feedback DRO & paired preference DPO"]
E --> F["Output High-Fidelity Watertight Meshes & Controllable Assets"]
Key Designs¶
1. Gaussian Curvature Prior Neural SDF Curation: Decoupling developable patches from topological ridges
Direct surface extraction from public datasets (such as Objaverse-XL, ShapeNet, and ABO) frequently yields non-manifold self-intersections and corrupted normals. This framework adapts the neural SDF formulation of NeurCADRecon, introducing a double-trough geometric energy landscape over Gaussian curvature: $\(\mathcal{L}_{\text{Gauss}} = \frac{1}{|\Omega|} \int_{\Omega} \text{DT}(|k_{\text{Gauss}}(\boldsymbol{x})|) \, \mathrm{d}\boldsymbol{x}\)$ Whereas standard second-order regularizers indiscriminately penalize curvature (\(|K| \to 0\)) and wipe out sharp edges, the double-trough prior \(\text{DT}(\cdot)\) defines two distinct attractors: a developable basin (\(|K| \approx 0\)) for smooth mesh patches and a singularity basin (\(|K| \approx \pi/2\)) for sharp topological creases. A steep penalty barrier separates these basins to suppress ambiguous, noisy curvatures. By varying the regularization coefficient \(\lambda_{\text{Gauss}}\) (e.g., from smooth surfaces at \(\lambda_{\text{Gauss}}=3\) to sharp features at \(\lambda_{\text{Gauss}}=100\)), the curation engine not only guarantees watertight meshes but also generates self-supervised pairs of smooth and sharp variations of the same object for preference training.
2. Automated Geometric Oracle & Direct Reward Optimization (DRO): Absolute quality feedback without pairwise annotations
Gathering large-scale pairwise human preference annotations in 3D is prohibitively expensive and prone to subjective ambiguity. The authors design an automated rule-based oracle \(o(x_0) \in \{0, 1\}\) that programmatically assesses mesh quality along three dimensions: topological watertightness (requiring a closed, manifold surface without holes), structural complexity (filtering degenerate shapes with face counts below 500,000), and surface smoothness (enforcing that the global variance of vertex-wise mean curvature remains strictly below 0.1). A sample meeting all criteria receives \(o(x_0)=1\), and \(0\) otherwise.
Using this binary scalar reward, Direct Reward Optimization (DRO) re-weights the rectified flow diffusion objective: $\(\mathcal{L}_{\text{DRO}} = \mathbb{E}_{t, x_0, \epsilon} \left[ w(t) (1 - 2 o(x_0)) \| \epsilon_\theta(x_t, t, I) - \epsilon \|^2 \right]\)$ The factor \((1 - 2 o(x_0))\) serves as an adaptive gradient modulator. When a sample satisfies the oracle (\(o(x_0)=1\)), the term becomes negative, reinforcing gradient steps that increase the likelihood of generating high-quality geometric latents; conversely, low-quality samples are penalized. Unlike traditional reinforcement learning methods requiring fragile reward model fitting or high-variance policy gradient updates, DRO optimizes directly within the flow-matching objective in a stable, closed-loop fashion.
3. Pairwise Curvature Perturbation & Direct Preference Optimization (DPO): Relative margin-based fine-grained alignment
When explicit relative preferences or stylistic control are desired, the Bradley-Terry preference model is extended to the 3D rectified flow framework. Utilizing curvature-perturbed pairs \((x_0^w, x_0^l)\)—where sharp, well-formed meshes act as the preferred winner \(x_0^w\) and over-smoothed variants act as the loser \(x_0^l\)—the 3D-DPO loss is formulated as: $\(\mathcal{L}_{\text{DPO}} = - \mathbb{E} \left[ \log \sigma \left( \beta \left( \Delta(x_t^l) - \Delta(x_t^w) \right) \right) \right]\)$ Here, the implicit error differential is \(\Delta(x_t) = \|\epsilon - \epsilon_\theta(x_t, t)\|_2^2 - \|\epsilon - \epsilon_{\text{ref}}(x_t, t)\|_2^2\) between the fine-tuned and frozen reference models, scaled by \(\beta=500\). DPO directly optimizes the model parameters to maximize the implicit reward margin between preferred and non-preferred geometries, effectively guiding denoising trajectories toward fine-grained sharp details.
4. Latent Set Rectified Flow & LoRA Fine-Tuning: Parameter-efficient adaptation without catastrophic forgetting
The generative backbone builds upon Hunyuan3D-2, comprising a 3D Shape VAE and a rectified flow transformer operating on unordered latent sets. Rectified flow learns a straight probability path between Gaussian noise and data distributions by predicting a constant velocity field \(v_\theta(x_t, t, c)\), ensuring stable and rapid sampling. To prevent full fine-tuning from overriding the diverse general 3D category priors learned during pre-training, the diffusion weights are frozen. Trainable low-rank adaptation matrices \(\Delta W = BA\) (\(r=64\)) are injected exclusively into the query and value projection layers of the coarse geometry transformer. By optimizing less than 1% of the total network parameters, LoRA drastically reduces training overhead while isolating stylistic and geometric surface control.
Key Experimental Results¶
Main Results¶
Quantitative evaluations are conducted on the unseen Toys4k benchmark. All predicted meshes are normalized to a unit cube and aligned with ground truth shapes using ICP before evaluation. Metrics comprise shape similarity via Chamfer Distance (CD, lower is better), surface detail accuracy via [email protected] (higher is better), and normal map rendering fidelity via PSNR-N (higher is better) and LPIPS-N (lower is better).
| Method | CD ↓ | F-Score ↑ | PSNR-N ↑ | LPIPS-N ↓ |
|---|---|---|---|---|
| InstantMesh | 0.0224 | 62.78 | 27.34 | 0.096 |
| 3DTopia-XL | 0.0228 | 65.90 | 31.87 | 0.090 |
| LN3Diff | 0.0299 | 59.17 | 27.10 | 0.094 |
| CLAY | 0.0124 | 72.59 | 35.35 | 0.035 |
| Trellis | 0.0083 | 73.12 | 36.11 | 0.024 |
| Hunyuan3D-2 (Baseline) | 0.0089 | 72.79 | 35.53 | 0.033 |
| Ours (DRO) | 0.0036 | 77.97 | 37.05 | 0.021 |
Ablation Study¶
The ablation investigates the original Hunyuan3D-2 baseline, standard supervised fine-tuning (SFT) on curated watertight assets, and preference-aligned fine-tuning via DPO and DRO.
| Config | CD ↓ | F-Score ↑ | Note |
|---|---|---|---|
| Baseline | 0.0079 | 73.79 | Pre-trained backbone without curated data fine-tuning |
| Baseline + SFT | 0.0045 | 76.28 | Supervised fine-tuning on curated clean meshes, resolving gross topology errors |
| Baseline + SFT/DPO | 0.0048 | 76.77 | DPO alignment using paired curvature-perturbed samples, sharpening fine details |
| Baseline + SFT/DRO | 0.0036 | 77.97 | Rule-based oracle absolute feedback offering broader solution space exploration |
Key Findings¶
- DRO absolute feedback surpasses paired DPO: DRO delivers the lowest CD (0.0036) and highest F-Score (77.97) on unseen test data. While pairwise DPO captures relative stylistic gradients, artificial curvature perturbation boundaries can introduce inductive bias. In contrast, the rule-based oracle in DRO supplies an objective geometric prior that enables the model to generalize across broader high-quality geometric manifolds.
- Data curation provides a vital foundation: Simply fine-tuning on high-quality neural SDF-curated meshes via SFT slashes CD from 0.0079 to 0.0045 (a 43% error reduction), proving that dirty web meshes constitute a major bottleneck in 3D synthesis.
- Consistent human perceptual preference: In a two-alternative forced-choice (2AFC) blind study with 90 participants evaluating 45 input images, Ours-DRO secured 60.0% of the total votes against the competitive baseline CLAY, receiving praise for surface cleanliness and structural plausibility.
Highlights & Insights¶
- First translation of LLM alignment to native 3D diffusion: Systematically establishes both absolute-reward (DRO) and relative-preference (DPO) pathways within latent rectified flow, moving beyond point-wise statistical regression.
- Double-trough Gaussian curvature data curation: Employs differential geometry invariants to cleanly separate developable patches from sharp ridges, providing both watertight training anchors and automated contrastive pairs.
- Lightweight, non-destructive LoRA control: Reallocates surface styles and geometric cleanliness by tuning less than 1% of transformer parameters, entirely preserving the pre-trained model's multi-category semantic generalization.
Limitations & Future Work¶
- Heuristic nature of the geometric oracle: Current DRO quality classification relies on fixed criteria (watertightness, face counts, curvature variance) rather than multi-view multimodal aesthetic judges (such as VLM-as-a-judge), limiting perceptual discrimination on complex artistic assets.
- Single-object focus: The neural SDF curation and latent set representations are engineered primarily for isolated objects, making extensions to dense, multi-object indoor/outdoor scenes challenging.
- Sensitivity of neural SDF to extreme sparsity: When raw point clouds exhibit severe occlusion or missing boundaries, neural SDF fitting can become computationally heavy or suffer from local topological collapse.
Related Work & Insights¶
- vs CLAY / Trellis / Hunyuan3D-2: Prior state-of-the-art 3D diffusion models focus primarily on scaling latent space architectures with unguided reconstruction losses. This framework introduces an alignment stage that substantially removes bumpy surface noise and topological defects under identical inference latency.
- vs 2D Diffusion-DPO: 2D alignment largely assesses image aesthetics and text-image alignment through subjective human ratings or VLM scoring. This work grounds preference alignment in rigorous 3D differential geometric properties like manifold watertightness and mean curvature variance.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Pioneers the systematic application of DRO and DPO to 3D rectified flow diffusion models, coupled with a double-trough curvature prior.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation encompassing quantitative geometric metrics, dual ablation setups, curvature controllability demonstrations, and a 2AFC blind human user study.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear narrative structure, sound mathematical formulation, and coherent alignment between motivation, methodology, and empirical evidence.
- Value: ⭐⭐⭐⭐⭐ Offers a practical, standardized paradigm for resolving non-watertight surfaces and geometric artifacts in industrial-grade 3D generative pipelines.