Skip to content

DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation

Conference: ECCV 2026
Paper: ECCV Official
Project: https://ejshim.github.io/diffgi/
Area: 3D Vision
Keywords: Geometry Images, Differentiable Marching Squares, Thin-Shell 3D Generation, Truncated Signed Distance Function, Latent Diffusion Models

TL;DR

DiffGI integrates a continuous 2D Truncated Signed Distance Function (TSDF) into geometry images with an analytical Differentiable Marching Squares algorithm and geometry-aware normal rendering loss, achieving high-fidelity, subpixel-boundary thin-shell and non-manifold 3D surface generation within an ultra-compact \(32 \times 32 \times 4\) latent space.

Background & Motivation

Contemporary 3D generative models predominantly build upon implicit volumetric field representations such as Signed Distance Fields (SDFs), volumetric occupancy grids, and Neural Radiance Fields (NeRFs). By learning continuous 3D spatial functions, these methods excel at generating watertight surfaces and smooth manifolds. However, for thin-shell and non-manifold open-boundary geometries ubiquitous in digital garments and furniture frames, volumetric implicit fields inherently enforce closed watertight topology, which frequently induces artificial thickness, volume bloat, or front-back surface blending. Furthermore, meshes extracted via Marching Cubes naturally lack UV coordinates, imposing severe friction when integrating into downstream industrial pipelines that demand cloth physics simulation and material authoring.

Multi-chart geometry images offer an attractive surface-centric alternative by parameterizing 3D surfaces onto regular 2D UV grids, treating coordinates and geometric attributes as image-like tensors that directly interface with mature 2D generative backbones (e.g., Omages, GIMDiffusion, GarmageNet). Nevertheless, existing geometry-image approaches rely on discrete binary occupancy maps to demarcate valid surface charts. Because binary masks define boundaries via discontinuous step functions, boundary localization becomes strictly resolution-dependent; downsampling to practical training resolutions inevitably causes staircase artifacts and destroys fine contour details. Crucially, post-processing steps like boundary snapping remain non-differentiable heuristics, structurally severing the learning loop between 2D latent representation and 3D surface optimization.

Addressing this fundamental disconnect, this work asks: why restrict differentiable iso-surface extraction to memory-intensive \(O(N^3)\) 3D voxel grids when it can be realized directly on 2D UV charts? Core idea: replace discrete binary occupancy with a continuous 2D Truncated Signed Distance Function (TSDF) to preserve subpixel boundary precision under downsampling, and devise an analytical Differentiable Marching Squares (DMS) operator paired with geometry-aware normal rendering supervision for end-to-end thin-shell mesh generation in an ultra-compact 2D latent space.

Method

Overall Architecture

The DiffGI framework establishes an end-to-end differentiable mapping: 3D Mesh ↔ 2D TSDF Geometry Image ↔ Compact Latent Space. In the offline preprocessing stage, input 3D meshes are parameterized into a \(256 \times 256 \times 4\) continuous geometry image tensor (a 3-channel 3D coordinate position map and a 1-channel continuous 2D TSDF) via AABB packing, barycentric surface sampling, boundary dilation, and 2D signed distance transformation. The DiffGI-VAE compresses this representation into a \(32 \times 32 \times 4\) latent space. During decoding and surface recovery, the Differentiable Marching Squares (DMS) module analytically computes boundary vertex positions from the predicted TSDF and position maps, allowing geometric losses from differentiable normal rendering to backpropagate seamlessly into the 2D network.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input 3D Thin-Shell Mesh"] --> B["Subpixel 2D TSDF Representation<br/>AABB UV packing + continuous 2D distance field"]
    B --> C["DiffGI-VAE Latent Compression<br/>SD 1.5 weight expansion + 32×32 compact latents"]
    C --> D["Differentiable Marching Squares Reconstruction<br/>Analytical linear interpolation + zero-crossing topology"]
    D --> E["Geometry-Aware Normal Rendering Loss<br/>nvdiffrast rasterization + 3D surface backpropagation"]
    E --> F["Latent Diffusion Downstream Generation<br/>DiT backbone + Flow-Matching multi-condition generation"]

Key Designs

1. Subpixel 2D TSDF Representation: Eliminating Resolution Sensitivity and Staircase Artifacts

Addressing the severe staircase artifacts and boundary destruction that binary occupancy maps suffer under downsampling, this design introduces a continuous 2D Truncated Signed Distance Function (TSDF) on the UV plane. The pipeline first packs UV charts with uniform global scaling and inter-patch padding using an AABB-based packing algorithm. Next, it samples 3D surface coordinates onto a \(1024 \times 1024\) grid via barycentric interpolation, applying edge dilation to prevent boundary bleeding. For every pixel on the grid, the signed Euclidean distance to the nearest chart contour is computed (positive inside, negative outside) and clamped at a truncation threshold of 15 pixels. When bilinearly downsampled to \(256 \times 256 \times 4\), the continuous TSDF retains subpixel zero-crossing locations, allowing precise recovery of open fabric boundaries without staircase aliasing.

2. Differentiable Marching Squares Reconstruction: Bridging 3D Geometric Loss to 2D Tensors

Addressing the discrete lookup table in classical Marching Squares that halts gradient backpropagation, this design derives an analytical linear interpolation formulation for continuous vertex extraction. On the 2D TSDF grid, when two adjacent pixel values \(\phi_A\) and \(\phi_B\) have opposite signs (\(\phi_A \cdot \phi_B < 0\)), the Intermediate Value Theorem ensures a zero-crossing boundary vertex on the connecting edge. To prevent gradient explosion when \(\phi_A \approx \phi_B\), a continuous interpolation formula with a regularization constant \(\epsilon = 10^{-5}\) is employed:

\[x = \frac{\phi_A}{\phi_A - \phi_B + \epsilon \cdot \operatorname{sgn}(\phi_B)}\]

Because \(\operatorname{sgn}(\cdot)\) is locally constant and treated as a non-differentiable constant during backpropagation, gradients propagate smoothly through the ratio \(\phi_A / (\phi_A - \phi_B)\) back to the 2D TSDF and position maps via the chain rule \(\frac{\partial \mathcal{L}}{\partial V}\). For topologically ambiguous saddle configurations (Case 6 and Case 9), the algorithm deterministically adopts the convention of treating the patches as separate independent regions. Implemented entirely with vectorized PyTorch tensor operations, DMS operates at \(O(N^2)\) computational complexity, scaling far more efficiently than \(O(N^3)\) 3D volumetric iso-surface extraction methods.

3. Geometry-Aware Normal Rendering Loss: Enforcing High-Frequency Curvature and Creases

Addressing the inadequacy of pixel-level L1 reconstruction losses in constraining surface curvature and sharp creasing—which frequently leads to washed-out fabric wrinkles—this design incorporates a differentiable surface rendering supervision loop. The mesh \(\hat{M}\) extracted differentiably via DMS is rendered into normal map images \(R(\hat{M})\) from four viewpoints using the high-performance differentiable rasterizer nvdiffrast, and compared directly against ground-truth rendered normal maps \(R(M^*)\):

\[\mathcal{L}_{\text{Normal}} = \|R(\hat{M}) - R(M^*)\|_1\]

This provides direct, dense gradient signals that guide the VAE encoder to prioritize high-frequency surface orientation and sharp boundary profiles during latent compression.

4. Latent Diffusion Downstream Generation: Decoupling Global Topology and Local Silhouettes

Addressing the irregular spatial layout of multi-chart geometry images where standard convolutional inductive biases struggle to capture cross-chart global topology, this design adopts a Diffusion Transformer (DiT) backbone with flow-matching scheduling on the \(32 \times 32 \times 4\) latent space. For label-conditioned generation, a DiT-B/2 model captures global shape variation. For single-view image-to-3D generation, semantic visual tokens extracted by DINOv2-Large are injected into DiT-L/2 via cross-attention to infer occluded back surfaces. For 2D pattern occupancy-conditioned draping, a lightweight UNet-Tiny architecture with skip connections is adopted to prioritize local silhouette correspondence. During training, random 90-degree rotations and re-packing augmentations at the chart level effectively eliminate fixed-layout positional bias.

Loss & Training

The overall training objective of DiffGI-VAE balances pixel-level reconstruction, distance field alignment, normal rendering supervision, and KL divergence regularization:

\[\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{Pos}} + \lambda_{\text{TSDF}}\mathcal{L}_{\text{TSDF}} + \lambda_{\text{Normal}}\mathcal{L}_{\text{Normal}} + \lambda_{\text{KL}}\mathcal{L}_{\text{KL}}\]

where \(\mathcal{L}_{\text{Pos}}\) and \(\mathcal{L}_{\text{TSDF}}\) are pixel-level L1 losses. The VAE backbone is initialized from pretrained Stable Diffusion 1.5 VAE weights, expanding the initial convolutional layer to 4 channels with zero-initialized extra weights to retain robust natural image spatial compression priors without catastrophic forgetting.

Key Experimental Results

Main Results

Quantitative evaluations on the commercial furniture benchmark ABO and physics-ready garment dataset GarmageSet compare DiffGI-VAE against representative geometry-image baselines Omages and GarmageNet. Metrics include Chamfer Distance (CD), Earth Mover's Distance (EMD), Jensen-Shannon Divergence (JSD), and Normal Consistency (NC):

Method Rep. Size ABO CD (×10⁻³) ↓ ABO EMD ↓ ABO JSD (×10⁻³) ↓ ABO NC ↑ GarmageSet CD (×10⁻³) ↓ GarmageSet EMD ↓ GarmageSet JSD (×10⁻³) ↓ GarmageSet NC ↑
Omages 64×64×4 0.89 0.25 0.92 0.89 1.31 0.17 1.79 0.95
GarmageNet (Tess.) N×72 - - - - 2.19 0.21 32.61 0.88
GarmageNet (Official) N×72 - - - - 1.89 0.17 5.51 0.90
Ours (DiffGI-VAE) 32×32×4 0.83 0.23 0.89 0.83 0.46 0.16 1.24 0.96

On the downstream single-view image-to-3D task on GarmageSet, DiffGI is evaluated against general-purpose 3D foundation models (TRELLIS, TRELLIS.2) and GarmageNet across CD, F1-score (threshold 0.01), Hausdorff Distance (HD), and Boundary Chamfer Distance (BCD):

Method #Vert. ↓ CD (×10⁻²) ↓ F1-Score ↑ HD (×10⁻²) ↓ BCD (×10⁻²) ↓
TRELLIS 109K 3.44 ± 6.97 0.28 ± 0.14 15.38 ± 27.57 N/A
TRELLIS.2 (512²) 380K 11.01 ± 7.64 0.27 ± 0.14 69.70 ± 50.00 12.44 ± 8.56
GarmageNet 526K 4.31 ± 3.03 0.20 ± 0.12 23.76 ± 11.12 5.64 ± 2.94
Ours (DiffGI) 23K 1.35 ± 0.47 0.48 ± 0.12 8.42 ± 3.25 2.91 ± 0.83

Ablation Study

The ablation study on GarmageSet systematically decomposes the contributions of the continuous TSDF representation versus binary occupancy (Occ.) and the geometry-aware normal rendering loss (\(\mathcal{L}_{\text{Normal}}\)). For fair comparison, occupancy variants are reconstructed with a differentiable version of Omages-style tessellation:

Rep. \(\mathcal{L}_{\text{Normal}}\) CD (×10⁻³) ↓ EMD ↓ JSD (×10⁻³) ↓ NC ↑
Occ. × 1.503 0.171 4.539 0.906
Occ. ✓ 1.313 0.166 4.257 0.947
TSDF × 0.595 0.165 2.169 0.921
TSDF (Ours) ✓ 0.461 0.160 1.244 0.961

Inference efficiency benchmarks demonstrate the lightweight footprint and fast generation speed of DiffGI:

Method & Task Hardware Peak VRAM (GB) ↓ Time (sec) ↓
TRELLIS-image RTX A6000 Ada 16.28 4.52
TRELLIS.2 (512²) RTX A6000 Ada 2.65 12.35
TRELLIS.2 (1024²) RTX A6000 Ada 6.41 45.14
Omages RTX A6000 Ada 2.49 52.00
GarmageNet RTX A6000 Ada 0.96 0.40
Ours-Label RTX A6000 Ada 1.18 0.50
Ours-Image RTX A6000 Ada 3.22 0.80
Ours-Image RTX 4070 (12GB) 3.22 1.21
Ours-Image MacBook M4 (CPU) - 8.52

Key Findings

  • TSDF representation is the fundamental driver of geometric accuracy: merely switching from binary occupancy to TSDF (without normal loss) cuts CD by over 60% (from \(1.503 \times 10^{-3}\) to \(0.595 \times 10^{-3}\)) and reduces JSD by more than 52%, proving that subpixel boundary encoding prevents downsampling degradation.
  • Normal rendering loss provides critical curvature regularization: adding \(\mathcal{L}_{\text{Normal}}\) on top of TSDF pushes NC to 0.961 while further lowering CD to \(0.461 \times 10^{-3}\), eliminating high-frequency surface ripple artifacts.
  • Foundation models suffer structural distortion on thin surfaces: volumetric foundation models (TRELLIS, TRELLIS.2) enforce watertightness and produce excessive thickness, resulting in high Boundary Chamfer Distance (\(12.44 \times 10^{-2}\)), whereas DiffGI achieves \(2.91 \times 10^{-2}\) BCD with only 23K vertices and a superior F1-score of 0.48.
  • Extreme computational efficiency: while Omages takes 52 seconds and volumetric models demand over 16 GB VRAM, DiffGI generates conditional 3D meshes in ~1.21 s on a consumer-grade RTX 4070 (3.22 GB VRAM) and runs in 8.52 s purely on an Apple M4 CPU.

Highlights & Insights

  • Dimensionality reduction of differentiable iso-surfacing: avoids the cubic \(O(N^3)\) computational and memory explosion of 3D Marching Cubes and DMTet by operating entirely on 2D UV charts with \(O(N^2)\) tensor operations.
  • Subpixel boundary encoding unlocks ultra-compact latents: overcomes the resolution bottleneck of geometry images, proving that continuous distance fields can preserve intricate open contours even when compressed into a \(32 \times 32\) latent grid.
  • High practical utility for industrial graphics: directly generates meshes with native UV parameterization, ready for physics simulation and texture authoring without non-differentiable remeshing post-processing.

Limitations & Future Work

  • Sharp mechanical edge rounding: linear interpolation in Marching Squares can cause minor rounding over extremely sharp mechanical angles and corners.
  • Cross-chart seam discontinuities: because UV charts are extracted independently and DMS separates saddle topologies, slight boundary discontinuities can appear along adjacent patch seams, posing challenges for airtight physical simulations.
  • Absence of texture and material generation: the current framework focuses solely on 3D geometry and normal generation, leaving RGB texture and PBR material synthesis to future work.
  • vs Omages: Omages operates on uncompressed \(64 \times 64\) geometry images with binary occupancy and non-differentiable tessellation, requiring 52 s per sample; DiffGI leverages continuous TSDF and differentiable Marching Squares within a compressed \(32 \times 32\) latent space, achieving higher boundary fidelity and sub-second generation.
  • vs GarmageNet: GarmageNet relies on fixed per-panel vectors with heuristic multi-stage extraction (~5.1 s/mesh) and fails to generalize to open furniture objects; DiffGI provides a unified 2D multi-chart representation supporting both garments and general open-surface meshes.
  • vs TRELLIS / DMTet: Volumetric representations enforce watertight topology, inflating thin shells into bloated volumes with hundreds of thousands of vertices; DiffGI retains a surface-centric paradigm, generating clean open boundaries with an order of magnitude fewer vertices.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering realization of continuous 2D TSDF and differentiable Marching Squares for multi-chart geometry images.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks spanning furniture and garments, rigorous ablations, efficiency analysis, and multiple conditioning modes.
  • Writing Quality: ⭐⭐⭐⭐⭐ Well-structured mathematical formulation, lucid problem articulation, and crisp empirical evidence.
  • Value: ⭐⭐⭐⭐⭐ Directly addresses the long-standing artificial thickness and staircase bottlenecks in thin-shell 3D generation with production-ready efficiency.