Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: 3D Vision
Keywords: 3D asset generation, super-resolution, progressive refinement, localized diffusion, sparse voxel latent
TL;DR¶
Addressing the cubic memory and compute bottleneck of full-object 3D diffusion generation, PLSR introduces a progressive and localized super-resolution framework on structured sparse voxel latents, decomposing global refinement into aligned patch triplets, lightweight low-resolution cross-attention LoRA fine-tuning, and per-step iterative denoising stitching to super-resolve coarse meshes up to \(2560^3\) resolution on a single GPU.
Background & Motivation¶
High-resolution 3D asset generation forms the bedrock of modern video games, cinematic visual effects, and industrial design. In recent years, native 3D diffusion and flow-matching foundation models (such as Craftsman, Hunyuan3D, and TRELLIS.2) have made significant strides by operating on compact structured sparse voxel latents, enabling direct feed-forward generation of meshes from single-view images. However, because 3D voxel grid representation and attention computations scale cubically (\(\mathcal{O}(N^3)\)) with spatial resolution, existing foundation models are predominantly pre-trained at moderate fixed latent resolutions (such as \(64^3\), corresponding to \(1024^3\) physical voxels). When production-grade fine surface details are required, naively scaling global voxel grids causes severe GPU out-of-memory errors, while training-free test-time upscaling suffers from the generalization bounds of the base model, yielding oversmoothed surfaces or distorted geometries.
The core tension lies in the mismatch between user demand for arbitrarily fine geometric micro-structures and the prohibitive compute and memory costs of whole-object high-resolution training and inference. Conversely, directly borrowing 2D patch-based super-resolution strategies into 3D space inevitably introduces severe boundary seam artifacts, neighbor inconsistencies, and cross-view occlusion ambiguities. Furthermore, prior geometric super-resolution paradigms typically require multi-scale paired high-resolution datasets for end-to-end retraining, which is exceedingly expensive given the scarcity of clean, high-polygon 3D data.
This paper's angle of attack is to decouple global macro-topology from localized high-frequency details: fine-grained surface geometry and appearance are predominantly determined by local neighborhood structures and corresponding close-up image textures, independent of the object's absolute spatial extent. Core idea: decompose the full-object 3D super-resolution task into localized latent patch flow-matching denoising sub-tasks, construct associatively aligned local triplets with lightweight low-resolution cross-attention for low-cost fine-tuning, and unify them in an iterative patch-wise denoising and center-weighted stitching pipeline that cascades progressively to extreme resolutions.
Method¶
Overall Architecture¶
PLSR takes as input a coarse mesh \(O_l\) generated by a pre-trained 3D generator (or the output of a preceding super-resolution stage) and one or more reference images \(I\) with known or estimated camera parameters, producing a high-resolution detailed mesh \(O_h\). The entire pipeline operates in the structured sparse latent space of TRELLIS.2 and consists of three stages: associative input decomposition, localized latent flow-matching denoising, and iterative patch-wise denoising with progressive cascading.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Coarse Mesh Ol + Reference Image I"] --> B["Associative Input Decomposition<br/>3D Voxel Slicing + Camera-Projected 2D Crops"]
B --> C["Localized Latent Diffusion Model<br/>Lightweight LR Attention + LoRA Fine-Tuning"]
C --> D["Iterative Patch-wise Denoising & Stitching<br/>Per-Step Center-Weighted Velocity Aggregation"]
D --> E["Progressive Cascading<br/>Current HR Latent Fed as Next Stage LR Input"]
E --> F["Output: High-Resolution Fine Mesh Oh"]
Key Designs¶
1. Associative Input Decomposition: Triplet Construction via Relative Coordinates and Camera Projection To circumvent the prohibitive memory overhead of global inputs and decouple resolution scaling from absolute global object positions, PLSR introduces associative input decomposition to partition the global task into localized sub-task triplets \((P_{h,t}, P_l, I_p)\). The high-resolution sparse voxel volume is partitioned into regular local 3D patches (e.g., \(32^3\) windows), while the corresponding low-resolution latent patch \(P_l\) is extracted from the coarse latent grid. Crucially, active voxel coordinates within \(P_{h,t}\) and \(P_l\) are strictly re-anchored relative to each patch's local bounding box origin, preventing the network from overfitting to global positional cues. For the conditioning image, the 3D patch bounding box is projected onto the reference image plane using camera parameters, yielding a tightly cropped close-up patch \(I_p\) from which localized DINOv3 visual features \(F_{I_p}\) are extracted. When multiple views are available, an occlusion-aware visibility score is derived via voxel rendering and depth testing; views are sampled probabilistically weighted by visibility during denoising, while patches occluded across all viewpoints have their visual feature tokens zeroed out to fall back safely on 3D geometric priors.
2. Localized Latent Diffusion Model: Lightweight LR Attention and LoRA Adaptation Because coarse mesh remeshing and simplification induce subtle coordinate shifts between low-resolution and high-resolution voxels, standard channel concatenation fails due to strict spatial misalignment. Rather than executing dense global cross-attention across full volumes, PLSR introduces a lightweight low-resolution conditioning branch on top of the frozen TRELLIS.2 Flow DiT. The low-resolution patch \(P_l\) is first passed through the initial two convolutional residual blocks \(\mathcal{D}_2\) of the pre-trained VAE decoder to double its spatial resolution, matching the dimensions of \(P_{h,t}\). Next, a lightweight LR attention layer calculates cross-attention strictly within a compact local radius between HR voxels and adjacent LR cells, adaptively resolving spatial discrepancies with minimal computational overhead. To maintain computational efficiency and preserve pre-trained priors, the LR attention layer is inserted only into a subset of DiT blocks and reused across consecutive blocks with affine layer normalization, while LoRA adapters are attached to all linear layers outside the input/output projections. The model is trained using the conditional flow matching objective: $$ \mathcal{L}{\mathrm{CFM}}(\theta) = \mathbb{E}}, \epsilon} \left| v_\theta\left(P_h(t), \mathcal{D2(P_l), F) \right|_2^2 $$ This allows the network to learn rich micro-structure synthesis entirely from diverse patch samples drawn from moderate-resolution data, bypassing the need for paired high-resolution full-object training sets.}, t\right) - (\epsilon - P_{h,0
3. Iterative Patch-wise Denoising & Stitching: In-Loop Synchronization and Center-Weighted Blending To eliminate boundary cracks and stitching artifacts characteristic of naive patch-based inference, PLSR rejects one-shot spatial blending on the final generated meshes and instead integrates patch decomposition and stitching inside each discrete denoising step \(\Delta t\). At timestep \(t\), the global noisy latent \(Z_{h,t}\) is partitioned into partially overlapping windows. The localized flow denoiser independently predicts local velocity fields for all patches. In overlapping boundaries, per-patch velocity predictions are aggregated via spatial center-weighted averaging, assigning high confidence to patch cores and linearly decaying weights toward edges where context is incomplete. The merged global velocity field \(v_{\mathrm{pred}}\) updates the shared global latent: $$ Z_{h, t - \Delta t} = Z_{h,t} + \Delta t \cdot v_{\mathrm{pred}} $$ Because the entire high-resolution latent is re-synchronized into a globally coherent state before entering the subsequent denoising iteration, neighboring patches continuously maintain mutual spatial awareness, fundamentally suppressing seams at their source.
4. Progressive Cascading: Seamless Arbitrary Resolution Scaling Because the localized super-resolution model is trained exclusively on local relative scale transforms rather than absolute dimensions, its learned refinement prior generalizes across different resolution levels. In deployment, PLSR adopts a progressive multi-round cascading workflow: starting from a coarse latent mesh at \(640^3\) resolution, the first pass refines it to \(1280^3\); this output latent is then directly treated as the low-resolution condition \(Z_l\) for a second pass, yielding an extreme resolution of \(2560^3\). This multi-stage hierarchical refinement avoids the degradation seen when models are forced into extreme single-step extrapolation, ensuring fine details emerge coherently while keeping peak memory strictly bounded on a single consumer GPU.
Loss & Training¶
The model is trained on 30K 3D assets selected from Objaverse-XL with aesthetic scores \(\ge 6.5\). Meshes are encoded into \(64^3\) latents using the TRELLIS.2 VAE and downsampled to \(32^3\) pairs. Patches are sampled at \(32^3\) for HR and \(16^3\) for LR, alongside \(1024^2\) image crops. Alongside \(\mathcal{L}_{\mathrm{CFM}}\), a minor regularization loss (\(\lambda = 0.005\)) prevents degeneration of the image-modulation pathway. The model is optimized using AdamW (learning rate \(1\times 10^{-5}\), weight decay 0.1, EMA 0.999) with color jitter, geometric warping, and patch dropout augmentations. Both geometry and texture SR models require only 2 epochs (approximately 50 GPU hours on a single NVIDIA A100 80GB), demonstrating exceptional resource efficiency.
Key Experimental Results¶
Main Results¶
The method is evaluated on two complementary benchmarks: the Sketchfab Pseudo-GT benchmark comprising 90 high-detail assets (evaluated with 2D render metrics and 3D geometric metrics), and an Open-World benchmark featuring 70 complex objects (evaluated via GPT-5.2 pairwise ELO scoring). Competing baselines include TRELLIS.2 (at default \(1536^3\) cascading and direct \(2560^3\) upscaling), UltraShape, DetailGen3D, and MVPaint UVR.
Table 1: Objective pseudo-GT evaluation on Sketchfab assets (CD scaled by \(10^3\), HFR scaled by \(10^2\))
| Method | Geom. LPIPS ↓ | Geom. FID ↓ | Geom. HFR ↑ | Geom. CD ↓ | Geom. F1 ↑ | Tex. LPIPS ↓ | Tex. FID ↓ | Tex. HFR ↑ |
|---|---|---|---|---|---|---|---|---|
| TRELLIS.2 (1536) | 0.122 | 60.37 | 0.981 | 0.0247 | 0.059 | 0.150 | 70.20 | 1.501 |
| TRELLIS.2 (2560) | 0.122 | 63.19 | 0.923 | 0.0225 | 0.036 | 0.158 | 72.77 | 1.482 |
| UltraShape | 0.167 | 72.03 | 0.876 | 1.3982 | 0.109 | - | - | - |
| DetailGen3D | 0.187 | 93.35 | 0.645 | 0.4600 | 0.083 | - | - | - |
| MVPaint UVR | - | - | - | - | - | 0.154 | 77.18 | 1.052 |
| Ours (2560) | 0.109 | 55.89 | 1.088 | 0.0139 | 0.145 | 0.125 | 59.89 | 1.646 |
Table 2: Open-world evaluation via GPT-5.2 pairwise ELO ratings (baseline 1500)
| Method | Geom. Detail Rich ↑ | Geom. Align w/ Ref. ↑ | Geom. Artifact Free ↑ | Geom. Overall Quality ↑ | Tex. Detail Rich ↑ | Tex. Align w/ Ref. ↑ | Tex. Artifact Free ↑ | Tex. Overall Quality ↑ |
|---|---|---|---|---|---|---|---|---|
| TRELLIS.2 (1536) | 1775 | 1774 | 1736 | 1790 | 1553 | 1621 | 1624 | 1599 |
| TRELLIS.2 (2560) | 1627 | 1592 | 1492 | 1572 | 1494 | 1426 | 1441 | 1416 |
| UltraShape / MVPaint | 1246 | 1385 | 1551 | 1356 | 1197 | 1369 | 1358 | 1350 |
| DetailGen3D | 869 | 927 | 1111 | 925 | - | - | - | - |
| Ours (2560) | 1982 | 1821 | 1608 | 1855 | 1756 | 1582 | 1575 | 1633 |
Ablation Study¶
The paper conducts an ablation study using the GPT-5.2 evaluation suite across three key configurations: (1) running the unmodified TRELLIS.2 base model under patch-wise inference without localized adaptation (w/o model); (2) removing window overlap during patch inference (w/o overlap); and (3) performing single-stage non-progressive super-resolution (w/o prog.).
Table 3: Quantitative ablation study (GPT-5.2 ELO ratings)
| Configuration | Geom. Detail Rich ↑ | Geom. Align w/ Ref. ↑ | Geom. Artifact Free ↑ | Geom. Overall Quality ↑ | Tex. Detail Rich ↑ | Tex. Overall Quality ↑ | Note |
|---|---|---|---|---|---|---|---|
| w/o model | 1291 | 1281 | 1209 | 1281 | 1288 | 1271 | Without localized fine-tuning, base model collapses on local patches |
| w/o overlap | 1611 | 1559 | 1550 | 1574 | 1624 | 1609 | Discarding overlaps triggers boundary seams and discontinuous steps |
| w/o prog. | 1424 | 1552 | 1665 | 1537 | 1441 | 1475 | Single-step scaling fails to synthesize multi-scale geometric subtleties |
| Full Model | 1672 | 1606 | 1575 | 1605 | 1645 | 1643 | Best overall balance between detail generation and artifact suppression |
Key Findings¶
- Failure of Direct Base Model Extrapolation: Directly evaluating TRELLIS.2 at \(2560^3\) causes geometric FID to worsen from 60.37 to 63.19 and overall ELO to drop from 1790 to 1572. This confirms that pre-trained foundation models cannot extrapolate beyond their native resolution, whereas PLSR's progressive localized approach achieves an ELO of 1855, establishing the power of scale-invariant local priors.
- In-Loop Stitching Eradicates Seams: Both qualitative figures and ablation tables demonstrate that non-overlapping patches leave severe visible seams. By embedding weighted averaging within each iterative velocity update, seam artifacts are completely eliminated while maintaining crisp geometric ridges.
- Multi-View Coverage Benefits: Increasing conditioning camera views from 1 to 4 reduces normal/color LPIPS from 0.109 / 0.125 to 0.103 / 0.119, demonstrating that visibility-weighted multi-view sampling effectively disambiguates occluded and peripheral regions.
Highlights & Insights¶
- Decoupling Global Scale from Local Detail: Formulating 3D super-resolution as relative patch-wise flow matching enables models trained on moderate datasets to generalize smoothly to extreme resolutions.
- Denoising-Synchronized Patch Blending: Integrating weighted velocity aggregation inside each flow ODE step rather than post-hoc mesh stitching ensures continuous latent alignment across adjacent blocks.
- Single-GPU Low-Cost Adaptation: By inserting lightweight cross-attention and LoRA adapters, a 4B parameter foundation model was transformed into a high-resolution detail engine in just 50 GPU hours.
Limitations & Future Work¶
- Dependence on Initial Macro-Topology: As a super-resolution framework, the model does not correct severe topological flaws or structural missing parts in the coarse input mesh.
- Sensitivity to Camera Pose and View Quality: Cropped 2D image condition extraction relies on camera parameters predicted by MapAnything; substantial pose estimation errors produce texture misalignments.
- Representation Specificity: The architecture is tailored to the sparse voxel latent structure of TRELLIS.2; extending to unconstrained NeRF or 3D Gaussian Splatting latents requires re-engineering the patch discretization mechanism.
Related Work & Insights¶
- vs TRELLIS.2: While TRELLIS.2 relies on full-object cascading that hits memory ceilings and exhibits artifacts at \(2560^3\), PLSR uses localized patch diffusion to deliver sharp, seam-free geometry and texture at \(\times 2.5\) higher resolution on a single GPU.
- vs SDF-Diffusion / SuperCarver: SDF-Diffusion is constrained by memory-intensive dense grids, while SuperCarver requires computationally heavy multi-view inverse rendering. PLSR generates geometry and PBR textures concurrently in a compact flow-matching latent space with much faster inference.
- vs UltraShape / LATTICE: UltraShape requires expensive full retraining on large-scale high-resolution datasets. In contrast, PLSR provides a lightweight, plug-and-play adapter framework on frozen foundation models.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Novel localized latent flow formulation with synchronized in-loop patch aggregation]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Exhaustive pseudo-GT metrics, GPT-5.2 ELO blind evaluations, and thorough ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, well-structured pipeline narrative, and honest limitation disclosures]
- Value: ⭐⭐⭐⭐☆ [Provides an accessible, single-GPU practical paradigm for industrial-grade high-resolution 3D asset generation]