DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding¶
Conference: NeurIPS2026
arXiv: 2609.34651
Area: 3D Vision
Keywords: UV texture generation, surface-aware positional encoding, multi-level attention, image conditioning, rectified flow
TL;DR¶
DirectUV generates mesh textures directly in the latent UV space of a frozen Flux VAE, embeds 3D surface coordinates into each attention head's rotary positional encoding, and uses the reference image and coarse UV only as conditions, achieving 25.13 PSNR and 0.9527 SSIM on GSO while improving consistency across UV islands and texture detail.
Background & Motivation¶
A common approach to texturing an existing 3D mesh is to generate multiple views with a 2D diffusion model and then project and bake their colors into a UV map. This approach benefits from strong image priors, but no view observes the entire surface, grazing-angle regions can become stretched, and different views may disagree about materials and illumination. Projection must reconcile these inconsistent observations, often leaving fragmented or blurry textures in occluded interiors, at seams, and on unseen backsides.
Direct UV generation avoids treating multi-view images as the final source of texture values, but another problem remains: UV unwrapping is a parameterization, not the object's actual adjacency structure. A continuous pattern crossing a mesh seam may occupy two distant UV islands, while adjacent UV pixels may represent unrelated surface regions. Methods such as TEXGen supplement the model with point cloud attention, but if the backbone's positional encoding still comes from the 2D UV grid, its spatial prior remains misaligned with the target surface.
DirectUV changes the positional encoding that shapes queryโkey interactions rather than adding another geometry branch. It also treats projected coarse UV as a layout cue rather than an initialization that must be preserved, allowing denoising to correct erroneous colors. Core Idea: replace UV-grid coordinates with 3D surface coordinates and assign different surface granularities to different attention heads, establishing cross-seam relationships and local detail within a single latent UV generation process.
Method¶
Overall Architecture¶
The inputs are one reference image and a mesh with an existing UV parameterization; the output is a complete UV texture that can be wrapped onto that mesh, not new geometry. During training, a frozen Flux VAE encodes the ground-truth UV texture, noise is added, and a DiT learns the rectified flow direction. Inference starts from Gaussian noise and ends with decoding through the same frozen VAE.
The method combines conditioning pathways, surface-aware positional encoding, and multi-level head allocation. The reference image supplies global appearance and local image features, while coarse UV supplies position-aligned layout cues. A Canonical Coordinate Map (CCM), precomputed from the mesh, supplies geometry only to positional encoding, not to the appearance conditioning bundle. SAPE and multi-level head allocation in the diagram both operate inside the same DiT attention layers; they are not separate serial generators.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Reference image + coarse UV"] --> C["Conditioning pathways"]
N["Noisy UV latent"] --> C
C -->|UV tokens and image conditions| S["Surface-aware positional encoding"]
G["Mesh โ CCM"] -->|Geometry coordinates, not appearance conditions| S
S -->|Position configuration inside attention| M["Multi-level head allocation"]
M --> D["DiT flow prediction"]
D -->|Inference: 28 Euler steps and VAE decoding| O["Complete UV texture"]
T["Training: clean UV latent and noise"] -.->|Flow-matching supervision| D
Key Designs¶
1. Conditioning pathways: give image semantics and UV layout coordinate-appropriate inputs
The reference image and target UV map do not share pixel coordinates, so they cannot be combined position by position. The pooled CLIP feature enters the modulation pathway originally used for timestep and text embeddings in Flux, affecting all tokens through adaptive layer normalization (AdaLN) and primarily conveying color, material tone, and object identity. Dense DINOv2 image patch tokens replace the original text tokens in the joint-attention stream, allowing UV tokens to retrieve reference-image details by content. CLIP and DINO features are therefore not simply interchangeable cross-attention inputs.
Coarse UV shares the target texture's UV coordinates and follows a position-aligned pathway: the same frozen Flux VAE encodes it, and its latent is channel-concatenated with the noisy target UV latent before patchification. The VAE downsamples spatially by a factor of 8 and produces 64 latent channels; concatenation yields 128 input channels, while the output remains 64 channels. The training pipeline also precomputes per-position, per-channel latent means and standard deviations and uses consistent normalization during training and inference.
Crucially, coarse UV is neither the generation starting point nor a set of fixed known pixels. The entire target latent is generated from noise, so the model can use projected layout while rejecting incorrect colors. In Appendix A's keyboard example, coarse UV incorrectly paints the underside white, yet the final result restores a dark metallic casing consistent with the reference and surrounding regions. This is qualitative evidence for the mechanism, not a guarantee that every erroneous condition will be corrected.
2. Surface-aware positional encoding: derive cross-island relationships from 3D coordinates
The CCM is a UV-space image whose valid pixels store corresponding mesh points' 3D coordinates rather than RGB colors. Basic SAPE downsamples it to the token grid, assigning one surface coordinate to each UV token. Queries and keys then use this 3D position for rotary encoding instead of the 2D UV grid. Each head's feature dimensions are divided into three axis groups and rotated independently; the implementation splits its 120 dimensions into 40, 40, and 40.
For one token's query and another token's key, the mechanism in the paper's Equation (1) is:
Here, \(\mathcal{S}_{i}\) is the 3D surface coordinate supplied by the CCM. The relative positional contribution depends on \(\mathcal{S}_{j}-\mathcal{S}_{i}\) rather than separation in the UV layout, providing a spatial prior for directly relating regions that are close in 3D but separated by a UV seam. Attention still depends jointly on query and key content; it does not become distance-only neighborhood averaging.
โ3D proximityโ here means relative offsets in ambient 3D coordinates, not geodesic distance along the mesh surface or explicit topological adjacency. Coordinates alone cannot distinguish two thin surfaces that are spatially close but far apart along the surface. The CCM is also not part of the conditioning bundle \(\mathbf{c}\): it enters exclusively through SAPE and should not be described as an auxiliary geometry-feature conditioning branch.
3. Multi-level head allocation: one position per head, multiple granularities per token
A token covers a UV region. One representative 3D coordinate may connect different UV islands but cannot express fine positional variation within that patch. Multi-Level SAPE reads progressively subdivided CCM coordinates within the same patch: level \(l\) uses a \(2^l\times2^l\) sub-grid with \(4^l\) sub-cells. Rather than superimposing every level's coordinates within every head, it assigns each head one sub-cell coordinate at one level and rotates that head's queries and keys using the basic SAPE mechanism.
The implementation uses 3 levels and 24 attention heads, organized into 4 groups of 6. In each group, one head reads the level-0 whole-patch coordinate, one reads the group's level-1 sub-cell, and four read its nested level-2 sub-cells. The head budget across levels is therefore 4:4:16. Each head first computes attention under a single-granularity positional relation, and the output projection then mixes the results. This changes the positional responsibilities of existing heads without adding learnable parameters or an auxiliary geometry network.
Section 3.4 says that every sub-cell at every level is read by exactly one head, but the Section 4.2 configuration actually repeats the same level-0 coordinate across four groups. The explicit configuration is retained here: four and sixteen heads cover level-1 and level-2 cells respectively, while four heads share level-0. The general statement should not be expanded into a duplicate-free allocation rule across all levels.
A Worked Example¶
Consider Appendix A's keyboard, with a reference image and a UV-parameterized keyboard mesh as inputs. During inference, Kiss3dGen first generates multi-view images that are projected into coarse UV. Although the underside is covered by a projection, it is incorrectly labeled white. CLIP supplies global material cues, DINOv2 supplies local image appearance, and the coarse UV latent indicates the corresponding layout rather than forcing the generator to copy white pixels.
Suppose adjacent casing regions occupy two different islands after UV unwrapping. The CCM still maps them to nearby 3D positions. SAPE gives the corresponding tokens this surface positional relation in their queryโkey interactions, while multi-level head allocation distinguishes sub-regions within each token. After 28 denoising steps from noise, the VAE decodes a complete UV texture whose underside can retain the same dark metallic appearance as the rest of the casing. This example explains how conditions can be corrected; it is not a claim that the paper provides per-head attention visualizations proving each interaction.
Loss & Training¶
Training uses rectified flow matching. The clean UV latent is \(\mathbf{x}\), Gaussian noise is \(\boldsymbol{\epsilon}\), and the uniformly sampled timestep is \(t\in[0,1]\). Linear interpolation gives \(\mathbf{z}_{t}=(1-t)\mathbf{x}+t\boldsymbol{\epsilon}\), and the target velocity is noise minus the clean latent. The paper's Equation (2) is:
Training conditions come from rendered views of known meshes, not from calling the inference-time multi-view generator at every step. Each step projects a randomly selected 0, 1, 2, or 4 rendered views; 0 means blank coarse UV. Some views use lighting and others use albedo colors, deliberately introducing inconsistency so that coarse UV becomes a potentially incomplete or erroneous cue. The reference image is the front rendering; the four canonical views use azimuths of 0, 90, 180, and 270 degrees and an elevation of 5 degrees.
The model contains 6 MMDiT blocks and 12 single-stream blocks, with DINOv2 ViT-L/14 with registers and CLIP ViT-B/32 as image encoders. Training uses 8 NVIDIA A800 GPUs, a per-GPU batch size of 16, and 170K steps. The optimizer is 8-bit AdamW with a learning rate of \(10^{-5}\), weight decay of \(10^{-4}\), betas of 0.9 and 0.999, gradient clipping at 1.0, bf16, and gradient checkpointing.
Inference holds the image conditions, coarse UV, and CCM fixed, integrates from noise toward the clean latent for 28 Euler steps, and decodes with the frozen VAE. โDirect UV generationโ refers to generating the final target in latent UV space, not to eliminating multi-view generation and projection from the entire system: standard inference still constructs coarse UV by projecting Kiss3dGen views.
Key Experimental Results¶
Main Results¶
The dataset section names TexVerse, PartNext, and Objaverse as training sources, retaining approximately 100K textured meshes after filtering. UV masks, ground-truth UV textures, and CCMs all have 1024ร1024 resolution. Evaluation uses the full Google Scanned Objects (GSO) dataset, disjoint from training. All methods receive the same mesh and single front-facing reference image; the generated textures are evaluated through multi-directional albedo renderings, including top and bottom views.
PSNR and SSIM measure pixel-level fidelity against ground-truth albedo images, while FID and KID measure distributional similarity. The following table preserves the values in the paper's Table 1. The supplied text does not specify an additional KID scaling convention, so no multiplier is inferred.
| Method | PSNR โ | SSIM โ | FID โ | KID โ |
|---|---|---|---|---|
| TEXGen | 24.45 | 0.9417 | 44.09 | 34.32 |
| FlexPainter | 23.21 | 0.9463 | 47.04 | 39.25 |
| Hunyuan3D-2.1 | 24.74 | 0.9450 | 39.872 | 33.34 |
| UniTEX | 24.23 | 0.9504 | 42.11 | 39.21 |
| DirectUV | 25.13 | 0.9527 | 39.17 | 30.03 |
DirectUV leads all four reported metrics, but the strongest comparator differs by metric: PSNR exceeds Hunyuan3D-2.1 by 0.39, while SSIM exceeds UniTEX by 0.0023. FID and KID are not direct error measurements for an individual seam region, and this table alone does not establish faster end-to-end inference.
Ablation Study¶
The following table comes from the paper's Table 2. Removing SAPE only replaces positional encoding with 2D UV-grid RoPE while keeping the rest of the model unchanged. Removing the multi-level design makes all heads share the level-0 coordinate.
| Config | PSNR โ | SSIM โ | FID โ | KID โ |
|---|---|---|---|---|
| Without SAPE | 23.82 | 0.9376 | 58.42 | 72.36 |
| Single-level SAPE | 24.36 | 0.9462 | 40.86 | 36.74 |
| Multi-Level SAPE (full model) | 25.13 | 0.9527 | 39.17 | 30.03 |
Key Findings¶
- Relative to removing SAPE, the full model increases PSNR by 1.31 and SSIM by 0.0151, while reducing FID by 19.25 and KID by 42.33. Figure 4 qualitatively attributes the improvement to fewer cross-island color mismatches, consistent with the positional-prior motivation.
- Single-level SAPE already substantially improves FID and KID. Adding the multi-level design further raises PSNR from 24.36 to 25.13 and SSIM from 0.9462 to 0.9527. The authors mainly explain this second improvement through finer sub-patch positional precision and reduced blur, not additional appearance conditions.
- The well and trash-bin interiors and the fan in Figure 3 illustrate advantages in occlusion completion and conflicting conditions. Appendix B shows plausible textures with coarse UV formed from 0, 1, 2, or 4 views, but provides no corresponding quantitative completeness curve.
Highlights & Insights¶
- Geometry can change how attention compares positions rather than only entering as an input feature. Even when a 2D layout separates a continuous surface, generation can use 3D relative positions instead of deferring seam correction until afterward.
- Multiple levels are allocated across heads rather than stacked into a mixed coordinate. Each head retains a single-granularity rotation rule, and the output projection integrates granularities, suggesting a reusable structure for other generation tasks with geometric parameterizations.
- Demoting projections to conditions and regenerating the entire texture leaves room to correct errors rather than only fill holes. The benefit relies on exposure to incomplete and inconsistent conditions during training; starting from noise alone does not guarantee it.
Limitations & Future Work¶
- The authors explicitly acknowledge dependence on UV parameterization quality. Appendix D's many tiny islands and narrow slivers make token-level CCM representations too coarse; multi-level subdivision cannot fully recover detail, causing blur and color leakage. Training also filters heavily fragmented layouts, so such inputs depart from the training distribution.
- Ambient 3D coordinates are not geodesic distances, so spatially close but topologically separate surfaces may require additional constraints. This is a potential mechanistic boundary; the paper offers no dedicated quantitative experiment on thin shells or closely contacting surfaces.
- Quantitative evaluation mainly uses one test dataset, GSO. Applications to Tripo- and Rodin-generated meshes are qualitative only. The text provides no end-to-end latency, memory use, repeated-run variance, or dedicated quantitative metrics partitioned by seams or occlusion.
- Training-data identity is inconsistent between the text and bibliography: Section 4.1 says TexVerse, whereas reference [27] lists Objaverse-XL. This note preserves the discrepancy rather than silently harmonizing the dataset name or source.
- Further work could test invariance to UV rearrangement and investigate UV-aware preprocessing or adaptive tokenization. Improvements on fragmented UV layouts require separate validation; adding more subdivision levels is not an established solution.
Related Work & Insights¶
- vs TEXGen / TexGarment: These methods introduce geometry through point cloud attention or cross-attention with 3D features; DirectUV expresses 3D relations in UV tokens' positional encoding. The distinction is where geometry influences interactions, not the first use of 3D information.
- vs RomanTex / Hunyuan3D-2.1: Both already use 3D-aware RoPE, mainly during multi-view image generation before projection and baking. DirectUV places this spatial idea in latent UV attention so that cross-island relations participate in generating the final texture.
- vs UniTEX: UniTEX combines projected partial textures with completion in a triplane feature space; DirectUV does not treat projected pixels as final values that must be preserved. Correctability is an advantage, while the absence of hard consistency constraints on observed regions is a trade-off.
- Research Direction: Keep the mesh and texture fixed while changing UV-island placement or cutting, compare SAPE with 2D RoPE, and add spatially close but topologically separate surfaces as controls. This would distinguish insensitivity to parameterization from actual surface-topology understanding; it is an untested direction, not a conclusion of this paper.
Rating¶
- Novelty: 4/5. Combines surface-coordinate positional encoding with per-head multi-granularity allocation for direct UV generation, with a clear contribution boundary despite prior 3D RoPE work.
- Experimental Thoroughness: 3/5. Main comparisons and two-stage ablations support the core design, but multiple datasets, dedicated local metrics, and resource comparisons are missing.
- Writing Quality: 4/5. Clearly separates appearance conditioning from positional geometry, but the head-allocation summary and dataset citation contain unresolved ambiguities.
- Value: 4/5. Useful for reference-image texturing of existing meshes, particularly occlusion and seams, while remaining constrained by UV quality and conditioning sources.