Skip to content

Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation

Conference: NeurIPS2026 (task-list assignment; the version read is arXiv v1)
arXiv: 2609.36918
Area: 3D Vision
Keywords: part-level generation, segmentation conditioning, structured latents, cross-part attention, canonical space

TL;DR

Seg3DParts injects externally supplied part segmentations into a two-stage 3D generator and exchanges structural information across part latents, directly producing separate meshes in a space that preserves training-time relative coordinates; its strength is controllable decomposition with less interpenetration, not automatic discovery of unknown parts.

Background & Motivation

A complete 3D mesh can support visualization without supporting editing or reassembly: when a chair's back, seat, and legs form one surface, they are not stable manipulation units. Existing pipelines often generate a whole object from an image, segment its surface, and complete each component, as in the combination of Hunyuan3D, PartField, and HoloPart. This provides explicit parts, but local completion sees incomplete surfaces. Under severe occlusion, individually plausible parts can still have inconsistent scales, conflicting positions, or interpenetration.

Joint generators instead couple multiple parts during generation. PartCrafter models their relationships through local and global attention, while PartPacker uses a dual-volume representation to preserve geometric separation; Seg3DParts therefore does not introduce part coordination for the first time. It emphasizes a remaining control problem: if latent slots implicitly determine part identity, users cannot easily specify which slot represents the chair back or request another decomposition granularity for the same image. A whole-object segmentation condition can still leave region-to-part allocation to the generator.

This paper makes the choice of parts an input constraint and leaves their structural compatibility to joint modeling. A 2D mask establishes correspondence to visible regions but cannot supply occluded 3D geometry; missing structure depends on learned priors, global image context, and other parts. Core Idea: explicitly specify generated part identities through per-part segmentation, preserve correspondence with local conditioning, and exchange information across parts in both stages' VAEs and DiTs to learn separate meshes that share an assembly-ready coordinate space.

Method

Overall Architecture

The inputs are one RGB image and a set of part segmentations supplied by an external model such as SAM or by user annotations, rather than parts discovered by Seg3DParts itself. The outputs are separate meshes corresponding to those conditions. The backbone follows TRELLIS's two-stage design: first generate coarse sparse voxel occupancy for each part, then generate fine geometry latents at those spatial locations and decode surfaces.

Both stages are made part-aware. Segmentation-grounded conditioning sends local image features into the corresponding part's AdaLN, while global image cross-attention provides shared context. Part-aware two-stage generation keeps distinct part representations in the VAEs and DiTs and exchanges information between them. Shared canonical-space decoding relies on training data retaining the original object's relative positions and scales, rather than predicting an assembly transform after generation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    I["RGB image + external part masks"] --> C["Segmentation-grounded conditioning"]
    C --> S["Part-aware two-stage generation<br/>Stage 1: sparse structure"]
    S --> L["Part-aware two-stage generation<br/>Stage 2: fine geometry latents"]
    L --> D["Shared canonical-space decoding"]
    D --> O["Separate part meshes"]
    T["Training supervision: aligned part meshes"] -.->|Occupancy supervision and target latents| S
    T -.->|Geometry supervision and target latents| L
    T -.->|Preserve relative scales and positions| D

Solid arrows show inference data flow; dashed arrows indicate training supervision only. Both stages use part-aware VAEs and rectified-flow DiTs, but their cross-part interaction schedules differ. Interaction is not merely a single attention operation immediately before final assembly.

Key Designs

1. Segmentation-grounded conditioning: a controllable image correspondence for each latent group

Each part mask selects its corresponding region from the full image to form a part-conditioning image. Frozen DINOv2 features are mapped to a compact appearance embedding by a lightweight trainable MLP. This embedding is combined with timestep and total-part-count embeddings to produce part-specific AdaLN scale, shift, and gating parameters. The appendix explicitly states that indexed broadcasting applies these parameters only to the corresponding part's tokens, rather than distributing one region's condition indiscriminately across all parts.

The original image also passes through DINOv2, and its global features enter the DiT through cross-attention. Local AdaLN and global image cross-attention serve different purposes: the former specifies the region associated with the part being generated, while the latter supplies shared object-category, silhouette, and visible-structure context. Both stages use this conditioning design, so identity constraints influence coarse occupancy as well as final geometric details.

This control grounds parts in 2D regions; it does not supply semantic names or complete 3D shapes. Changing masks can change decomposition boundaries and granularity, but the paper provides no guarantee that arbitrary masks yield meaningful decompositions. Fully invisible parts, in particular, must already have their existence and part count specified externally. A blank image condition is not a mechanism for discovering unknown parts.

2. Part-aware two-stage generation: allocate each part's occupancy before jointly refining surfaces

The stage-1 VAE represents coarse structure with \(64^{3}\) voxel occupancy grids. Each part is a separate element along the batch dimension, with shared encoderโ€“decoder weights. Hierarchical 3D convolutions compress occupancy into Gaussian latents, and the decoder reconstructs binary occupancy. Cross-part interaction is restricted to the compact latent bottleneck, while higher-resolution encoding and decoding remain independent per part. This exchanges information about spatial extent without continuously mixing all local voxels globally. The corresponding DiT models coarse structure on a \(16^{3}\) grid representation; this latent-space resolution must not be mistaken for the final occupancy resolution.

Conditioned on sparse structure, stage 2 generates sparse latents containing fine geometry for each part. The paper replaces TRELLIS's voxel-feature projection with mesh-surface encoding. The appendix specifies a sparse Transformer encoder and decoder, each with 12 blocks, with cross-part interaction inserted symmetrically in both. The stage-2 DiT processes sparse features in a \(64^{3}\) coordinate space, rather than generating a dense full volume for every part. Surface encoding supports geometric detail, but its contribution is not isolated by a dedicated ablation.

Both DiTs have 24 Transformer blocks, hidden dimension 1024, and 16 attention heads. Most blocks apply intra-part self-attention, global image cross-attention, and feed-forward transformations; a smaller subset inserts cross-part attention immediately after intra-part self-attention. The paper describes a fixed sparse interleaving schedule but does not list layer indices, so a particular interaction interval cannot be assumed.

Cross-part attention lets the current part query tokens from other parts, instead of first merging all representations into an identity-free whole. Its residual mechanism is:

\[ \mathbf{X}^{k}_{\text{out}}=\mathbf{X}^{k}+\mathrm{Attn}\!\left(\mathbf{Q}=\mathbf{X}^{k},\;\mathbf{K}=[\mathbf{X}^{j}]_{j\neq k},\;\mathbf{V}=[\mathbf{X}^{j}]_{j\neq k}\right). \]

Brackets denote concatenation along the token dimension, and \(\mathbf{X}^{k}\) denotes the current part's latent tokens. Intra-part attention first maintains the part's surface or occupancy information, then other parts help update its scale, position, and contact relationships. This explains why independently generating parts and concatenating them afterward is not equivalent: other components influence the geometry as it forms. The appendix also specifies a self-attention fallback for a single part. This is an implementation special case, not direct evaluation of the equation with an empty key/value set.

3. Shared canonical-space decoding: learn existing relative positions instead of solving assembly poses afterward

During dataset construction, parts are separated from complete objects but retain their original scales and positions. Encoding, generation, and decoding use these aligned representations, and stage 2 finally extracts each part's mesh through FlexiCubes. Although the output files are separate, they refer to one canonical space, without subsequent per-part translation, scaling, or optimization-based alignment.

Separate outputs do not imply independently chosen coordinates. Cross-part interaction helps learn relationships during training and generation, while shared training coordinates specify the target placement. The appendix describes normalization of decoded values for stable supervision and explicitly says canonical alignment is preserved. This should not be interpreted as independently centering and rescaling every part while somehow retaining assembly positions.

Spatial consistency is learned rather than imposed by hard geometric constraints. The method does not explicitly guarantee collision-free meshes, exact contact, or manufacturability; its main-experiment part-overlap IoU is also nonzero. Thus, the absence of post-hoc alignment describes the inference procedure, not a guarantee of correct assembly for every input.

A Worked Example

Consider an illustrative chair image with occlusion, where external input specifies a seat, a back, and four legs: 6 parts in total. This is a pipeline illustration, not an additional experiment from the paper. Each visible part receives its own conditioning image and AdaLN parameters, while all share global image context. Stage 1 jointly determines coarse occupancy for the 6 parts, and stage 2 generates geometry at those locations while consulting the states of other parts.

If one rear leg is fully invisible, Appendix F retains a condition for that part and supplies an all-black conditioning image. Fully hidden training parts are handled in the same way, teaching the model that black means no visual evidence rather than no part. The leg's shape and position must therefore be inferred from visible structure and learned priors, not determined by its mask. The resulting 6 meshes share coordinates. If the user merges the seat and back into one region, the interface permits another granularity, but the paper demonstrates such changes qualitatively rather than with a separate quantitative controllability metric.

Loss & Training

The two VAEs first learn latent geometry representations; the two DiTs then use flow matching to predict velocity fields from noise toward target latents. The stage-1 VAE uses Dice occupancy reconstruction with KL regularization weighted by \(10^{-3}\). The stage-2 VAE supervises rendered silhouettes, depth, truncated signed distance fields, and normals, with KL weight \(10^{-6}\). Normal supervision combines pixelwise \(\ell_{1}\), SSIM, and LPIPS terms to capture numerical and perceptual consistency.

The stage-2 depth, TSDF, normal-SSIM, and normal-LPIPS weights are 10.0, 0.01, 0.2, and 0.2, respectively. Depth is specified as Smooth-\(\ell_{1}\), but the source does not give complete computational details for silhouette and TSDF losses or such flow-matching details as the time-sampling distribution. No guessed exact loss definition is added here.

The VAEs are trained on 8 A800 GPUs for approximately 2 days, and the two stages' DiTs on 16 A800 GPUs for approximately 1 week. The source does not clearly separate training duration per model. Inference uses rectified-flow integration with 50 and 30 steps for the respective stages and image-conditioning classifier-free guidance (CFG) scale 3.0, without test-time optimization or post-processing.

Key Experimental Results

Main Results

PartObjectNet combines Objaverse, Texverse, and PartNet, retaining objects with 2โ€“15 parts after automatic filtering and manual curation. The dataset section describes approximately 200K objects, whereas the abstract claims over 200K objects and 1M annotated parts; an exact verifiable count is not provided. An internal test set randomly holds out 500 objects. The paper also evaluates PartObjaverse-Tiny, described as independent of training data, but the full text read does not state the size of this external test set.

Global geometry concatenates all part meshes and compares them with ground truth in a common \([-1,1]^{3}\) space, uniformly sampling 16K points from each surface. [email protected] is the harmonic mean of surface-point matching precision and recall at distance threshold 0.1. CD measures bidirectional nearest-neighbor surface distance, but the paper does not specify its squared form or any additional scaling. Part-level metrics compare explicitly corresponding parts, with IoU computed on \(64^{3}\) voxel grids. Part-overlap IoU averages pairwise volumetric IoU between generated parts and is lower-is-better, unlike part-to-ground-truth IoU.

The following representative comparisons come from main-text Table 1. Better global geometry should not automatically be equated with better part identity.

Method PartObjectNet [email protected] โ†‘ CD โ†“ Overlap IoU โ†“ PartObjaverse-Tiny [email protected] โ†‘ CD โ†“ Overlap IoU โ†“
PartCrafter 0.698 0.248 0.050 0.731 0.182 0.049
PartPacker 0.876 0.115 0.033 0.802 0.138 0.033
OmniPart 0.885 0.108 0.057 0.763 0.168 0.041
Hunyuan3D2.1 + PartField + HoloPart 0.858 0.121 0.067 0.796 0.143 0.041
Seg3DParts 0.917 0.085 0.012 0.810 0.128 0.018

Against OmniPart on the internal test set, global FS increases by 0.032 and CD decreases by 0.023. Against PartPacker on the external test set, global FS increases by only 0.008. Overall fidelity, part control, and overlap should therefore be considered separately rather than summarized by one improvement magnitude.

Table 2 compares only OmniPart and Seg3DParts, which support explicit part correspondence. Main-text Figure 3 states that both receive ground-truth part segmentations, so these results should not be read as end-to-end performance with noisy masks from an external segmentation model.

Dataset Method Part [email protected] โ†‘ Part CD โ†“ Part IoU โ†‘
PartObjectNet OmniPart 0.565 0.417 0.551
PartObjectNet Seg3DParts 0.774 0.192 0.781
PartObjaverse-Tiny OmniPart 0.455 0.418 0.429
PartObjaverse-Tiny Seg3DParts 0.639 0.310 0.702

Internal part FS increases by 0.209, and external part IoU by 0.273. These results support more accurate generation given established correspondence, not automatic part-discovery accuracy. The absence of other methods from this table does not imply zero part quality.

Ablation Study

All variants in main-text Table 3 are trained for 5K steps on only 1K randomly sampled objects. This is a shared small-scale setting: its full-model row is not the fully trained model above, and the numbers must not be treated as one experiment.

Config Global FS โ†‘ Global CD โ†“ Overlap IoU โ†“ Part FS โ†‘ Part CD โ†“ Part IoU โ†‘
Full model 0.856 0.129 0.054 0.631 0.467 0.649
Without cross-part interaction 0.838 0.138 0.103 0.531 0.534 0.434
Single whole-object segmentation condition 0.840 0.137 0.126 0.488 0.589 0.444

The interaction ablation replaces all multi-part blocks with single-part blocks; the whole-object variant replaces per-part conditioning with one whole-object segmentation map. The former increases overlap IoU by 0.049 and reduces part FS by 0.100; the latter increases overlap IoU by 0.072 and reduces part FS by 0.143. This shows how overall shape can conceal part-level failure, but does not establish exact contribution magnitudes under full-scale training.

Key Findings

  • Global FS falls by only 0.018 and 0.016 in the two ablations, while part metrics deteriorate more substantially. Similar silhouettes do not establish correct decomposition or absence of interpenetration.
  • Per-part conditioning supports identity grounding, and cross-part interaction supports geometric coordination. The ablations suggest complementary roles but do not isolate interaction contributions from each stage's VAE and DiT.
  • Figure 5 qualitatively demonstrates different segmentations of the same image; Appendix F illustrates recovery of fully occluded parts. Neither provides dedicated sample counts, error statistics, or controlled comparisons, making their evidence weaker than the main geometry evaluation.

Highlights & Insights

  • Input can determine control granularity instead of relying on fixed latent-slot interpretations. Per-part AdaLN maintains correspondence between regions and token groups, while global image cross-attention retains shared context needed to understand the object.
  • Coordination belongs inside representation learning and generation. The VAE learns relational latents and the DiT exchanges states during generation, influencing missing geometry more directly than estimating poses only after independent completion.
  • Shared coordinates are a joint responsibility of data and modeling. Retaining relative scales and positions during training means separate output meshes need not require reassembly; this principle can transfer to compositional asset generation with explicit regional conditions.

Limitations & Future Work

  • The authors acknowledge sensitivity to input segmentation quality and show failures with low-quality or semantically ambiguous boundaries in Appendix G. Main quantitative results use ground-truth segmentation; future evaluation should separately cover automatic masks, manual masks, and controlled mask perturbations.
  • Fully hidden parts still require external specification of their existence and a black-image condition signaling absent visual evidence. The appendix offers only qualitative examples without statistics by occlusion level, category, or geometric error; it does not establish autonomous discovery of all hidden parts.
  • Dataset filtering covers 2โ€“15 parts. Reliability outside that range, inference latency, and memory scaling with part count are untested. Concatenating other parts' tokens may incur scaling costs, but the full text contains no measured complexity analysis.
  • Low overlap does not establish correct contact or functional usability. Contact-surface error, assembly stability, and consistency after part editing would complement global FS and volumetric IoU.
  • Reproduction details remain incomplete: exact interaction-layer schedules, all reconstruction-loss definitions, inference cost, and a code link are absent from the version read. The main text and appendix call the mesh encoder TripoSF [34], whereas reference [34] is titled Sparseflex; the baseline is called Hunyuan3D 2.1, whereas [14] is Hunyuan3D 2.0. Experimental names are retained as reported rather than silently harmonizing versions.
  • vs TRELLIS: The sparse-structure and structured-latent two-stage backbone is inherited. Changes concern per-part representations, local conditioning, and cross-part coordination; neither the backbone nor rectified flow itself should be credited as new here.
  • vs PartCrafter / PartPacker: These methods already model multiple parts jointly. Seg3DParts emphasizes explicit region correspondence through external segmentation, at the cost of stronger input dependence. Dual-volume representations or part attention are not inherently incompatible with segmentation grounding.
  • vs OmniPart: Both can use segmentation, but Seg3DParts injects per-part image embeddings into corresponding tokens through AdaLN, rather than relying primarily on global segmentation and part-space prediction for allocation. The ground-truth-mask part-level comparison is its most direct quantitative control-related evidence.
  • vs Hunyuan3D + PartField + HoloPart: Whole-object generation followed by segmentation and completion differs from coordination during generation. The transferable lesson is to introduce part relationships while geometry is forming rather than assuming independent completions will assemble consistently afterward.

Rating

  • Novelty: 4/5. Explicit per-part image grounding combined with two-stage cross-part coordination is a clear contribution, although the backbone and joint modeling have established precedents.
  • Experimental Thoroughness: 3/5. Internal and external geometry comparisons are substantial, while control, full occlusion, and mask-noise evaluation remain primarily qualitative.
  • Writing Quality: 3/5. The method is coherent and the appendix supplies architecture and weights, but reference-model naming and some implementation details remain unclear.
  • Value: 4/5. Useful for editable part-level assets, provided input segmentation quality and the limits of learned spatial consistency are respected.