Skip to content

CAST3D: Customizing Arbitrary 2D Assets into 3D World

Conference: ECCV 2026
Paper: ECCV Official
Area: 3D Vision
Keywords: 3D Asset Customization / Compositional Generation / Diffusion Models / Stochastic Trajectory Manipulation / Training-Free

TL;DR

CAST3D presents a training-free framework for customized 3D composition, bridging 2D assets and 3D generation through 3D layout hinting, stochastic trajectory manipulation, and connectivity-based pruning to synthesize coherent, faithful 3D objects under textual guidance.

Background & Motivation

Recent breakthroughs in diffusion models have led to an abundance of high-quality, easily editable 2D assets across digital art, design, and virtual environments. However, while 2D content creation has become increasingly accessible, generating coherent 3D objects and scenes remains constrained by intricate geometric representations and heavy optimization costs. Leveraging accessible 2D visual assets (such as character models, costumes, and props) and composing them into custom 3D assets represents a natural and promising frontier. Yet, contemporary 3D diffusion models such as Trellis and Hunyuan3D are predominantly trained and optimized for isolated single-object synthesis, lacking the compositional reasoning and spatial awareness needed to assemble multiple assets coherently.

Existing attempts to tackle this problem typically collapse into two flawed paradigms. The first is "Uplift-then-Compose", where each 2D asset is reconstructed independently into 3D and then arranged spatially according to text prompts. Because assets are generated in isolation, they exhibit misaligned coordinate frames, conflicting scales, and arbitrary orientations, inevitably resulting in severe clipping and geometric collisions. The second paradigm is "Customize-then-Generate", which relies on multi-concept 2D diffusion models or vision-language models to merge the 2D assets into a composite image before lifting it into 3D. This pipeline is severely bottlenecked by the 2D model's ability to maintain identity consistency and suffers from catastrophic occlusion, perspective distortion, and multi-view inconsistencies when converting 2D images into 3D space.

The core tension lies in the fact that discrete 2D observations lack consistent 3D spatial priors, making direct spatial assembly in the 2D domain or post-hoc 3D rigid placement prone to geometric failure. The authors recognize that layout reasoning and asset composition should occur natively within the latent space of 3D diffusion models to exploit their rich geometric priors. Core idea: CAST3D establishes a training-free two-stage pipeline—3D layout hinting and compositional generation—leveraging Stochastic Trajectory Manipulation (STM) and connectivity-based pruning to guide pre-trained 3D flow models toward geometrically consistent and appearance-faithful multi-asset synthesis.

Method

Overall Architecture

Given a target textual description \(c\), a set of 2D asset pairs \(\{(o_i, I_i)\}_{i=1}^N\) (matching text concept \(o_i\) with reference image \(I_i\)), and a designated anchor asset index \(i_B\), CAST3D outputs a complete 3D asset whose components strictly reflect their corresponding 2D visual identities while adhering to the global composition. The workflow consists of two consecutive stages: 3D Layout Hinting and Compositional Generation. In the first stage, the base geometry of the anchor asset is modulated via STM into a coarse 3D layout proposal, from which component-level voxel layout hints are extracted using 3D back-projection of 2D open-vocabulary segmentation masks. In the second stage, fine component geometries are generated according to their layout hints and reference images, assembled using connectivity-based pruning to eliminate intersecting artifacts, and seamlessly blended across appearance boundaries.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: 2D Assets + Text Guidance"] --> B["3D Layout Hinting<br/>Anchor deformation and 3D Gaussian backprojection"]
    B --> C["Stochastic Trajectory Manipulation (STM)<br/>Reverse SDE formulation and noise mapping"]
    C --> D["Connectivity-based Pruning & Composition<br/>6-connected component filtering and soft-boundary blending"]
    D --> E["Output: Coherent & Faithful 3D Asset"]

Key Designs

1. Stochastic Trajectory Manipulation (STM): Balancing Geometric Preservation and Semantic Modification Standard Rectified Flow models perform deterministic sampling by solving a Probability Flow ODE (PF-ODE): \(\dif \boldsymbol{x}_t = \boldsymbol{v}(\boldsymbol{x}_t, t)\dif t\). Although fast and high-fidelity, deterministic inversion accumulates linearization errors and causes noticeable structural drift when conditioned on strong editing prompts. To overcome this, CAST3D converts the flow matching reverse process into an equivalent reverse-time Stochastic Differential Equation (SDE) that preserves the identical marginal density: $$ \dif \boldsymbol{x}_t = \left[\boldsymbol{v}(\boldsymbol{x}_t, t) - \frac{1}{2} g^2(t)\nabla \log p_t(\boldsymbol{x}_t)\right]\dif t + g(t)\dif\bar{\boldsymbol{w}} $$ where \(g(t) = \sqrt{2\lambda_{\text{diff}} t}\) is the diffusion coefficient. STM extracts the applied noise from a reference sample trajectory and maps it to the target trajectory via an identity mapper \(\mathcal{F}\) while injecting prior states through a trajectory generator \(\mathcal{G}\). Mathematically, STM induces an integration kernel on velocity differences that suppresses structural drift during early timesteps while amplifying semantic modification at later timesteps. This allows the model to firmly anchor the overall 3D spatial layout while leaving sufficient generative flexibility to assimilate new component features.

2. 3D Layout Hinting: Native 3D Spatial Reasoning via Cross-Modal Back-Projection Rather than relying on error-prone 2D layout compositions, this design directly executes layout reasoning within the 3D domain. The anchor image \(I_B\) is first processed by Trellis's sparse structure flow transformer to generate base geometry \(\boldsymbol{p}_B\). STM is then applied under prompt \(c\) to yield a voxelized layout proposal \(\boldsymbol{p}_C\). To disentangle this holistic 3D layout into component-level spatial bounding hints, \(\boldsymbol{p}_C\) is decoded into 3D Gaussian Splatting (3DGS) representations and rendered across multiple viewpoints. A 2D open-vocabulary segmenter (LangSAM) segments the projected views conditioned on the component labels \(o_i\), and the resulting masks \(M_i^v\) cast weighted votes back to the 3D Gaussian primitives: $$ m_j^i = \text{sgn} \sum_{v \in \mathcal{V}} \sum_{p \in P_j^v} \alpha T_j(p)\left[\mathbb{I}(p \in M_i^v) - \mathbb{I}(p \notin M_i^v)\right] $$ By thresholding the density ratio of masked Gaussians within each voxel at \(\tau_{\text{mask}}\), the framework produces clean 3D voxel layout hints \(\boldsymbol{h}_i\) for each asset, completely circumventing multi-view perspective discrepancies.

3. Connectivity-based Pruning and Composition: Resolving Geometric Collisions and Smoothing Transitions When individual component geometries \(\boldsymbol{p}_i\) are generated independently using their respective 2D assets and layout hints, direct boolean union causes disruptive clipping and disconnected artifacts. CAST3D isolates the spatial difference \(\boldsymbol{p}_{\text{diff}} = (\boldsymbol{p}_B \oplus \boldsymbol{p}_C) \odot \sum_{i \ne i_B} \text{Expand}(\boldsymbol{p}_i)\) to form a composition anchor \(\hat{\boldsymbol{p}}_B = \boldsymbol{p}_B \oplus \boldsymbol{p}_{\text{diff}}\). It then evaluates the residual voxels under 6-connectivity, pruning disjoint debris and preserving only the dominant connected component alongside stable topological structures. The assembled geometry is refined with a lightweight STM pass to produce a smooth surface \(\boldsymbol{p}_{\text{GC}}\). During appearance synthesis, component SLAT features \(\boldsymbol{z}_i\) are mapped to corresponding geometric regions with optimal velocity fields, softened by Gaussian kernel transitions, and left unguided during the final \(\tau_{\text{fine}}\) timesteps to allow diffusion dynamics to naturally seal the seams.

Loss & Training

CAST3D is entirely training-free and operates at inference time on pre-trained 3D diffusion backbones. Experiments are executed on a single NVIDIA RTX 4090 GPU with 25 sampling steps and a rescaling ratio of 3. Text classifier-free guidance (CFG) is set to 7.5, and image CFG is set to 5.0. In the layout hinting phase, STM step bounds \(n_{\text{max}}\) range within \([15, 21]\) based on structural transformation magnitude, and \([9, 14]\) during compositional generation. Hyperparameters are set to \(\lambda_{\text{diff}} = 0.6\), voxel mask threshold \(\tau_{\text{mask}} = 0.9\), and refinement cutoff \(\tau_{\text{fine}} = 0.15\).

Key Experimental Results

Main Results

The evaluation benchmark comprises 45 diverse 2D assets assembled into 27 challenging customized 3D composition scenarios. Baselines include "Uplift-then-Compose" (UtC) using Trellis with centroid and volume alignment, and "Customize-then-Generate" (CtG) using Qwen-Image-Edit-2511 across single-view and multi-view setups. Evaluation is conducted using Gemini 3 Pro across Faithfulness, Aesthetic quality, and Structural Rationality (scale 1–10), complemented by multi-view CLIP textual alignment (CLIP-T) and component-level visual identity alignment (CLIP-I).

Method Faithfulness ↑ Aesthetic ↑ Rationality ↑ Average ↑ CLIP-T ↑ Compon.-lvl CLIP-I ↑
Uplift-then-Compose (UtC) 8.33 7.25 5.29 6.96 22.80 0.762
Customize-then-Generate (CtG, Single View) 5.73 5.75 6.87 6.12 22.78 0.733
Customize-then-Generate (CtG, Multiple Views) 6.77 5.89 6.03 6.23 21.64 0.695
CAST3D (Ours) 9.19 8.68 8.65 8.84 22.84 0.789

Ablation Study

The ablation study analyzes the role of Stochastic Trajectory Manipulation (STM) and connectivity-based pruning in sustaining structural integrity. A comprehensive human user study was also conducted to benchmark perceptual faithfulness and spatial plausibility.

Configuration / Study Variant Faithfulness Rationality Average Note
CAST3D Full Model 8.62 8.49 8.55 Highest human preference; coherent shape with seamless seams
w/o STM 6.81 6.15 6.48 Severe layout drift; components mislocated relative to base
w/o Pruning 7.54 6.73 7.13 Noticeable clipping artifacts and disjoint geometric debris
User Study: UtC Baseline 7.41 5.92 6.66 Decent component appearance but unnatural structural clipping
User Study: CtG (Single View) 6.49 6.34 6.41 Suffer from rear occlusion and degraded visual fidelity
User Study: CtG (Multiple Views) 6.85 6.60 6.72 Geometry collapses due to view inconsistency

Key Findings

  • Superiority of Native 3D Reasoning over 2D Intermediates: Intermediate 2D composition approaches consistently degrade during 3D lifting—single-view conditioning creates dark unobserved occlusions, while multi-view conditioning suffers from slight cross-view hallucinations that collapse 3D geometry. Proposing layouts natively in 3D eliminates these pitfalls.
  • STM Velocity Integration Dynamics: STM effectively stabilizes the macro-geometry early in the reverse trajectory while accommodating fine-grained semantic modifications later. Removing STM leads to dramatic positional detachment between accessories and base subjects.
  • Human vs. VLM Evaluation Divergence: In CtG baselines, advanced VLMs penalize multi-view generation more heavily due to localized geometric collapses, whereas human evaluators favor multi-view results over single-view ones for overall silhouette completeness.

Highlights & Insights

  • Unified SDE Framework for Flow Trajectory Editing: Derives an equivalent stochastic reverse formulation for Rectified Flow models, unifying prior methods like FlowEdit and FlowAlign under an integration kernel perspective and demonstrating effective 3D trajectory control without fine-tuning.
  • Seamless Cross-Dimensional Spatial Grounding: Synergizes 3D Gaussian Splatting rendering with 2D open-vocabulary segmentation voting, establishing accurate 3D voxel hints from unstructured text and single-view images with minimal computation.
  • Topology-Aware Geometric Cleansing: Employs 6-connectivity domain filtering coupled with Gaussian soft boundaries, replacing crude CSG boolean unions and solving boundary seam degradation in multi-component 3D assembly.

Limitations & Future Work

  • Dependency on Base Anchor Selection: The framework requires designating a primary anchor asset \(i_B\) to establish the base coordinate layout. Synthesizing egalitarian multi-entity interactions without a dominant base remains non-trivial.
  • Resolution Constraints under Extreme Scale Variations: When composing assets with drastic scale disparities (e.g., miniature accessories on massive creatures), fixed-grid voxelization risks smoothing out fine topological nuances, suggesting the need for octree-based adaptive representations.
  • vs. Uplift-then-Compose (UtC): UtC isolates 3D reconstruction and enforces rigid spatial placement, failing to accommodate deformation at contact surfaces. CAST3D informs component generation with contextual layout hints, ensuring organic surface adaptation.
  • vs. Multi-Concept 2D Diffusion Models (e.g., MS-Diffusion, Qwen-Image-Edit): Pure 2D models lack genuine depth reasoning and occluded surface hallucination. CAST3D lifts conditioning into pre-trained 3D latent spaces, operating with intrinsic geometric consistency.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates the task of customized 3D composition from arbitrary 2D assets and resolves it within native 3D diffusion latent spaces in a training-free manner.
  • Experimental Thoroughness: ⭐⭐⭐⭐☆ Thorough evaluation blending advanced VLM judging, CLIP feature alignment, and rigorous human user blind tests.
  • Writing Quality: ⭐⭐⭐⭐⭐ Mathematically elegant SDE derivation with coherent architectural articulation and intuitive visual breakdowns.
  • Value: ⭐⭐⭐⭐⭐ Offers immediate practical utility for automated virtual asset generation, game asset kitbashing, and personalized avatar design.