Skip to content

Penetration-Free Compositional 3D Generation via Gaussian Surface Offset

Conference: ECCV 2026
Paper: ECCV Official
Area: 3D Vision
Keywords: Compositional 3D Generation, 3D Gaussian Splatting, Penetration-Free, Signed Distance Function, Score Distillation Sampling

TL;DR

To tackle severe geometric collisions and inter-object penetrations in compositional text-to-3D generation, CompSAG introduces a Gaussian Surface Offset mechanism that computes surface-normal repulsive forces via continuous proxy surfaces to shift penetrating 3D Gaussians, combined with relation-aware score distillation and normal rectification for penetration-free, high-fidelity multi-object scenes.

Background & Motivation

Compositional text-to-3D scene generation aims to synthesize multi-object 3D environments from text prompts with intricate spatial arrangements and entity interactions, playing an essential role in virtual reality, embodied robotics environments, and spatial intelligence simulations. Recently, explicit 3D Gaussian Splatting (3DGS) has rapidly emerged as a dominant alternative to Neural Radiance Fields (NeRF) due to its real-time differentiable rendering capabilities and flexible object-level spatial manipulation. However, 3DGS parameterizes scenes via unstructured discrete volumetric densities and spherical harmonic color fields, fundamentally lacking explicit physical surface boundaries.

In existing multi-object joint optimization pipelines, unconstrained optimization frequently causes severe inter-object penetrations and entangled representations. Prior collision mitigation strategies largely rely on coarse bounding-sphere constraints (such as tolerant collision loss) or simple point-distance penalties. These formulations merely suppress interior point concentrations without identifying where contact and collisions physically occur on the geometry. Consequently, practitioners face an intractable trade-off: objects either remain deeply embedded within one another (e.g., a rose buried inside an open diary or a watering can penetrating a plant pot), or are pushed unnaturally far apart with distorted rotations, disrupting intended textual contact relationships. Concurrently, standard Score Distillation Sampling (SDS) under full-scene supervision introduces cross-entity semantic leakage (such as wooden textures bleeding from a desk onto a globe) and coarse, spiky geometric artifacts.

Resolving this tension requires empowering discrete 3D Gaussians with explicit surface awareness to establish physically grounded, contact-guided repulsive forces. Core idea: maintain a continuous Signed Distance Function (SDF) proxy surface for each entity to localize penetrating primitives and project them onto pseudo-contact boundaries, applying aggregated surface-normal repulsive forces via momentum-based offsets while unifying relation-aware SDS and Gaussian normal rectification to completely eliminate penetrations, geometric spikes, and attribute confusion.

Method

Overall Architecture

CompSAG executes across three sequential stages: first, a layout-guided scene initialization driven by a fine-tuned LLM and feedforward 3D lifting, accompanied by Reference Depth Calibration (RDC) to rectify macroscopic camera and viewpoint discrepancies; second, a Gaussian Surface Offset (GSO) stage that uses lightweight implicit proxy surfaces to detect colliding Gaussians, project them onto pseudo-contact boundaries, and apply momentum-buffered repulsive forces along the separation axis; and third, a geometry-constrained appearance refinement stage that decomposes scenes into interacting entity pairs for relation-aware SDS distillation while enforcing surface adherence and normal alignment.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Scene Text Prompt"] --> B["Reference Depth Calibration & Scene Initialization<br/>Fine-tuned LLM Layout + SAM3D Lifting + PnP/Depth Alignment"]
    B --> C["Proxy Surface Collision Localization & Normal Projection<br/>Lightweight SDF Proxy + Zero-Level Set Detection + Orthogonal Projection"]
    C --> D["Momentum-Based Surface Offset & Repulsive Force Aggregation<br/>Pseudo-Contact Repulsion + Separation Axis Filtering + Momentum Position Update"]
    D --> E["Relation-Aware Distillation & Gaussian Surface Rectification<br/>Interacting Pair SDS + SDF Center Attraction + Thin-Axis Normal Alignment"]
    E --> F["Penetration-Free High-Fidelity Compositional 3D Gaussian Scene"]

Key Designs

1. Reference Depth Calibration & Scene Initialization: Establishing Blueprint-Aligned Collision-Reduced Geometry

Directly optimizing multi-object 3D scenes from ungrounded text prompts leads to severe multi-face Janus artifacts and disordered spatial scales. CompSAG anchors 3D generation onto an explicit 2D visual blueprint. Addressing the spatial reasoning limitations of off-the-shelf language models, the framework fine-tunes a Qwen2.5-3B model on the Synthetic Visual Genome (SVG) dataset using LoRA, directing it to output structured CSS-formatted 2D bounding boxes that guide box-conditioned diffusion models to synthesize a high-fidelity reference blueprint \(I_{ref}\). Each entity is then lifted into an initial 3D Gaussian primitive set \(\mathcal{G}_i = \{\mathbf{p}_k, \mathbf{s}_k, \mathbf{q}_k, \mathbf{c}_k, \alpha_k\}\) using feedforward SAM3D.

Because feedforward lifting introduces per-object pose discrepancies, the initial 3D assembly exhibits substantial depth misalignment. The system samples views around the analytic camera, matches 2D-3D keypoints via SuperGlue, and refines the camera extrinsics \((R, \mathbf{t})\) via Perspective-n-Point (PnP). Rendered object depth maps \(D_{ren}\) are then aligned with the monocular metric depth \(D_{ref}\) estimated from \(I_{ref}\). For each object, the depth discrepancy \(\delta_i\) is computed over foreground masks: $\(\delta_i = \frac{1}{|\mathcal{M}_i^s|} \sum_{\mathbf{u} \in \mathcal{M}_i^s} s_{d,i} \cdot D_{ref}(\mathbf{u}) - \frac{1}{|\mathcal{M}_i^r|} \sum_{\mathbf{u} \in \mathcal{M}_i^r} D_{ren}(\mathbf{u})\)$ Each object's Gaussians are translated along the viewing ray by \(\delta_i\) and uniformly scaled by mask area ratios. Finally, a structural base object's canonical \(+Z\) axis is aligned with the world \(+Z\) axis and centered within \([-1, 1]^3\), providing a geometrically coherent macro layout for subsequent fine-grained collision resolution.

2. Proxy Surface Collision Localization & Normal Projection: Continuous Anchoring for Discrete Gaussians

Discrete 3D Gaussians represent continuous volumetric density fields without explicit geometric boundaries. Directly calculating interaction forces on noisy Gaussian centers leads to severe numerical instability and diverging optimizations. CompSAG overcomes this by maintaining a continuous Signed Distance Function (SDF) \(f_{\phi_i}: \mathbb{R}^3 \to \mathbb{R}\) parameterized by a lightweight MLP for each object \(o_i\). Supervised by SAM3D meshes under an Eikonal regularizer, the zero-level set \(\mathcal{S}_i = \{\mathbf{x} \mid f_{\phi_i}(\mathbf{x}) = 0\}\) defines a smooth geometric boundary proxy.

For any interacting pair of intruder \(o_i\) and obstacle \(o_j\), colliding Gaussians are identified as primitives penetrating the obstacle boundary: \(\mathcal{C}_{i \to j} = \{\mathbf{p}_i^{(m)} \in \mathcal{G}_i \mid f_{\phi_j}(\mathbf{p}_i^{(m)}) < 0\}\). To suppress high-frequency point-cloud noise, each colliding center \(\mathbf{p}_i^{(m)}\) is orthogonally projected onto its own proxy surface along the local surface normal: $\(\hat{\mathbf{p}}_i^{(m)} = \mathbf{p}_i^{(m)} - f_{\phi_i}(\mathbf{p}_i^{(m)}) \frac{\nabla f_{\phi_i}(\mathbf{p}_i^{(m)})}{\|\nabla f_{\phi_i}(\mathbf{p}_i^{(m)})\|}\)$ This projection rectifies unstructured Gaussians onto a smooth zero-level set, forming a well-defined pseudo-contact boundary that stabilizes subsequent repulsive force derivations.

3. Momentum-Based Surface Offset & Repulsive Force Aggregation: Simulating Physical Repulsion for Smooth Decoupling

Across the pseudo-contact boundary, for each projected point \(\hat{\mathbf{p}}_i^{(m)}\) with closest obstacle point \(\hat{\mathbf{p}}_{ij}^{(m)}\) on \(\mathcal{S}_j\), the repulsive force magnitude is defined by projecting the penetration vector onto the obstacle's outward surface normal. Mathematically, this evaluates directly to the obstacle's absolute signed distance value: \(|\mathbf{F}_i^{(m)}| = |f_{\phi_j}(\hat{\mathbf{p}}_i^{(m)})|\), providing a physical property where deeper penetrations generate stronger repulsive forces.

The force direction follows the inward surface normal of the intruder \(\mathbf{v}_i^{(m)} = -\mathbf{n}_i^{(m)}\). To prevent local concave geometry from erroneously pushing objects deeper, the system filters forces against the centroid separation axis \(\mathbf{d}_{ij} = \boldsymbol{\mu}_i - \boldsymbol{\mu}_j\), retaining only aligned directions to form the aggregated object-level force: $\(\mathbf{F}_i = \frac{1}{|\mathcal{C}_{i \to j}|} \sum_{\mathbf{p}_i^{(m)} \in \mathcal{C}_{i \to j}} \mathbf{1}\left(\mathbf{v}_i^{(m)} \cdot \mathbf{d}_{ij} > 0\right) |\mathbf{F}_i^{(m)}| \mathbf{v}_i^{(m)}\)$ To prevent rigid oscillations, updates are integrated through a momentum buffer \(\Delta \boldsymbol{\mu}_i = \beta \mathbf{F}_{prev, i} + (1 - \beta) \mathbf{F}_i\) (\(\beta = 0.3\)) under an ease-in-out learning rate schedule. As penetrations vanish, the depth-proportional forces decay naturally to zero, guaranteeing bounded and stable convergence without heuristic clipping, while grounding constraints zero out vertical forces for base-contact objects.

4. Relation-Aware Distillation & Gaussian Surface Rectification: Disentangling Semantics and Suppressing Spiky Artifacts

Once physical overlaps are decoupled, visual appearance and fine geometric structures must be refined. Applying global SDS optimization across full scenes causes diffusion attention confusion and semantic attribute leakage. CompSAG parses the input prompt into a scene graph using the fine-tuned LLM, isolating interacting subject-relation-object pairs. Each centered pair is rendered from sampled camera viewpoints \(\pi\) and supervised using a dedicated relational prompt \(y_{rel}\) under a relation-aware SDS objective \(\mathcal{L}_{\text{SDS}_{rel}}\).

To eliminate the surface concavities and spiky artifacts commonly produced by SDS, the framework introduces Gaussian Surface Rectification (GSR). This mechanism establishes bi-directional geometric coupling between explicit Gaussians and the implicit proxy surface: it pulls Gaussian centers toward the frozen SDF zero-level set via \(\mathcal{L}_{GS}^i = \sum_k |f_{\phi_i}(\mathbf{p}_i^{(k)})|\), and aligns each Gaussian's thinnest axis \(\mathbf{n}_{p_i}^{(k)}\) (the rotation column corresponding to the minimum scale factor) with the proxy surface normal: $\(\mathcal{L}_{norm}^i = \sum_{k=1}^{N_i} \left\| \frac{\nabla f_{\phi_i}(\mathbf{p}_i^{(k)})}{\|\nabla f_{\phi_i}(\mathbf{p}_i^{(k)})\|} \cdot \mathbf{n}_{p_i}^{(k)} - 1 \right\|^2\)$ The squared cosine formulation remains invariant to normal orientation. Alternating iterative optimization across relation-aware SDS (400 steps), SDF surface tracking (200 steps), and Gaussian rectification (400 steps) flattens Gaussians into smooth surface-conforming splats, suppressing geometric irregularities while regularizing SDS appearance distillation.

Key Experimental Results

Main Results

On multi-object compositional 3D generation benchmarks, CompSAG was evaluated against non-compositional baselines (ProlificDreamer, MVDream, GaussianDreamer) and state-of-the-art compositional frameworks (Set-the-Scene, GraphDreamer, CG3D, GALA3D, CompGS, Layout-your-3D). Metrics comprise runtime, CLIP semantic similarity, B-VQA for multi-object completeness and visual quality, GPT-CoT for spatial relation grounding, and Colliding Volume Ratio (CVR) evaluated over a \(256^3\) discretized voxel grid.

Method 3D Rep. Time CLIP B-VQA GPT-CoT CVR ↓
ProlificDreamer NeRF 240min 28.91 42.16 27.27 N/A
MVDream NeRF 30min 31.05 28.29 30.98 N/A
GaussianDreamer 3DGS 15min 28.24 31.87 31.06 N/A
Set-the-Scene NeRF 50min 29.43 41.20 37.74 0.265
GraphDreamer NeRF 120min 32.11 40.29 45.69 0.278
CG3D 3DGS 60min 31.56 46.19 43.25 0.147
GALA3D 3DGS 60min 32.63 50.48 52.03 0.205
CompGS 3DGS 70min 31.67 53.62 51.23 0.186
Layout-your-3D 3DGS 30min 31.79 52.57 50.93 0.103
CompSAG (Q) 3DGS 30min 32.72 56.17 53.69 0.042
CompSAG (Q*) 3DGS 30min 33.64 57.28 54.38 0.034
CompSAG (Q*+U) 3DGS 30min 33.87 58.11 56.64 0.038

Note: Q indicates base Qwen2.5-3B, Q* denotes the model fine-tuned on SVG layouts, and U represents human-adjusted reference layouts. Minor residual CVR (e.g., 0.034) reflects shared boundary voxels at finite grid resolutions during legitimate contact rather than perceptual interpenetration.

On the T3Bench multi-object benchmark, CompSAG also achieved leading scores in visual quality and user preference:

Method Quality ↑ Alignment ↑ Average ↑ User Study (1-5) ↑
Fantasia3D 22.7 14.3 18.5 2.42
LatentNeRF 21.7 19.5 20.6 3.17
Magic3D 26.6 24.8 25.7 3.46
ProlificDreamer 45.7 25.8 35.8 3.65
DreamGaussian 12.3 9.5 10.9 2.54
GaussianDreamer 38.8 30.2 34.5 3.21
VP3D 49.1 31.5 40.3 2.93
Set-the-Scene 20.8 29.9 25.4 3.82
CompGS 54.2 37.9 46.1 4.15
CompSAG (Ours) 55.5 41.9 48.7 4.42

Ablation Study

To isolate the contribution of each module toward resolving physical intersections, progressive ablations were conducted across varying scene complexities (\(M=2, 3, 5\) objects) evaluating CVR:

Config \(M=2\) \(M=3\) \(M=5\) Average CVR ↓ Note
Baseline 0.214 0.291 0.317 0.276 Feedforward initialization via SAM3D lifting
+ RDC 0.172 0.260 0.278 0.239 Resolves camera and viewpoint depth offset
+ TCL 0.085 0.109 0.113 0.103 Layout-your-3D coarse bounding-sphere repulsion
+ GSO 0.028 0.033 0.040 0.034 Full continuous proxy surface repulsive offset

Ablation on appearance quality and geometric regularization measured via B-VQA:

Setting B-VQA ↑ Mechanism & Analysis
Full setting 57.28 Complete integration of GSO, GSR, and relation-aware SDS
\(-\)GSO 50.33 6.95-point drop; persistent collisions cause entangled appearance during joint distillation
\(-\)GSR 52.19 5.09-point drop; lack of normal alignment leads to spiky surfaces and hollow artifacts
\(-\)SDSrel 54.47 2.81-point drop; unstructured full-scene SDS causes cross-object texture bleeding

Key Findings

  • Surface-normal repulsion is the primary driver for collision removal: Compared to bounding-sphere heuristics (TCL achieving 0.103 CVR), GSO drives collision ratios down to 0.034, eliminating structural entanglement and yielding a substantial 6.95-point boost in B-VQA visual quality.
  • Decoupled pipeline yields high computational efficiency: Resolving collisions (text-to-image, SAM3D lifting, RDC, and GSO) takes only ~8 minutes, leaving ~22 minutes for appearance refinement. Due to the strict Gaussian count cap (800k), scaling from 2 to 5 objects increases total runtime only modestly from 27 to 34 minutes.
  • Fine-tuned layout LLMs bridge 2D-to-3D spatial reasoning: Using the SVG-fine-tuned Q* model lifts GPT-CoT from 53.69 to 54.38 and improves CLIP scores by nearly 1.0, proving that structured CSS-based grounding successfully reduces multi-object hallucination.

Highlights & Insights

  • Discrete Gaussian Projection onto Continuous Implicit Proxy Surfaces: Evaluating repulsive forces directly on discrete point sets is ill-posed and noisy. CompSAG elegantly maps colliding Gaussians onto continuous SDF zero-level sets, translating penetration depth directly into the SDF magnitude to enable stable, differentiable physical repulsion.
  • Gaussian Minimum-Axis Normal Alignment Regularization: To eliminate the spiky artifacts inherent to SDS optimization, aligning the Gaussian covariance matrix's thinnest axis with the surface normal ensures that Gaussians flatten into tight, surface-hugging splats, greatly improving surface regularity and mesh extractability.
  • Dual Operating Modes for Flexible Deployment: Applications requiring rapid prototyping can terminate after the GSO stage (~8 minutes) to obtain clean, penetration-free 3D layouts, providing an ideal starting geometry for downstream simulation pipelines.

Limitations & Future Work

  • Primitive Count Ceiling in Dense Scenes: To maintain strict memory and compute bounds, total scene Gaussians are capped at 800k, which restricts per-object geometric and textural detail in scenes containing more than five dense objects.
  • Force Cancellation in Symmetric Collision Geometries: Because GSO relies on projected surface normals, symmetric configurations (such as two objects compressing a central object from opposite sides) can cause repulsive vectors to cancel out, leaving central penetrations unresolved.
  • Rigid-Body Assumption and Lack of Physical Dynamics: All entities are modeled as rigid objects, ignoring elastoplastic deformations (e.g., soft plush toys deforming against a sofa) and potential floating artifacts without gravity. Coupling CompSAG's collision-free layouts with differentiable physics solvers (e.g., MPM/FEM) presents a promising avenue for dynamic 4D generation.
  • vs Set-the-Scene / GraphDreamer: Early NeRF-based compositional pipelines require prolonged training times (e.g., 120 minutes for GraphDreamer) and lack surface boundaries, leaving severe interpenetrations unresolved; CompSAG leverages 3DGS and SDF proxies to deliver penetration-free generation within 30 minutes with nearly an order-of-magnitude lower CVR.
  • vs Layout-your-3D / CompGS: Existing compositional 3DGS methods rely on coarse bounding spheres or unconstrained joint optimization, resulting in artificial repulsions or deformed shapes; CompSAG's proxy surface accurately identifies local contact regions, maintaining natural contact boundaries.
  • vs SuGaR / GS-Pull: While SuGaR and GS-Pull pioneered SDF-Gaussian alignment for single-object reconstruction, CompSAG extends surface-aware Gaussians to multi-entity collision detection and physical repulsion.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant integration of continuous SDF surface proxies with discrete 3DGS to model physically grounded repulsive offsets.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across CVR, T3Bench, B-VQA, GPT-CoT, progressive ablations, and runtime breakdowns.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulations, clear pipeline visualizations, and transparent technical insights.
  • Value: ⭐⭐⭐⭐⭐ Establishes a standard benchmark for penetration-free compositional 3D generation and offers clean initial states for embodied physics simulators.