Skip to content

Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/NWUzhouwei/PDM
Area: 3D Vision
Keywords: Single-View 3D Reconstruction, Diffusion Models, State-Space Models, Mamba, Limited-Data Regime

TL;DR

To tackle ill-posed single-view 3D reconstruction and memory bottlenecks under data scarcity, this paper proposes Point Diffusion Mamba (PDM), unifying local geometric aggregation and dual-directional Mamba diffusion with a hierarchical feature integration network and dynamic weighted sampling, achieving state-of-the-art geometry recovery with only 0.39 GB GPU memory.

Background & Motivation

Recovering complete, fine-grained 3D geometric shapes from a single 2D image observation is an intrinsically ill-posed inverse problem in computer vision. Because lifting 2D projections to 3D space suffers from inherent depth ambiguities, a model must not only capture visible surface details but also hallucinate coherent structures across severely occluded regions. Denoising Diffusion Probabilistic Models (DDPM) have recently demonstrated remarkable generative capacity in 3D shape synthesis and surface completion. However, existing diffusion-based 3D reconstruction methods are notoriously data-hungry, requiring vast amounts of paired 3D CAD ground-truth models. In real-world applications where 3D annotations are difficult and expensive to acquire, 3D reconstruction under the severely constrained data-scarce regime remains a critical yet underexplored bottleneck.

Existing diffusion-based 3D reconstruction architectures suffer from fundamental trade-offs between computational tractability and geometric fidelity. MLP-based architectures struggle to characterize complex topological manifolds. 3D voxel-based convolutional networks (such as PC2 and Buildiff) scale with prohibitive \(O(n^3)\) cubic computational and memory complexity, constraining voxel resolution and missing long-range contextual relationships. Conversely, Transformer-based diffusion backbones (such as DiT-3D) incur an \(O(n^2)\) quadratic attention overhead when processing long point patch sequences. While State-Space Models (SSMs / Mamba) deliver linear computational complexity in sequence processing, directly adapting them to 3D diffusion reconstruction encounters two primary hurdles: the unordered, non-Euclidean nature of point clouds creates pseudo-sequential dependencies that fail to capture fine local geometry, and high-level abstract tokens extracted by Mamba blocks cannot directly support the precise, per-point noise displacement prediction required by diffusion denoising.

This paper's angle of attack is to break the computational bottleneck of quadratic attention and voxelization by systematically decoupling local neighborhood geometry extraction from dual-directional global sequence modeling, while harnessing generative priors to counteract data sparsity. Core idea: develop a unified diffusion-state-space framework combining local geometric aggregation, dual-directional Mamba sequence modeling, and hierarchical point feature propagation, coupled with dynamic weighted sampling between generative priors and reconstruction predictions to achieve high-fidelity 3D reconstruction under strict data scarcity.

Method

Overall Architecture

The end-to-end objective of PDM is to reconstruct a clean 3D point cloud from a randomly perturbed Gaussian noise point cloud \(X_t \in \mathbb{R}^{N \times 3}\), conditioned on a single 2D image \(J\). First, a pre-trained ViT-32 extracts 2D multi-scale semantic features, which are back-projected onto the noisy point cloud via a differentiable rasterization function, yielding an enhanced point cloud \(X_f^t \in \mathbb{R}^{N \times (3+d_y)}\) that couples spatial coordinates with projected image cues. Farthest Point Sampling (FPS) selects \(s\) geometric center points, and K-Nearest Neighbors (KNN) groups neighboring points into \(s\) localized patches, which are encoded into patch tokens \(F \in \mathbb{R}^{s \times d}\) through lightweight convolutions and max-pooling. Within the core denoising stage, features pass through the Local Geometric Aggregation (LGA) module to distill fine-grained neighborhood geometry, enter the Dual-Directional State Space Module (DDSM) with diffusion timestep adaptive normalization to model global contextual structure, and are upsampled by the Hierarchical Feature Integration Network (HFINet) to compute precise per-point noise predictions. Finally, during reverse sampling, a Dynamic Weighted Sampling (DWS) strategy adaptively blends unconditional generative priors with reconstruction predictions.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Single-view image J and noisy point cloud Xt"] --> B["Rasterized Feature Projection & FPS-KNN Patch Partition"]
    B --> C["1. Local Geometric Aggregation (LGA)<br/>K-Norm neighborhood standardization + K-Pool dual pooling"]
    C --> D["2. Dual-Directional State Space Module & Timestep Conditioning (DDSM & Time Conditioning)<br/>Bi-directional scan + diffusion timestep adaptive modulation"]
    D --> E["3. Hierarchical Feature Integration Network (HFINet)<br/>Multi-scale token aggregation + point feature propagation to per-point space"]
    E --> F["4. Dynamic Weighted Sampling Strategy (DWS)<br/>Adaptive geometry-consistent fusion of generative prior & reconstruction"]
    F --> G["High-Fidelity Reconstructed 3D Point Cloud"]

Key Designs

1. Local Geometric Aggregation (LGA): Standardizing local neighborhood scale and balancing prominent details with surface smoothness Standard Mamba sequence modeling on discrete point tokens struggles to capture delicate local geometric manifolds. The LGA module resolves this limitation through the coordinated operation of K-Norm and K-Pool. K-Norm constructs a KNN neighborhood graph (of size \(n\)) around each center token \(F_i\), computing the local mean \(\mu_i\) and standard deviation \(\sigma_i\) to normalize neighbor features, thereby eliminating distortion caused by density variations. The normalized neighbor deviations are concatenated with \(F_i\) to produce normalized local feature descriptors \(G_{ij} \in \mathbb{R}^{2d}\). K-Pool then aggregates these neighborhood features through hybrid pooling: $\(H_i = \phi\left(\max_{j \in \mathcal{N}(i)} G_{ij} + \frac{1}{n}\sum_{j \in \mathcal{N}(i)} G_{ij}\right)\)$ Here, max-pooling preserves high-frequency structural elements such as sharp corners, thin boundaries, and contours, while mean-pooling suppresses stochastic high-frequency noise and stabilizes smooth planar surfaces. This unified operation establishes robust local geometric embeddings \(H \in \mathbb{R}^{s \times d}\) prior to global sequence modeling.

2. Dual-Directional State Space Module & Timestep Conditioning (DDSM & Time Conditioning): Eliminating pseudo-order biases and adaptively regulating diffusion dynamics Standard Mamba models rely on causal 1D sequential scans, which impose artificial directional dependencies when applied to unordered 3D point clouds. To resolve this, the Dual-Directional State Space Module (DDSM) models bidirectional context by processing both the forward sequence copy \(H_c\) and the channel-reversed sequence \(H_r\): $\(\text{DDSM}(H) = H + \zeta(H_c) + \eta(H_r)\)$ This bidirectional state propagation neutralizes orientation artifacts and captures omnidirectional spatial relations across the point cloud. To dynamically guide diffusion denoising across different noise levels, time-conditional adaptive normalization (adaLN) generates scaling parameter \(\beta(t)\), shifting parameter \(\gamma(t)\), and gating parameter \(\alpha(t)\) from the diffusion timestep embedding \(t\): $\(H_{mb} = H + \alpha(t) \cdot \text{DDSM}(H \cdot (1 + \beta(t)) + \gamma(t))\)$ This adaptive modulation dynamically scales feature representations according to the denoising phase (macro structure formation early on versus micro refinement in later steps), effectively curbing error accumulation in multi-step sampling.

3. Hierarchical Feature Integration Network (HFINet): Bridging sparse high-level semantic tokens to dense per-point noise predictions While Mamba blocks output rich high-level semantic tokens at \(s\) downsampled patch centers, diffusion-based point cloud reconstruction demands precise 3D vector predictions \(\epsilon_\theta\) for all \(N\) perturbed points. HFINet resolves this spatial-semantic resolution gap. Multi-level intermediate features \(H'_{mb} = \{H_{mb1}, \dots, H_{mbM}\}\) extracted across Mamba layers are concatenated and compressed via max-pooling and mean-pooling to derive a global context vector \(H_g \in \mathbb{R}^{2d'}\). Concurrently, an inverse-distance weighted Point Feature Propagation operator \(\tau\) interpolates multi-layer patch features back into the full per-point coordinate space of \(X_f^t\), generating localized feature field \(H_l = \tau(H', X_f^t) \in \mathbb{R}^{N \times d''}\). Finally, per-point local features \(H_l\) and global context \(H_g\) are concatenated and passed through lightweight convolutional layers to output dense point displacements: $\(X_{\text{pred}} = \phi(\text{concat}(H_l, H_g))\)$ This hierarchical aggregation prevents structural discretization and hole artifacts inherent to direct token decoding, markedly elevating reconstruction precision.

4. Dynamic Weighted Sampling Strategy (DWS): Adaptively fusing generative priors and reconstruction trajectories via geometric consistency Under extreme data scarcity (such as 10% or fewer training samples), single-view reconstruction models tend to overfit visible regions and fracture across occluded geometry. To address this, an identical unconditional 3D generative Mamba model is pre-trained to capture comprehensive geometric priors. During reverse diffusion sampling, guided fusion between the reconstruction model's point cloud \(X_r\) and generative prior \(X_g\) is performed every 32 timesteps. Instead of static averaging, DWS adaptively computes per-point fusion weights based on localized geometric consistency: $\(w_j = \frac{1}{1 + \exp\left(-\delta \|x_{g,j} - x_{r,j}\|_2^2\right)}\)$ $\(\hat{X} = \{w_j x_{g,j} + (1 - w_j) x_{r,j}\}_{j=1}^N\)$ When both trajectories agree closely (indicating confident reconstruction aligned with generative distribution), weights prioritize reconstruction outputs to preserve faithful alignment with the 2D conditioning image. Conversely, where reconstruction models diverge or collapse in unseen regions, the Sigmoid formulation dynamically shifts weight toward generative priors, maintaining smooth surface continuity and topological completeness.

Loss & Training

The framework minimizes the standard diffusion denoising mean squared error objective: $\(\mathcal{L} = \mathbb{E}_{t, X_0, \epsilon}\left[\|\epsilon - \epsilon_\theta(X_t, t, J)\|^2\right]\)$ where \(\epsilon_\theta\) denotes the noise predicted by PDM. The unconditional generative model and the single-view reconstruction model are trained completely independently for 140,000 iterations without shared parameters. During sampling, core backbone layers remain frozen, and guided trajectory fusion is executed via DWS with sensitivity coefficient \(\delta\). The entire training workflow executes on a single commodity NVIDIA GeForce RTX 3090 GPU.

Key Experimental Results

Main Results

PDM is benchmarked on the synthetic ShapeNet-R2N2 dataset (Chair, Airplane, Car) and the real-world Pix3D dataset (Chair, Sofa, Table) using Chamfer Distance (CD \(\times 10^4 \downarrow\)) and [email protected] (F1 \(\uparrow\)).

The table below reports quantitative comparisons across varying data availability regimes (10% data scarcity, 50% medium scale, 100% full scale) on ShapeNet-R2N2:

Method Venue/Source Chair (10%) Chair (50%) Chair (100%) Airplane (10%) Airplane (50%) Airplane (100%) Car (10%) Car (50%) Car (100%)
PC2 CVPR 2023 97.25 / 0.393 73.58 / 0.437 65.57 / 0.464 88.00 / 0.605 76.39 / 0.628 65.97 / 0.655 64.99 / 0.524 62.59 / 0.542 64.36 / 0.547
CCD-3DR arXiv 2023 89.79 / 0.418 63.13 / 0.474 58.47 / 0.498 81.29 / 0.612 72.46 / 0.635 62.77 / 0.651 63.13 / 0.531 62.25 / 0.550 61.88 / 0.562
BDM-M CVPR 2024 94.94 / 0.395 71.56 / 0.446 64.48 / 0.468 87.75 / 0.604 73.19 / 0.629 65.16 / 0.653 63.53 / 0.524 60.71 / 0.549 64.16 / 0.554
BDM-B CVPR 2024 94.67 / 0.410 69.99 / 0.463 64.21 / 0.485 83.62 / 0.612 68.66 / 0.641 59.04 / 0.660 60.48 / 0.539 62.58 / 0.554 65.85 / 0.559
MESC-3D CVPR 2025 101.51 / 0.381 73.47 / 0.427 65.69 / 0.458 74.31 / 0.611 51.28 / 0.619 50.54 / 0.714 56.28 / 0.522 44.48 / 0.532 51.99 / 0.602
PDM (Ours) ECCV 2026 82.41 / 0.419 68.85 / 0.465 62.14 / 0.488 60.64 / 0.614 50.14 / 0.663 48.66 / 0.719 55.23 / 0.542 55.53 / 0.558 51.54 / 0.607

The table below demonstrates zero-shot real-world transfer results on Pix3D (CD \(\downarrow\) / F1 \(\uparrow\)):

Method Venue/Source Chair (CD↓ / F1↑) Sofa (CD↓ / F1↑) Table (CD↓ / F1↑)
PC2 CVPR 2023 115.94 / 0.443 47.17 / 0.445 202.77 / 0.397
CCD-3DR arXiv 2023 111.42 / 0.456 44.91 / 0.450 196.28 / 0.418
BDM-M CVPR 2024 113.40 / 0.449 44.50 / 0.451 202.08 / 0.413
BDM-B CVPR 2024 110.60 / 0.455 45.05 / 0.455 186.46 / 0.429
MESC-3D CVPR 2025 91.36 / 0.370 41.98 / 0.284 206.16 / 0.306
PDM (Ours) ECCV 2026 79.28 / 0.499 41.43 / 0.463 184.35 / 0.422

The table below compares inference computational efficiency and resource footprints (Batch size = 1):

Method Parameters (M) Runtime (s) GPU Memory (GB)
PC2 / CCD-3DR 47.41 48.46 1.73
BDM-M 74.82 52.47 2.01
BDM-B 73.78 50.81 1.93
MESC-3D 74.84 34.05 0.86
PDM (Ours) 56.34 20.73 0.39

Ablation Study

Ablations are conducted on the Pix3D Chair dataset. The table below evaluates the individual contribution of architectural modules and the impact of input patch sequence lengths:

Architectural Configuration F1-Score↑ CD↓ Note Sequence Length F1-Score↑ CD↓
Full Model 0.499 79.28 Full PDM model 128 (Default) 0.499 79.28
w/o LGA 0.478 91.14 Removes local geometric aggregation; F1 drops by 0.021 32 0.485 84.31
w/o HFINet 0.435 126.57 Removes hierarchical integration; CD degrades severely by 47.29 64 0.474 91.15
Self-Attention 0.481 85.46 Replaces Mamba with standard self-attention; worse speed & accuracy 256 0.491 80.57
One-SSM 0.474 88.72 Unidirectional SSM fails to mitigate point cloud disorder bias 384 0.487 82.29

The table below analyzes model scaling behaviors and sampling strategy effectiveness:

Model Variant Architecture Specification F1-Score↑ CD↓ Sampling Strategy F1-Score↑ CD↓ Note
PDM-S 9 layers, 192 hidden dim 0.479 90.87 Direct Sampling 0.483 85.27 Vanilla unguided sampling
PDM-B 12 layers, 384 hidden dim 0.499 79.28 BDM Sampling 0.492 80.23 Bayesian prior fusion baseline
PDM-L 18 layers, 768 hidden dim 0.492 80.23 Our DWS 0.499 79.28 Adaptive spatial-consistency weights

Key Findings

  • Unquestioned superiority under data scarcity: In the critical 10% data regime on ShapeNet Chairs, PDM reduces CD to 82.41, outperforming second-best CCD-3DR (89.79) by 9% and decimating MESC-3D (101.51) by 23%. Under extreme stress tests (1% and 5% data), baseline methods experience catastrophic topological breakdown, whereas PDM reliably maintains global shape integrity.
  • HFINet propagation is the architectural cornerstone: Removing HFINet causes the largest performance drop across all ablations, causing CD error to spike from 79.28 to 126.57. This confirms that abstract Mamba tokens alone cannot supply the high-resolution geometric signals required for per-point diffusion noise prediction.
  • Bidirectional scanning counteracts pseudo-sequential order: Reverting DDSM to unidirectional scanning (One-SSM) degrades F1 by 0.025, validating that bidirectional state propagation is crucial for neutralizing spurious sequence biases in unordered 3D point sets.
  • Unprecedented memory and runtime reduction: PDM requires only 0.39 GB VRAM—a 55% reduction compared to MESC-3D (0.86 GB) and an 80% reduction compared to BDM-M (2.01 GB). Its inference latency of 20.73s is nearly \(2.5\times\) faster than PC2 (48.46s).

Highlights & Insights

  • First comprehensive adaptation of Mamba to 3D diffusion reconstruction: Transcends conventional discriminative uses of SSMs by tailoring bidirectional scanning, timestep conditioning, and multi-scale feature propagation for generative 3D diffusion tasks.
  • Geometry-consistent generative prior arbitration: DWS effectively tames severe overfitting in low-data regimes by dynamically modulating the fusion ratio between reconstruction and generative trajectories using local Euclidean distance.
  • Voxel-free and attention-free lightweight 3D pipeline: By eliminating cubic voxel grids and quadratic attention matrices, PDM provides a practical blueprint for running high-fidelity 3D generative diffusion models on commodity edge hardware.

Limitations & Future Work

  • Independent dual-model training footprint: The unconditional generative prior model and the reconstruction model are trained separately, requiring two distinct training runs and joint model memory during inference.
  • Fine surface fidelity on delicate geometries: Discretized 4096-point representations cannot fully depict continuous ultra-thin manifolds or micro-textures. Future work could integrate PDM with continuous 3D Gaussian Splatting or neural SDF implicit fields.
  • End-to-end distillation integration: Exploring offline knowledge distillation to embed generative structural priors directly into the reconstruction backbone would eliminate the dual-model inference requirement during DWS.
  • vs PC2 (CVPR 2023): PC2 relies on voxelization and 3D convolutions with high memory overhead (1.73 GB VRAM). PDM replaces voxelization with FPS-KNN patch tokens and Mamba modeling, slashing VRAM to 0.39 GB while improving 10% low-data Chair CD by 14.84.
  • vs BDM (CVPR 2024): BDM couples top-down prior and bottom-up data-driven diffusion with quadratic attention backbones, leading to 50s+ inference latencies. PDM introduces linear bidirectional state-space modeling and fine-grained spatial DWS, reducing runtime to 20.73s.
  • vs MESC-3D (CVPR 2025): MESC-3D relies heavily on full-scale semantic text prompts and suffers severe overfitting in low-data regimes (Chair 10% CD of 101.51). PDM excels in data-scarce conditions, obtaining 82.41 CD on 10% data through its robust geometric priors.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering integration of local geometric aggregation, bidirectional Mamba, and adaptive prior-guided sampling for 3D diffusion.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Extensive multi-scale data experiments on ShapeNet and Pix3D with exhaustive ablation of modules, sequence length, and model scaling.
  • Writing Quality: ⭐⭐⭐⭐⭐ Crisp mathematical formulation, clear architectural illustrations, and self-contained empirical analysis.
  • Value: ⭐⭐⭐⭐⭐ Provides an efficient, highly practical blueprint for single-view 3D reconstruction under extreme data scarcity on commodity GPUs.