Skip to content

Mind2Cloud: EEG-to-Point Cloud Generation with Two-Granularity Diffusion Decoding

Conference: ECCV 2026
Paper: ECCV Paper
Code: https://github.com/duasoi/Mind2Cloud
Area: Medical Imaging
Keywords: brain-computer interface, EEG decoding, 3D point cloud generation, diffusion models, two-granularity decoding

TL;DR

Addressing the evolving semantic discrepancy between brain neural representations and diffusion denoising trajectories, Mind2Cloud introduces a two-granularity point cloud decoder coupling a global Transformer with a local PVCNN branch via a timestep-adaptive fusion mask and adversarial refinement, substantially outperforming prior art in geometric fidelity and semantic alignment at 2048-point resolution across all 12 subjects.

Background & Motivation

Reconstructing 3D object shapes from non-invasive electroencephalogram (EEG) signals represents a fundamental frontier in neural decoding and interactive brain-computer interfaces (BCIs). Compared to functional magnetic resonance imaging (fMRI), which offers high spatial resolution but suffers from excessive cost, immobility, and poor temporal resolution, EEG provides millisecond-level temporal precision, affordability, and practical portability. However, decoding neural oscillations into coherent 3D geometry is severely challenged by the substantial domain gap between noisy, low-dimensional scalp potentials and irregular, continuous 3D coordinate spaces. The pioneer framework Neuro-3D demonstrated the feasibility of EEG-to-3D shape generation using static-dynamic EEG encoders paired with a diffusion-based generator. Nonetheless, prior approaches adopt a uniform decoding architecture throughout the reverse diffusion trajectory, neglecting the dynamically shifting semantic requirements of both brain neural signals and diffusion generation.

From a cognitive neuroscience perspective, EEG signals convey visual stimuli across distinct hierarchical scales: low-frequency neural rhythms predominantly encode coarse object-level semantics and holistic categories, whereas high-frequency components correspond to localized contours, edges, and surface details. This hierarchy naturally mirrors the coarse-to-fine progression of diffusion generation. During early reverse steps under heavy Gaussian noise, the model faces acute uncertainty and must synthesize global structural topology; during later steps, the priority shifts toward refining fine geometric surfaces and local curvatures. Fixed-granularity decoders struggle in both regimes, as they lack sufficient long-range relational reasoning in early stages and fail to enforce local manifold consistency in late stages.

The core motivation of this work is to explicitly synchronize the multi-scale abstraction of EEG signals with the coarse-to-fine timeline of diffusion denoising. Core idea: introduce a timestep-aware Two-Granularity point cloud diffusion decoder that deploys a global Transformer branch for coarse structural synthesis during early high-uncertainty denoising, smoothly shifts influence to a local Point-Voxel CNN (PVCNN) branch for detail sculpting during later stages via a learnable dynamic fusion mask, and applies a point cloud discriminator for adversarial refinement.

Method

Overall Architecture

The Mind2Cloud framework is an end-to-end conditional 3D generative pipeline comprising three coordinated modules: the Two-Stream EEG Encoder, the Two-Granularity Diffusion Decoder, and the Adversarial Refinement Module.

Initially, subject-specific EEG signals evoked by static images and rotating video sequences are passed through temporal Transformer backbones, merged via cross-attention, and processed by 1D spatial convolutions to model cortical channel topography. The resulting latent EEG representation is mapped into the shared CLIP visual embedding space using InfoNCE and MSE objectives. Next, this neural embedding conditions the two-granularity diffusion decoder to iteratively denoise a noisy point cloud. The decoder leverages an 8-layer Transformer branch endowed with learned spatial interaction matrices for global structural cohesion and a multi-stage PVCNN branch for local geometric feature extraction, coordinated dynamically by a timestep-dependent channel mask. Finally, a 1D residual convolutional discriminator distinguishes generated point clouds from ground-truth shapes to refine surface manifolds.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input: Static and Dynamic EEG Signals + Noisy Point Cloud $X_t$"] --> Enc["Two-Stream EEG Encoder: Cross-Attention Fusion + Spatial Convolutions + CLIP Alignment"]
    Enc --> TF["Global Transformer Branch: Position-Aware Spatial Attention for Coarse Structure"]
    Enc --> PV["Local PVCNN Branch: Point-Voxel Set Abstraction for Local Manifold Details"]
    TF --> Gate["Time-Aware Feature Fusion: Timestep-conditioned Learnable Mask $M_t$"]
    PV --> Gate
    Gate --> Noise["Noise Residual Prediction and Reverse Denoising Steps"]
    Noise --> Adv["Adversarial Refinement: 1D Residual Convolutional Discriminator Supervision"]
    Adv --> Out["Output: High-Fidelity 3D Point Cloud (2048 points)"]

Key Designs

1. Two-stream EEG Encoder with Multi-Modal Contrastive Alignment: Capturing Temporal Dynamics and Scalp Topography

EEG signals exhibit substantial noise and inter-channel functional dependencies. The encoder processes preprocessed static responses \(\mathbf{x}_s \in \mathbb{R}^{C \times T_s}\) and dynamic rotation responses \(\mathbf{x}_d \in \mathbb{R}^{C \times T_d}\) with dual Transformer blocks to obtain temporal features \(\mathbf{z}_{\text{stc}}\) and \(\mathbf{z}_{\text{dyn}}\). A cross-attention mechanism treats static features as queries and dynamic features as key-value pairs: $$ \mathbf{z}{\text{fuse}} = \text{CrossAttn}(Q\mathbf{z}}}, K\mathbf{z{\text{dyn}}, V\mathbf{z}) $$ Subsequently, residual 1D spatial convolutions model inter-electrode dependencies across scalp regions (frontal, central, temporal, and occipital lobes) to yield a compact neural embedding in }\(\mathbb{R}^{1024}\). To bridge the modality divide, the encoder is supervised with a symmetric InfoNCE loss \(\mathcal{L}_{\text{clip}}\) alongside an MSE regression term \(\mathcal{L}_{\text{img}}\) against CLIP visual embeddings, imbuing the neural representation with strong geometric and category discriminability.

2. Position-Aware Global Transformer Branch: Anchoring Coarse Structure Under High Uncertainty

During early diffusion steps where coordinate coordinates are corrupted by heavy noise, localized neighborhood operators frequently collapse. The global branch operates on latent point features \(\hat{\mathbf{X}}_t \in \mathbb{R}^{M \times d}\), normalizes them via Dual PatchNorm and an MLP into token embeddings \(\mathbf{T}_0 \in \mathbb{R}^{M \times D}\), and augments them with continuous 3D positional embeddings \(\text{Pemb} \in \mathbb{R}^{M \times D}\). To prevent spatial geometry from attenuating in deeper layers, a learned spatial interaction matrix \(\mathbf{H}\) is constructed via projection matrix \(\mathbf{W}_p \in \mathbb{R}^{D \times Z}\): $$ \mathbf{H} = \text{Softmax}\left((\text{Pemb}\mathbf{W}_p)(\text{Pemb}\mathbf{W}_p)^\top\right) $$ Across \(L=8\) Transformer layers, this structural prior modulates self-attention via Hadamard product: $$ \mathbf{T}_l^* = \text{Softmax}\left(\frac{(\mathbf{T}_l\mathbf{W}_q)(\mathbf{T}_l\mathbf{W}_k)^\top \odot \mathbf{H}}{\sqrt{Z}}\right)(\mathbf{T}_l\mathbf{W}_v) $$ This design explicitly preserves spatial topological coherence across distant points, ensuring robust global object semantics and part arrangements early in the generation trajectory.

3. Hierarchical Point-Voxel CNN (PVCNN) Branch: Regularizing Manifolds and Refining Local Curvatures

Unstructured point sets provide fine coordinate granularity but lack structural regularity, whereas voxel grids provide uniform spatial neighbourhoods. The local branch employs multi-stage Set Abstraction combining point-wise MLPs (aggregating \(k\)-NN neighborhoods for micro-contours) and sparse voxel convolutions (imposing volumetric geometric priors and suppressing high-frequency spatial noise). Intermediate coordinates and activations are retained across downsampling and upsampled via Feature Propagation (FP) layers. This branch concentrates on capturing localized physical structures—such as edges, slender components, and curvature variations—yielding high-frequency geometric features \(\mathbf{F}_{\text{cnn}}\).

4. Timestep-Aware Dynamic Feature Fusion: Adaptive Scheduling Across Diffusion Trajectories

Rather than relying on static concatenation or fixed weighting, this module adaptively arbitrates between global semantics and local details across the diffusion steps \(t\). The timestep \(t\) is mapped via sinusoidal encodings and an MLP with a Sigmoid activation into a channel-wise fusion mask \(\mathbf{M}_t \in (0, 1)^D\). The blended feature representation is derived as: $$ \mathbf{F}{\text{out}} = \mathbf{M}_t \odot \text{Conv}(\mathbf{F}}}) + (1 - \mathbf{Mt) \odot \text{TF}(\mathbf{F}) $$ When }\(t\) is large, the network dynamically assigns higher weight to the global Transformer feature \(\mathbf{F}_{\text{tr}}\) to establish structural foundations; as \(t \to 0\), the gating mechanism shifts capacity toward the PVCNN feature \(\mathbf{F}_{\text{cnn}}\) to finalize local surface fidelity.

5. Adversarial Refinement Module: Enforcing Coherent Point Cloud Surface Distributions

Standard mean-squared-error diffusion objectives can yield over-smoothed surfaces and floating outlier points. To penalize unrealistic coordinate scatter, a lightweight discriminator \(D\) composed of 4 stacked residual 1D convolutional layers, adaptive global max-pooling, and a binary classifier is incorporated. It is supervised using binary cross-entropy on generated point clouds \(\tilde{\mathbf{X}}\) and real ground-truth shapes \(\mathbf{X}_0\): $$ \mathcal{L}_G^{\text{adv}} = \text{BCE}(D(\tilde{\mathbf{X}}), 1) $$ Adversarial feedback provides gradient signals that enforce sharp surface continuity, part connectivity, and natural spatial compactness.

Loss & Training

The complete framework is optimized end-to-end. The EEG encoder is constrained by \(\mathcal{L}_{\text{EEG}} = \mathcal{L}_{\text{clip}} + \mathcal{L}_{\text{img}}\), while the diffusion backbone predicts noise residuals via \(\mathcal{L}_{\text{diff}} = \|\hat{\boldsymbol{\epsilon}} - \boldsymbol{\epsilon}\|^2\). These are combined using balance weight \(\alpha = 0.95\): $$ \mathcal{L}{\text{joint}} = \alpha \mathcal{L}}} + (1 - \alpha) \mathcal{L{\text{EEG}} $$ Incorporating the adversarial fine-tuning loss with \(\lambda_{\text{adv}} = 0.05\) yields the total objective: $$ \mathcal{L}}} = \mathcal{L{\text{joint}} + \lambda $$ Training was executed on a single NVIDIA RTX A6000 GPU over approximately 2.5 days using AdamW with an initial learning rate of }} \mathcal{L}_G^{\text{adv}\(1 \times 10^{-3}\). Ground-truth meshes were dynamically downsampled using Farthest Point Sampling (FPS) from 8192 points to 2048 points to optimize memory throughput while maintaining uniform spatial density.

Key Experimental Results

Main Results

Experiments were conducted on the benchmark EEG-3D dataset, encompassing 12 subjects observing 72 object classes from Objaverse. Evaluation metrics include Chamfer Distance (CD \(\times 10^2\)), Earth Mover's Distance (EMD \(\times 10^2\)), and \(N\)-way Top-\(k\) classification accuracy derived from an independently trained PointNet++ evaluator.

Table 1 presents geometric reconstruction fidelity comparisons: | Method | Subjects | CD ↓ (\(\times 10^2\)) | EMD ↓ (\(\times 10^2\)) | Note | | :--- | :--- | :---: | :---: | :--- | | Neuro-3D (CVPR 2025) | S01-S05 | 4.32 | 16.31 | Baseline with uniform diffusion decoder | | Mind2Cloud (Ours) | S01-S05 | 2.53 | 14.02 | CD reduced by 41.4% | | Mind2Cloud (Ours, Full Cohort) | S01-S12 | 2.65 | 14.43 | Averaged across all 12 participants |

Table 2 presents semantic reconstruction accuracy across 2-way and 10-way evaluation setups: | Method | Subjects | 2-way/top-1 (Avg) | 10-way/top-3 (Avg) | 2-way/top-1 (Top-1 of 5) | 10-way/top-3 (Top-1 of 5) | | :--- | :--- | :---: | :---: | :---: | :---: | | Neuro-3D (CVPR 2025) | S01-S05 | 50.80% | 31.72% | 68.33% | 53.89% | | Mind2Cloud (Ours) | S01-S05 | 52.92% | 32.47% | 70.83% | 56.67% | | Mind2Cloud (Ours, Full Cohort) | S01-S12 | 51.68% | 31.30% | 70.53% | 52.84% |

Ablation Study

Ablations on subjects S01-S05 validate the necessity of each architectural component: | Config | CD ↓ (\(\times 10^2\)) | EMD ↓ (\(\times 10^2\)) | 2-way/top-1 (Top-1 of 5) | 10-way/top-3 (Top-1 of 5) | Note | | :--- | :---: | :---: | :---: | :---: | :--- | | Full Model (Mind2Cloud) | 2.53 | 14.02 | 70.83% | 56.67% | Complete two-granularity + GAN model | | w/o Discriminator | 3.26 | 15.93 | 66.94% | 52.22% | Removal of adversarial refinement | | w/o Two-Granularity | 3.20 | 15.93 | 67.22% | 51.95% | Single uniform decoder baseline | | w/o PVCNN | 3.71 | 17.10 | 50.56% | 37.50% | Transformer-only, severe detail collapse | | Fixed Fusion | 3.00 | 15.42 | 66.39% | 52.22% | Static equal-weight branch blending |

Key Findings

  1. Critical Role of PVCNN Local Geometry Branch: Removing the PVCNN branch leads to catastrophic degradation; CD surges from 2.53 to 3.71 and the 10-way Top-1 accuracy collapses from 56.67% to 37.50%. This demonstrates that global self-attention alone cannot regularize irregular point set manifolds without localized spatial convolutions.
  2. Dynamic Timestep Fusion Outperforms Static Stacking: Replacing the learnable timestep mask with a fixed 50/50 fusion increases CD by 0.47 and reduces top-1 accuracy by over 4%, confirming that shifting computational focus from global topology to localized detail across the denoising timeline is foundational to performance gains.
  3. Correlation with Cortical Activation Topography: Scalp activation analysis demonstrates that high-performing participants (e.g., Sub01 with 72.22% top-1 accuracy) exhibit central suppression combined with prominent occipital-temporal activation, reflecting structured visual engagement. The overall 12-subject average (CD 2.65, Top-1 70.53%) closely mirrors the initial 5-subject benchmark, confirming cross-subject generalizability.

Highlights & Insights

  • Cognitive-Generative Trajectory Alignment: Elegantly pairs neuroscientific hypotheses of hierarchical EEG frequency encoding with the coarse-to-fine nature of reverse diffusion, providing a principled foundation for multimodal neural decoding.
  • Resource-Efficient 2048-Point Generation: Replaces computationally heavy 8192-point formulations with an efficient 2048-point representation while delivering superior structural continuity, crisper contours, and intact part connectivity on intricate shapes (e.g., helicopters, skateboards).
  • Persistent Positional Interaction in Latent Space: The explicit outer-product spatial interaction matrix \(\mathbf{H}\) ensures continuous coordinate context across all 8 Transformer layers without decaying through residual additions.

Limitations & Future Work

  • Lack of Color and Photorealistic Textures: Mind2Cloud currently synthesizes geometric point coordinates only, lacking RGB color, reflectivity, or surface material properties. Integrating 3D Gaussian Splatting or neural radiance field priors represents an important future step.
  • Inter-Subject Neural Domain Shift: Individual variations in skull thickness, electrode impedances, and perceptual strategies cause minor performance variances across participants. Incorporating unsupervised domain adaptation or subject-conditioned meta-learning could further elevate cross-subject robustness.
  • vs. Neuro-3D (CVPR 2025): While Neuro-3D introduced the first EEG-3D benchmark, its generator relies on a static diffusion backbone that struggles with thin components and parts disconnection. Mind2Cloud introduces timestep-adaptive two-granularity decoding, cutting Chamfer Distance by 41.4% at one-quarter of the point resolution.
  • vs. TIGER (CVPR 2024): TIGER investigated time-varying convolutional-transformer hybrids for unconditional 3D generation. Mind2Cloud adapts this insight to the noisy, cross-modal neural decoding domain, integrating contrastive multimodal alignment and adversarial point refinement.

Rating

  • Novelty: ⭐⭐⭐⭐ [Principled alignment of hierarchical neural representations with time-adaptive diffusion decoding]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluation spanning all 12 subjects, rigorous geometric and semantic metrics, detailed component ablations, and EEG scalp mapping]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical formulations, well-organized methodology, and rigorous empirical validation]
  • Value: ⭐⭐⭐⭐ [Establishes a strong technical benchmark for non-invasive, lightweight 3D neural decoding in BCI systems]