MedGSSR: Generalizable Medical Image Super-Resolution 3D Reconstruction via Hierarchical Feed-forward Gaussian Splatting¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://william2ai.github.io/medgssr
Area: Medical Imaging
Keywords: 3D Medical Image Super-Resolution, Feed-forward Gaussian Splatting, Arbitrary-Scale Reconstruction, Hierarchical Gaussian Projector, Differentiable Gaussian Voxelizer
TL;DR¶
Addressing the core limitations of per-subject optimization, spectral bias in coordinate-based neural implicit representations, and slice-wise inconsistencies of 2D proxy supervision, MedGSSR introduces a generalized, purely volumetric feed-forward 3D Gaussian Splatting framework that factorizes anatomy into coarse structure and fine texture via sub-voxel decomposition, enabling fast, test-time optimization-free, arbitrary-scale 3D super-resolution with exceptional cross-dataset generalizability.
Background & Motivation¶
High-resolution (HR) isotropic volumetric medical imaging, such as high-field MRI and thin-slice CT, provides an indispensable foundation for accurate diagnostic decision-making, clinical lesion delineation, and quantitative morphological analysis. However, acquiring high-resolution scans in routine practice is severely impeded by scanner hardware availability, prolonged acquisition durations that heighten vulnerability to patient motion artifacts, and radiation dose constraints in X-ray computed tomography. While Medical 3D Super-Resolution (Med3DSR) offers a software-driven computational paradigm to reconstruct isotropic high-resolution volumes from observed low-resolution (LR) inputs without modifying imaging equipment, existing approaches encounter persistent representational and computational bottlenecks when scaling to clinical deployment.
Coordinate-based Implicit Neural Representations (INRs) enable continuous, arbitrary-scale spatial querying, yet their shared global MLP parameterization inherently suffers from spectral bias, heavily prioritizing low-frequency components and yielding over-smoothed tissue interfaces where fine vascular pathways and trabecular microstructures are washed out. Concurrently, due to the scarcity of paired isotropic 3D volumetric ground truth, explicit neural reconstruction frameworks have largely retreated to either isolated per-subject self-supervised optimization (e.g., CuNeRF) or 2D proxy supervision via planar projection or slice-wise diffusion models. Per-subject optimization requires several minutes of gradient descent per scan and fundamentally fails to accumulate transferable anatomical priors across patients; meanwhile, 2D slice-wise proxies lack full 3D spatial awareness, inevitably inducing inter-slice inconsistencies, boundary stepping artifacts, and geometric projection ambiguities.
The core insight of this paper is that medical volumes physically represent continuous scalar fields (such as Hounsfield units in CT or proton densities in MRI), which can be explicitly modeled by the superposition of continuous 3D Gaussian primitives with localized compact support, bypassing global spectral bias while preserving physical density continuity. Core idea: reformulate Med3DSR as a generalized, purely volumetric feed-forward regression from sparse low-resolution voxel grids to an explicit continuous 3D Gaussian Splatting field, hierarchically decoupling macroscopic anatomical geometry from microscopic textural details via sub-voxel decomposition and differentiable 3D voxelization to achieve high-fidelity arbitrary-scale super-resolution without test-time optimization.
Method¶
Overall Architecture¶
The MedGSSR pipeline consists of three tightly coupled components: the Pyramid Anatomical Encoder (PAE), the Hierarchical Gaussian Projector (HGP), and the Differentiable Gaussian Voxelizer (DGV). Given a low-resolution volumetric patch \(V_{LR} \in \mathbb{R}^{H \times W \times D}\), the PAE utilizes a 3D U-Net backbone stripped of Batch Normalization layers to extract multi-scale volumetric features. It generates a coarse feature grid \(F_{coarse}\) at the base resolution \(1\times\) and a fine feature grid \(F_{fine}\) at \(2\times\) upsampled resolution via an input-residual concatenation and 3D PixelShuffle. Subsequently, the HGP processes these dual-resolution features through parallel lightweight MLPs with sub-voxel Gaussian decomposition, regressing 11-dimensional attributes for both structural and textural Gaussian primitives within each anchor voxel. Finally, the DGV aggregates the resulting continuous 3D Gaussian field into the target high-resolution spatial coordinate grid via differentiable analytical superposition and Mahalanobis truncation, rendering the super-resolved volume \(V_{HR}\) at an arbitrary upsampling factor \(\kappa\).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["LR Volume V_LR"] --> B["Pyramid Anatomical Encoder<br/>3D U-Net dual-stream feature extraction"]
B --> C["Hierarchical Decoupling & Sub-voxel Decomposition<br/>Structural Gaussians (Coarse) + Textural Gaussians (Fine)"]
C --> D["Explicit 3D Gaussian Primitive Set G<br/>K = N_coarse + N_fine"]
D --> E["Differentiable Gaussian Voxelizer<br/>3σ Mahalanobis continuous scalar superposition"]
E --> F["Arbitrary-Scale HR Volume V_pred"]
Key Designs¶
1. Pyramid Anatomical Encoder: Dual-Stream Feature Decoupling and Absolute Intensity Preservation
In medical imaging, macroscopic organ morphologies and microscopic tissue textures operate across drastically disparate spatial scales; conventional single-scale feature encoders struggle to preserve contextual topology while maintaining high-frequency acuity. To address this, the PAE adopts a 3D U-Net backbone while deliberately omitting all Batch Normalization (BN) layers across the entire architecture, thereby preventing sample-wise normalization statistics from corrupting absolute physical intensity values (such as Hounsfield Units in CT or calibrated proton density in MRI). From the final decoder representation \(H_{dec}\), the coarse branch applies a 3D convolutional layer to produce \(F_{coarse} \in \mathbb{R}^{H \times W \times D \times C}\) at the original \(1\times\) resolution. Each voxel in \(F_{coarse}\) spans an 8× larger physical volume than its fine-branch counterpart, providing a broad receptive field to capture macro-level organ geometry. Simultaneously, the fine branch concatenates the raw input volume \(V_{LR}\) with \(H_{dec}\) as an explicit residual signal \([H_{dec}, V_{LR}]\), which is channeled through a 3D convolution and a 3D PixelShuffle operator with a factor of 2, generating a high-resolution feature grid \(F_{fine} \in \mathbb{R}^{2H \times 2W \times 2D \times C}\) dedicated to resolving fine-grained tissue textures.
2. Hierarchical Decoupling & Sub-voxel Decomposition: Continuous Basis Overcoming Discrete Grid Bounds
When reconstructing continuous 3D volumes from discrete voxel grids, allocating merely a single primitive per voxel severely limits the local expressive capacity, introducing discretization artifacts around intricate tissue interfaces. The HGP resolves this by deploying parallel three-layer MLPs equipped with a Sub-voxel Gaussian Decomposition mechanism: for each anchor voxel \(v \in \{v_c, v_f\}\), the projector predicts \(m\) distinct Gaussian primitives (default \(m=4\)) with learnable local offsets and anisotropic scales, yielding a unified primitive count of \(K = m(HWD) + m(8HWD) = 9mHWD\). For the \(j\)-th sub-Gaussian primitive, the network regresses its 11-dimensional parameter set relative to the voxel anchor center \(v\):
- Center coordinate \(\boldsymbol{\mu}_j = \boldsymbol{v} + \delta \cdot \tanh(\Delta \boldsymbol{\mu}_j)\), where \(\delta\) restricts positional deviation within a localized neighborhood to enforce spatial coherence and avert primitive collapse;
- Anisotropic scaling \(\boldsymbol{s}_j = \text{softplus}(\Delta \boldsymbol{s}_j + s_{base})\);
- Rotation unit quaternion \(\boldsymbol{q}_j = \boldsymbol{\rho}_j / \|\boldsymbol{\rho}_j\|\) guaranteeing a positive definite covariance matrix;
- Intensity amplitude \(\alpha_j = \text{sigmoid}(\rho_{\alpha, j} + \alpha_{base})\).
To enforce structural-textural role differentiation, Structural Gaussians are initialized with a smaller scale base \(s_{base} = 4.0\) and higher amplitude bias \(\alpha_{base} = 2.0\) to model dense macroscopic volume bodies, whereas Textural Gaussians are initialized with a larger scale base \(s_{base} = 6.0\) and lower amplitude bias \(\alpha_{base} = 4.0\) to capture sparse high-frequency boundary variations.
3. Differentiable Gaussian Voxelizer: Direct 3D Scalar Field Superposition and Arbitrary-Scale Querying
Conventional 3DGS frameworks are tailored for view synthesis, projecting primitives onto 2D camera sensor planes via ray-based alpha blending. In volumetric medical imaging, this projection mechanism creates geometric ambiguity and view-dependent artifacts. The DGV reformulates volumetric rendering directly in 3D Euclidean space by modeling the continuous intensity field \(I(\boldsymbol{x}): \mathbb{R}^3 \to \mathbb{R}\) at an arbitrary coordinate \(\boldsymbol{x}\) as the direct analytical superposition of all \(K\) Gaussian primitives:
$\(I(\boldsymbol{x}) = \sum_{i=1}^K \alpha_i \exp\left( -\frac{1}{2} (\boldsymbol{x} - \boldsymbol{\mu}_i)^\top \boldsymbol{\Sigma}_i^{-1} (\boldsymbol{x} - \boldsymbol{\mu}_i) \right)\)$
where the covariance matrix is parameterized as \(\boldsymbol{\Sigma}_i = \boldsymbol{R}(\boldsymbol{q}_i) \operatorname{diag}(\boldsymbol{s}_i^2) \boldsymbol{R}(\boldsymbol{q}_i)^\top\). For computational scalability, the DGV evaluates primitives within a \(3\sigma\) truncation radius in Mahalanobis space, strictly retaining primitives satisfying \(d_M(\boldsymbol{x}, \boldsymbol{\mu}_i) = \sqrt{(\boldsymbol{x} - \boldsymbol{\mu}_i)^\top \boldsymbol{\Sigma}_i^{-1} (\boldsymbol{x} - \boldsymbol{\mu}_i)} < 3\). During inference at any arbitrary upsampling scale \(\kappa = (\kappa_h, \kappa_w, \kappa_d)\), a dense coordinate grid of dimensions \((\lceil \kappa_h H \rceil \times \lceil \kappa_w W \rceil \times \lceil \kappa_d D \rceil)\) is instantiated and queried against \(I(\boldsymbol{x})\), completing full 3D continuous super-resolution in a single forward pass without iterative optimization.
Loss & Training¶
Because the DGV process is entirely analytical and differentiable, MedGSSR is trained in a fully end-to-end manner directly in the 3D domain. Reconstructed voxel intensities \(V_{pred}\) at the target ground-truth grid resolution are supervised using a voxel-wise \(L_1\) reconstruction loss: $\(\mathcal{L}_{rec} = \| V_{GT} - V_{pred} \|_1\)$ The model is trained for 1,000,000 iterations using the Adam optimizer across four NVIDIA RTX 4090 GPUs. The learning rate starts at \(1 \times 10^{-4}\) and decays logarithmically to \(1 \times 10^{-5}\). The framework operates strictly on 3D volumetric patches, eliminating any dependence on pretrained 2D diffusion models or auxiliary 2D single-image super-resolution networks.
Key Experimental Results¶
Main Results¶
Quantitative evaluations were conducted across two core medical modalities: MRI (MSD T1-weighted brain volumes) and CT (MELA chest CT volumes), benchmarked against traditional interpolation, self-supervised per-subject optimization (CuNeRF), 2D-diffusion-guided projection (NAB-GS), and neural implicit coordinate networks (ArSSR) across \(2\times\), \(3\times\), and \(4\times\) upsampling factors.
| Dataset (Modality) | Scale | Paradigm | Method | PSNR (dB) ↑ | SSIM ↑ | LPIPS ↓ | Inference Time |
|---|---|---|---|---|---|---|---|
| MSD (MRI) | 2× | Traditional | Trilinear | 33.31 | 0.9670 | 0.1134 | - |
| Traditional | Cubic | 33.99 | 0.9725 | 0.0981 | - | ||
| Self-supervised | CuNeRF | 32.60 | 0.9715 | 0.0513 | ~311 s | ||
| Neural Implicit | ArSSR | 32.87 | 0.9742 | 0.0581 | ~80 s | ||
| FF-3DGS | MedGSSR (Ours) | 35.91 | 0.9821 | 0.0337 | ~8 s | ||
| 3× | Traditional | Cubic | 31.13 | 0.9422 | 0.2083 | - | |
| Neural Implicit | ArSSR | 29.56 | 0.9442 | 0.0730 | ~80 s | ||
| FF-3DGS | MedGSSR (Ours) | 33.97 | 0.9700 | 0.0706 | ~8 s | ||
| 4× | Traditional | Trilinear | 29.08 | 0.9100 | 0.2558 | - | |
| Traditional | Cubic | 29.27 | 0.9118 | 0.2711 | - | ||
| Self-supervised | CuNeRF | 28.26 | 0.9309 | 0.1271 | ~311 s | ||
| Neural Implicit | ArSSR | 28.51 | 0.9315 | 0.0853 | ~80 s | ||
| FF-3DGS | MedGSSR (Ours) | 32.10 | 0.9541 | 0.1246 | ~8 s | ||
| MELA (CT) | 2× | Neural Implicit | ArSSR | 38.53 | 0.9644 | 0.0914 | ~80 s |
| FF-3DGS | MedGSSR (Ours) | 42.08 | 0.9738 | 0.0705 | ~8 s | ||
| 4× | Traditional | Cubic | 34.09 | 0.8888 | 0.3263 | - | |
| Self-supervised | CuNeRF | 33.94 | 0.8957 | 0.2456 | ~311 s | ||
| Self-supervised (2D Prior) | NAB-GS | 34.13 | 0.9518 | - | - | ||
| Neural Implicit | ArSSR | 34.63 | 0.9244 | 0.2261 | ~80 s | ||
| FF-3DGS | MedGSSR (Ours) | 37.31 | 0.9362 | 0.1472 | ~8 s |
In cross-dataset generalization experiments, models trained on MSD were evaluated zero-shot on the unseen Human Connectome Project (HCP) MRI dataset: at 4× upsampling, MedGSSR achieved 35.84 dB PSNR, outperforming ArSSR (31.87 dB) by a remarkable +3.97 dB margin. On the Ultra-High-Resolution CT (UHRCT) benchmark evaluated at an extreme 8× upsampling scale under domain shift, MedGSSR registered 35.24 dB PSNR and 0.9233 SSIM, surpassing ArSSR (29.38 dB) by +5.86 dB and CuNeRF (30.50 dB) by +4.74 dB. Furthermore, downstream clinical evaluation via automated brain tissue segmentation using nnUNet on 4× super-resolved HCP volumes confirmed that MedGSSR reached an average Dice score of 0.8178, well exceeding ArSSR (0.7908) and CuNeRF (0.3618, which experienced catastrophic boundary failure).
Ablation Study¶
Ablation studies on the HCP dataset for 4× 3D SR comprehensively dissect the impact of structure-texture decoupling, sub-voxel count \(m\), truncation radius, and auxiliary supervision.
| Category | Configuration | PSNR (dB) ↑ | SSIM ↑ | LPIPS ↓ | Note |
|---|---|---|---|---|---|
| Branch Decoupling | Coarse Only (Base) | 35.24 | 0.9802 | 0.0789 | Captures macroscopic geometry but produces blurred cortical folds |
| Coarse + Fine (Full Model) | 36.60 | 0.9870 | 0.0363 | Fine branch adds high-frequency residuals (+1.36 dB PSNR, >50% lower LPIPS) | |
| Sub-voxel Count \(m\) | \(m = 1\) | 35.98 | 0.9760 | 0.0423 | Limited capacity; single Gaussian per voxel struggles at tissue boundaries |
| \(m = 2\) | 36.06 | 0.9771 | 0.0388 | Increased expressiveness with consistent quantitative gains | |
| \(m = 4\) (Default) | 36.60 | 0.9870 | 0.0363 | Optimal trade-off between anatomical accuracy and memory efficiency | |
| \(m = 8\) | 36.52 | 0.9862 | 0.0365 | Performance plateaus with slight noise sensitivity and double compute | |
| Gaussian Truncation & Loss | Full Model (3σ Truncation) | 35.84 | 0.9533 | 0.1253 | Cross-domain evaluation baseline |
| Aggressive 1σ Truncation | 31.05 | 0.9142 | 0.1608 | Severely compromised Gaussian overlap leads to field discontinuity (-4.79 dB) | |
| Auxiliary FFT Loss | 35.72 | 0.9472 | 0.1078 | Modest perceptual edge sharpness gain at the cost of voxel-wise fidelity | |
| 75% Training Data Scale | 35.49 | 0.9512 | 0.1283 | Retains competitive accuracy, confirming robust prior learning | |
| 50% Training Data Scale | 34.26 | 0.9405 | 0.1452 | Notable drop due to insufficient anatomical distribution coverage |
Key Findings¶
- Hierarchical structure-texture decoupling drives fine-detail recovery: Relying solely on the coarse branch provides adequate macroscopic organ layout (35.24 dB) but fails to resolve cortical sulci and fine vessels. Incorporating the fine branch slashes LPIPS from 0.0789 to 0.0363, proving that dual-scale feature splitting and differentiated parameter initialization effectively disentangle frequency bands.
- Sub-voxel primitive density exhibits clear saturation dynamics: Expanding sub-voxel count from \(m=1\) to \(m=4\) delivers a steady +0.62 dB PSNR boost. However, further scaling to \(m=8\) causes slight metric saturation (36.52 dB) due to parameter redundancy, confirming \(m=4\) as the optimal Pareto front for clinical volumetric representation.
- Spatial overlap continuity is non-negotiable for scalar field fidelity: Restricting the Mahalanobis truncation radius to \(1\sigma\) causes catastrophic performance degradation (-4.79 dB PSNR), validating that continuous medical density fields require sufficient multi-Gaussian spatial support to prevent voxelization stepping artifacts.
Highlights & Insights¶
- Shifting from per-scene optimization to direct feed-forward 3DGS: MedGSSR breaks away from the conventional reliance on time-consuming per-subject optimization or fragile 2D proxy projections in neural field medical reconstruction, achieving an order-of-magnitude acceleration (inference reduced from 311s to 8s) while demonstrating unprecedented cross-subject generalization.
- Differentiated physical initialization for frequency factorization: Rather than letting the network discover structural boundaries unconstrained, the architecture injects inductive biases by assigning distinct spatial scale (\(s_{base}\)) and amplitude (\(\alpha_{base}\)) offsets to structural versus textural branches, establishing a principled multi-band Gaussian decomposition.
- Tangible benefits for downstream clinical segmentation: By preventing the tissue bleeding and edge blur inherent to coordinate-based INRs, MedGSSR delivers sharp anatomical boundaries that improve automated brain tissue segmentation (average Dice reaching 0.8178 vs. 0.7908 for ArSSR and 0.3618 for CuNeRF), bridging the gap between perceptual metrics and clinical utility.
Limitations & Future Work¶
- Limitations acknowledged by authors: While MedGSSR demonstrates remarkable generalization across standard datasets, handling extreme pathology cases or major anatomical resections at large scale factors (e.g., >8×) may still require exposure to broader heterogeneous training distributions. Furthermore, current models are trained on individual modalities rather than a single unified multi-modal network spanning MRI and CT.
- Potential future directions: Although \(3\sigma\) Mahalanobis truncation provides efficient rendering for standard patches, full-organ isotropic super-resolution at massive grid dimensions (e.g., \(1024^3\)) incurs non-trivial GPU memory footprints. Future work could integrate octree-based spatial chunking or hierarchical LOD primitives. Additionally, incorporating anatomical topology priors (such as vascular connectivity loss) could further enhance structural consistency in microvascular regions.
Related Work & Insights¶
- vs CuNeRF [ICCV 2023]: CuNeRF requires slow, isolated test-time self-supervision on every single patient scan and cannot accumulate population-level anatomical priors; MedGSSR learns generalized feed-forward mapping that executes in 8 seconds and delivers robust zero-shot cross-domain generalization.
- vs ArSSR [CVPR 2024] / Coordinate INRs: Coordinate-based implicit MLPs suffer from spectral bias that blurs fine vascular boundaries and cortical folds; MedGSSR exploits explicit 3D Gaussian primitives with compact local support, naturally preserving sharp, high-frequency physical transitions.
- vs NAB-GS [arXiv 2025]: NAB-GS relies on projecting 3D volumes into 2D X-ray planes to exploit 2D diffusion priors, introducing inter-slice inconsistencies and projection ambiguity; MedGSSR operates strictly in the continuous 3D volumetric domain, maintaining strict spatial and physical consistency.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First end-to-end feed-forward 3D Gaussian Splatting architecture formulated for generalized, arbitrary-scale medical 3D super-resolution, with highly principled sub-voxel decomposition and 3D differentiable voxelization.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive cross-modality (MRI & CT) validation, intra-domain comparisons across multiple upsampling scales, challenging cross-dataset evaluations up to 8× SR, downstream clinical segmentation testing, and detailed ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear technical prose, rigorous mathematical formulations, informative architectural diagrams, and well-structured empirical analyses.
- Value: ⭐⭐⭐⭐⭐ Resolves the long-standing conflict between clinical real-time inference speed and high-frequency structural fidelity, offering a foundational framework for isotropic medical volumetric enhancement.