Skip to content

Gaussian Volumetric Representation for Efficient Shear–Warp Visualization

Conference: ECCV 2026
Paper: ECCV Official
Code: Project Page / Code
Area: Medical Imaging
Keywords: Gaussian volume representation, shear-warp volume rendering, Monte Carlo estimation, curriculum learning, medical volume visualization

TL;DR

To overcome the prohibitive cost of ray marching and the surface limitation of standard Gaussian splatting on dense medical volumes (MRI and Visible Korean Cryosection), this paper presents a continuous Gaussian volumetric representation optimized via unbiased Monte Carlo importance sampling and slice-based curriculum learning, coupled with an adaptive single-stack GPU shear–warp renderer that achieves up to 43.86 FPS rendering and an 11.31:1 compression ratio.

Background & Motivation

Three-dimensional medical volumetric data such as high-field MRI and Cryosection serial cross-sections embody intricate internal anatomical structures and tissue textures essential for clinical diagnosis, anatomical education, and surgical planning. However, direct volume rendering techniques predominantly rely on ray marching, which traces rays through dense voxel grids with repeated trilinear interpolations. As volumetric resolution scales (such as typical \(512^3\) grids) alongside high-definition display demands, massive memory bandwidth consumption and dense floating-point evaluations render interactive exploration computationally prohibitive. Conventional voxel compression schemes (such as tensor decomposition, discrete cosine transforms, or octree encodings) merely approximate discrete signals without tightly coupling into the rendering pipeline; conversely, implicit neural representations (NeRF, 3D CNNs, or Transformers) suffer from expensive multi-layer perceptron (MLP) forward evaluations or quadratic attention overheads, making real-time frame rates unattainable.

Meanwhile, 3D Gaussian Splatting has recently emerged as an efficient surface-oriented paradigm, projecting anisotropic ellipsoids onto screen space via visibility-aware rasterization. Nevertheless, 3DGS natively assumes surface-like scene geometries with viewpoint-dependent radiance, failing to model continuous internal volumetric densities where meaningful anatomical boundaries permeate the interior. While formulating Gaussians as continuous volumetric density fields under radiative transfer equations can achieve physical consistency, evaluating volumetric ray integrals reintroduces heavy sampling burdens, sacrificing the real-time speed advantage.

The angle of attack in this paper is to treat Gaussian kernels as continuous basis functions approximating the dense volumetric signal, tailored specifically to the slice-based shear–warp rendering algorithm whose computational complexity is significantly lower than ray marching. Directly optimizing Gaussian mixtures against complete dense voxel grids is intractable, whereas naive random sampling severely impairs sharp anatomical boundaries. Core idea: parameterize dense volumetric signals into a compact Gaussian mixture field, optimize Gaussian parameters from highly sparse supervision via importance-weighted unbiased Monte Carlo estimation combined with smooth planar slice curriculum learning, and visualize internal structures via an adaptive single-stack GPU shear–warp renderer for real-time frame rates and high compression.

Method

Overall Architecture

The input consists of a dense medical voxel grid (such as Cryosection RGB or MRI scalar fields), and the output is a compact set of continuous Gaussian volumetric kernels together with interactive shear–warp rendering views. The overall pipeline operates in three cooperative stages: first, a sparse voxel subset occupying less than 25% of the original memory footprint (biased toward high-gradient anatomical boundaries) is selected to initialize the Gaussian kernels; second, an optimization phase schedules Monte Carlo importance sampling alongside randomly oriented 2D planar slice pixel supervision via curriculum learning, establishing global volume coverage in early iterations and smoothing local slice continuity later; third, during inference, the trained Gaussian field dynamically generates view-aligned 2D slice textures and streams them into a custom GPU shear–warp compositing pipeline for rapid affine shearing and front-to-back alpha compositing.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Dense Medical Volume<br/>(MRI / Cryosection Voxel Grid)"] --> B["Sparse Gaussian Field Initialization<br/>(Gradient-guided <25% Voxel Memory)"]
    B --> C["Monte Carlo Volumetric Estimation<br/>(Importance-weighted Unbiased Loss)"]
    C --> D["Curriculum Structured Sampling<br/>(Smooth Sigmoid Shift: Voxels → Planar Slices)"]
    D --> E["Continuous Gaussian Volumetric Field<br/>(Mean μ, Covariance Σ, Color C, Opacity α)"]
    E --> F["Adaptive Single-Stack Shear–Warp GPU Renderer<br/>(View-aligned Dynamic Slicing + Angle-cached Shearing)"]
    F --> G["Real-time Medical Volume View<br/>(>40 FPS Interactive Rendering)"]

Key Designs

1. Monte Carlo Volumetric Estimation: Unbiased Sparse Approximation to Dense Reconstruction

High-resolution grids (e.g., \(512^3\)) encompass hundreds of millions of voxels. Directly evaluating and back-propagating the full mean squared error loss \(L_{\text{dense}} = \frac{1}{n} \sum_{\mathbf{x} \in \Omega} \|V_\Theta(\mathbf{x}) - V(\mathbf{x})\|_2^2\) over the entire voxel set \(\Omega\) is computationally and memory-wise intractable. Naive random subsampling of \(m \ll n\) voxels either squanders capacity on empty background under uniform sampling or introduces intrinsic distribution bias when sampling more frequently around high-gradient boundaries, which diverts stochastic gradients from the true volume objective. To resolve this dilemma, the authors introduce an importance-weighted Monte Carlo estimator:

\[ \hat{L}_{\text{MC}} = \frac{1}{m} \sum_{i=1}^{m} \frac{1}{n \, p(\mathbf{x}_i)} \|V_\Theta(\mathbf{x}_i) - V(\mathbf{x}_i)\|_2^2, \quad \mathbf{x}_i \sim p(\mathbf{x}) \]

where the sampling probability balances global uniform coverage with reconstruction-error importance: \(p(\mathbf{x}) = \lambda \frac{1}{n} + (1 - \lambda) p_{\text{importance}}(\mathbf{x})\), with \(p_{\text{importance}}(\mathbf{x}) \propto \ell(\mathbf{x}; \Theta) + \epsilon\). This mathematical formulation guarantees an unbiased objective expectation \(\mathbb{E}_{p}[\hat{L}_{\text{MC}}] = L_{\text{dense}}(\Theta)\) as well as unbiased stochastic gradients \(\mathbb{E}[\nabla_\Theta \hat{L}_{\text{MC}}] = \nabla_\Theta L_{\text{dense}}\). Periodically refreshing high-gradient boundary samples ensures rapid convergence toward the true dense volume optimum using only a minuscule fraction of voxels.

2. Curriculum Structured Sampling: Smooth Transition from Global Coverage to Planar Continuity

Isolated 3D voxel sampling provides unbiased statistical guidance across the volume domain, yet lacks local spatial coherence constraints among adjacent spatial coordinates. This often leads anisotropic Gaussian ellipsoids to deform irregularly along unconstrained axes, resulting in high-frequency pepper-like noise on arbitrary cross sections. Furthermore, shear–warp compositing relies fundamentally on coherent 2D slice textures. To bridge this gap, structured planar supervision on randomly oriented 2D slice planes is incorporated via a smooth Sigmoid-based curriculum learning schedule:

\[ \lambda_s(t) = \lambda_{s,\max} \cdot \sigma\left(\frac{t - \gamma_s}{\beta_s}\right), \quad \lambda_v(t) = 1 - \lambda_s(t) \]

The overall training loss is defined as \(\mathcal{L}(t) = \lambda_v(t) \hat{L}_{\text{MC}} + \lambda_s(t) L_{\text{slice}}\). In initial iterations (\(t \ll \gamma_s\)), \(\lambda_v(t) \approx 1\), allowing Monte Carlo voxel supervision to anchor Gaussians onto macro organ envelopes; as iterations advance, the slice supervision weight smoothly climbs. In each step, an arbitrary 3D plane is sampled and sparse 2D pixel coordinates are mapped back to 3D voxel space. Coplanar geometric constraints enforce adjacent Gaussians to align smoothly along anatomical boundaries and tissue interfaces, eliminating slice-wise discontinuities.

3. Adaptive Single-Stack Shear–Warp GPU Renderer: Eliminating Multi-Stack Storage via View-Aligned Generation

The classical shear–warp volume rendering algorithm decomposes viewing transformations into a shear in object space and a 2D warp on the intermediate image, replacing 3D trilinear interpolation with efficient 2D bilinear sampling. However, traditional implementations require precomputing and storing three axis-aligned slice stacks parallel to the \(X, Y, Z\) Cartesian planes to avoid grazing-angle degeneracies, tripling volumetric memory consumption.

Because the learned continuous Gaussian field represents space without rigid coordinate-axis dependencies, the authors design an adaptive single-stack GPU renderer. Instead of maintaining three static stacks, parallel 2D slice textures are dynamically extracted from the continuous Gaussian field perpendicular to the instantaneous camera viewing vector. The renderer incorporates an angular rotation threshold: as long as camera motion remains within this threshold, intermediate slice textures are reused and only the shear matrix and front-to-back alpha compositing are updated. When camera rotation exceeds the threshold, the slice stack is refreshed and re-aligned. This eliminates redundant multi-stack memory overhead while delivering rendering speeds substantially faster than ray marching.

Loss & Training

The parameters are optimized using the Adam optimizer with an initial learning rate of \(5 \times 10^{-4}\) over 120 epochs. The initial Gaussian count is bounded within 25% of the raw volume memory footprint, comprising 60% high-gradient boundary voxels and 40% uniform random volume samples. The Monte Carlo distribution \(p(\mathbf{x})\) is updated periodically, while slice loss weight ramps up following the Sigmoid schedule, ensuring stable optimization without ever materializing the dense voxel grid in GPU memory.

Key Experimental Results

Main Results

The framework is comprehensively evaluated across multi-modal medical datasets, including grayscale T2-weighted MRI from BraTS, neonatal and embryonic mouse high-resolution histology MRI, and full-color RGB Cryosection datasets from the Visible Korean Project (including isolated brain, liver, heart, and lung organ volumes). Baselines encompass coordinate MLPs (Voxel MLP, Slice MLP), 3D convolutional networks (3D UNet), volumetric Transformers, and advanced neural radiance / Gaussian models (CVT-xRF, FreeSplatter, GNT-MOVE).

The table below summarizes volumetric reconstruction fidelity, rendering speed, and volume compression ratios across diverse datasets:

Dataset Modality PSNR (dB) SSIM MS-SSIM FPS Compression Ratio
Brain (Cryosection) Color Cryo RGB 32.10 0.943 0.991 45.50 9.34:1
Liver (Cryosection) Color Cryo RGB 43.16 0.981 0.996 50.18 9.45:1
Heart (Cryosection) Color Cryo RGB 32.19 0.958 0.990 48.57 10.09:1
Lungs (Cryosection) Color Cryo RGB 34.93 0.952 0.989 39.77 4.73:1
BraTS Grayscale MRI 40.88 0.988 0.998 44.82 11.31:1
Mouse Neonatal Grayscale MRI 33.48 0.956 0.987 46.56 9.94:1

The performance comparison against high-capacity deep learning volumetric representations on Cryosection data is detailed below:

Method Recon. PSNR↑ Recon. SSIM↑ Render PSNR↑ Render SSIM↑ Training Time (h) Memory (GB) FLOPs Render FPS↑
Voxel MLP 23.63 0.772 46.61 0.968 6.7 15 \(9.50 \times 10^{12}\) 18.99
Slice MLP 30.02 0.920 36.32 0.973 35.0 14 \(1.88 \times 10^{13}\) 14.27
Transformer 9.60 0.631 29.73 0.873 33.3 30 \(1.91 \times 10^{15}\) 0.11
UNet 26.70 0.798 42.10 0.982 16.7 36 \(2.05 \times 10^{12}\) 15.00
CVT-xRF 18.98 0.591 15.09 0.536 19.3 22 \(2.84 \times 10^{15}\) 30.17
FreeSplatter 14.74 0.483 16.24 0.569 0.13 11 \(1.27 \times 10^{12}\) 36.01
GNT-MOVE 11.28 0.046 15.11 0.084 39.2 12 \(1.05 \times 10^{15}\) N/A†
Ours 34.81 0.956 54.87 0.995 6.0 28 \(9.17 \times 10^9\) 43.86

Ablation Study

To rigorously evaluate the individual impact of Monte Carlo estimation and the curriculum sampling strategy, the authors benchmark five ablation configurations: - A1 (Random voxel supervision): Random sampling and high-gradient voxels without MCE importance weight correction. - A2 (Curriculum: random voxels \(\rightarrow\) random slices): Curriculum transition without MCE importance weighting. - A3 (Curriculum: random slices \(\rightarrow\) random voxels): Inverted curriculum sequence starting with slices before voxels without MCE. - A4 (Only slice supervision): Slice supervision exclusively without global voxel guidance. - A5 (MCE voxel-based supervision): Unbiased MCE voxel estimation without slice supervision. - Ours (Full model): Unbiased MCE voxel estimation combined with smooth slice curriculum learning.

Quantitative ablation metrics across cross-sectional reconstruction and rendered views are presented below:

Config Description Recon. PSNR↑ Recon. SSIM↑ Recon. MS-SSIM↑ Render PSNR↑ Render SSIM↑ Render MS-SSIM↑
A1 Random voxel supervision (No MCE correction) 33.46 0.956 0.990 50.70 0.962 0.964
A2 Curriculum: random voxels \(\rightarrow\) slices (No MCE) 12.16 0.060 0.214 27.10 0.826 0.789
A3 Curriculum: random slices \(\rightarrow\) voxels (No MCE) 22.64 0.727 0.850 40.59 0.932 0.937
A4 Only slice supervision (No voxel anchor) 14.35 0.646 0.512 33.18 0.878 0.887
A5 MCE voxel supervision only (No slice coherence) 32.40 0.943 0.988 52.34 0.968 0.991
Ours Full model (MCE voxels + slice curriculum) 34.81 0.956 0.993 54.87 0.995 0.993

Key Findings

  • Synergy of MCE and Curriculum Scheduling: The contrast between A2/A3 and Ours reveals that introducing slice supervision naively without MCE importance weights severely destabilizes training (A2 reconstruction PSNR drops catastrophically to 12.16 dB). Unbiased importance weighting anchors optimization, enabling slice supervision to inject planar structural coherence cleanly (+2.53 dB rendering PSNR over A5).
  • Critical Progression Direction: Inverting the curriculum sequence (A3, slice-first) leads to severe local minima (PSNR 22.64 dB vs 34.81 dB), establishing that Gaussians require initial 3D volumetric anchoring before planar refinement can be successfully assimilated.
  • Drastic Computation Reduction: Ours requires only \(9.17 \times 10^9\) FLOPs, lowering evaluation cost by 5 to 6 orders of magnitude compared to Transformer and NeRF baselines, boosting rendering throughput from 15.05 FPS (ray marching) to 43.86 FPS (a 2.91× speedup).

Highlights & Insights

  • Adapting Continuous Gaussians to Classical Shear–Warp: Elegantly leverages the continuous orientation freedom of Gaussian kernels to overcome the multi-stack storage bottleneck of classic shear–warp algorithms, generating view-aligned slices on the fly.
  • Monte Carlo Importance Sampling for Dense Volume Fitting: Demonstrates that continuous volume fitting does not require evaluating dense voxel grids; an unbiased importance-weighted estimator converges accurately to the dense optimum while training on a sparse voxel budget.
  • Cross-Modality Anatomical Fidelity: Validated on both scalar grayscale MRI scans and RGB anatomical cryosections, showing remarkable color preservation and high compression ratios up to 11.31:1.

Limitations & Future Work

  • Smooth Kernel Blurring: Because Gaussian primitives rely on smooth spatial radial basis kernels, subtle over-smoothing or blurring can manifest at ultra-sharp density transitions, such as thin bone cortices or fine bronchial walls.
  • Sparse Sampling Dot Artifacts: Under extreme sparsity budgets (<25% memory), isolated bright dot-like artifacts can occasionally appear in regions with sparse spatial coverage.
  • Future Improvements: Exploring compact-support Gaussian basis kernels to preserve hard edge discontinuities, along with anatomical geometric regularizations to eliminate sparse dot artifacts.
  • vs 3D Gaussian Splatting (Kerbl et al.): 3DGS optimizes surface-oriented splats via screen projection rasterization; this work models internal 3D continuous volume fields and integrates with slice-based volume compositing.
  • vs Classical Shear–Warp (Lacroute & Levoy): Standard shear–warp precomputes three orthogonal axis-aligned voxel stacks; this approach synthesizes view-directed slice textures dynamically from the continuous field, saving substantial memory.
  • vs NeRF / Coordinate MLPs (Mildenhall et al. / CVT-xRF): Coordinate networks require dense ray marching with repeated MLP evaluations; this work uses analytical Gaussian mixtures and GPU shear rasterization to slash rendering FLOPs by orders of magnitude.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Combines continuous Gaussian volume fields with shear–warp volume rendering, powered by an unbiased MCE and curriculum sampling strategy.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous validation across BraTS, mouse histological MRI, and multiple Korean Cryosection organs against comprehensive neural and splatting baselines.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulations of the unbiased estimator and well-reasoned curriculum design.
  • Value: ⭐⭐⭐⭐☆ Provides a practical, high-throughput solution for real-time medical volume exploration on standard workstation GPUs.