Skip to content

RaPTGS: Render-Agnostic Post-Training Compression of 3D Gaussian Splatting

Conference: ECCV 2026
Paper: ECCV 2026
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Post-Training Compression, Render-Agnostic Pruning, Voxel Mass Compensation, Vector Quantization

TL;DR

Under a strict model-only constraint where training images, camera poses, and rendering pipelines are unavailable, RaPTGS introduces multi-criteria render-agnostic pruning (opacity, shape anisotropy, scale-aware appearance energy), training-free voxel mass compensation, degree-wise SH vector quantization, and spatial-reordering entropy coding to deliver a 30x compression ratio in seconds while preserving high rendering fidelity.

Background & Motivation

3D Gaussian Splatting (3DGS) has revolutionized novel view synthesis by marrying explicit geometric primitives with highly optimized tile-based rasterization, achieving photorealistic visual fidelity at real-time rendering speeds. However, capturing complex real-world scenes typically requires millions of multi-attribute Gaussians—each parameterizing 3D position, 3D scale, quaternion rotation, opacity, and high-degree spherical harmonic (SH) coefficients. This yields prohibitive storage footprints ranging from several hundred megabytes to multiple gigabytes, creating a severe bottleneck for cloud asset transmission, on-device caching, and edge deployment. Consequently, compressing 3DGS has emerged as an active research frontier.

Despite notable progress, existing compression schemes almost uniformly rely on access to the original training setup: multi-view training images, camera trajectories, or active rendering loops. Sensitivity-driven and view-dependent pruning approaches (e.g., LightGaussian, MesonGS, POTR) require extensive forward and backward rasterization passes across training viewpoints to measure splat contributions, while fine-tuning pipelines (e.g., PUP 3D-GS, MesonGS-FT) demand costly retraining epochs. Yet, in realistic asset archival and distribution scenarios, only the trained 3DGS model parameters are accessible; training viewpoints and SfM/COLMAP metadata are frequently proprietary, lost, or impractical to ship. Under this strict model-only constraint, falling back to naive global opacity thresholding severely discards fine structures such as foliage, thin wires, and text textures, leading to catastrophic degradation in visual quality.

The key insight of this paper is that Gaussian significance can be reliably extracted from the intrinsic geometry and appearance parameters of the primitives themselves, completely divorcing compression from rendering loops and retraining. Core idea: construct a render-agnostic importance metric via percentile-ranked maximum pooling over opacity salience, shape anisotropy, and scale-aware appearance energy to guide pruning, coupled with an analytical voxel mass compensation to conservatively restore local spatial coverage, followed by degree-wise SH vector quantization and spatial-reordered entropy coding to attain ultra-fast, 30x high-fidelity compression.

Method

Overall Architecture

RaPTGS operates strictly under the model-only setting, processing a trained 3DGS scene through four consecutive, feed-forward stages without any iterative fine-tuning or view-dependent supervision: 1. Render-Agnostic Pruning: Evaluates each primitive using three view-independent proxies (opacity salience, shape anisotropy strength, and scale-aware appearance energy), maps them to percentile ranks, and takes the element-wise maximum to retain the global Top-\(K\) Gaussians; 2. Post-Pruning Refinement (VMC): Voxelizes the scene into a grid whose element count matches the retained Gaussian budget, computes the ratio of lost Spatial Support Mass within each non-empty voxel, and inflates the retained primitives' scales subject to conservative pre-pruning local maximum scale bounds; 3. Degree-wise SH Quantization: Groups high-order AC spherical harmonic coefficients by band degree \(\ell \in \{1, 2, 3\}\), optimizes global bit allocations across bands under a total bit budget, and clusters coefficients via K-Means to construct compact codebooks and index mapping tables; 4. Attribute Compression: Optimizes Gaussian serialization locality via 3D Hilbert curve sorting followed by parallel windowed 2-opt trajectory refinement, uniformly quantizes remaining attributes, and encodes them with lossless entropy coders.

The complete compression pipeline is illustrated below:

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Pretrained 3DGS Model<br/>(Position, Scale, Rotation, Opacity, SH)"] --> B["Render-Agnostic Pruning<br/>Percentile Max-Pooling of OS, SAS, and SAE"]
    B --> C["Voxel Mass Compensation (VMC)<br/>Local Mass Conservation with Conservative Clamping"]
    C --> D["Degree-wise SH Quantization<br/>Band-Grouped AC SH Vector Quantization"]
    D --> E["Attribute Spatial Reordering & Coding<br/>Hilbert Sorting + Parallel Windowed 2-Opt"]
    E --> F["Output: 30x Compact Compressed Model<br/>(Entropy-Coded Streams and Codebook Indices)"]

Key Designs

1. Render-Agnostic Multi-Criteria Pruning: Decoupling Significance from View Rasterization

Prior pruning paradigms rely heavily on view-aggregated visibility or reconstruction sensitivity, which are unavailable in a model-only environment; relying purely on opacity culls semi-transparent yet geometrically crucial surface points and thin structures. The authors define the linear geometric extent of each Gaussian as \(e_i = s_{i,x} + s_{i,y} + s_{i,z}\), avoiding the degenerate zero-volume issue of highly flattened splats. Pruning decisions are then guided by three complementary intrinsic proxies: - Opacity Salience (OS): \(\mathrm{OS}_i = \alpha_i\), capturing the baseline contribution under volumetric alpha compositing; - Shape Anisotropy Strength (SAS): Captures geometric specificity to safeguard thin splats and structural edges, defined by the scale aspect ratio weighted by opacity with logarithmic dampening: \(\mathrm{SAS}_i = \alpha_i \log\left(\frac{\max(s_{i,x}, s_{i,y}, s_{i,z})}{\min(s_{i,x}, s_{i,y}, s_{i,z}) + \varepsilon}\right)\); - Scale-Aware Appearance Energy (SAE): Quantifies high-frequency view-dependent detail encoded in higher-order spherical harmonic coefficients (\(\ell \ge 1\)). To avoid prematurely culling tiny Gaussians that encode sharp textures, the geometric extent is lower-bounded by \(e_{\min}\) (calibrated such that primitives below \(e_{\min}\) account for 15% of total scene extent), yielding clamped extent \(\tilde{e}_i = \max(e_i, e_{\min})\) and proxy \(\mathrm{SAE}_i = \alpha_i \tilde{e}_i \sum_{\ell \ge 1} \|c_i^{(\ell)}\|_2\).

To reconcile these heterogeneous metrics, each proxy is converted into its percentile rank \(\pi(\cdot) \in [0, 1]\) over the entire model, and combined using a maximum pooling operator: $\(\mathcal{I}_i = \max\Big(\pi(\mathrm{OS}_i), \pi(\mathrm{SAS}_i), \pi(\mathrm{SAE}_i)\Big)\)$ This maximum pooling functions as a logical OR gate: a primitive is protected as long as it excels in visibility, geometric structural anisotropy, or high-frequency appearance complexity, ensuring that only universally redundant primitives are culled.

2. Voxel Mass Compensation (VMC): Analytical Coverage Restoration Without Rendering Feedback

Aggressive pruning inevitably creates localized spatial voids and semi-transparent thinning. Without camera views or gradient backpropagation, fidelity must be restored analytically. The scene is partitioned into a uniform voxel grid of size \(h^*\), optimized via binary search such that the number of active voxels \(V(h^*)\) matches the retained Gaussian budget \(|G'|\). Defining the Spatial Support Mass of an individual Gaussian as \(\mathrm{SSM}_i = \alpha_i e_i\), the aggregate mass of voxel \(v\) is: $\(\mathcal{M}_v = \sum_{g_i \in \mathcal{V}_v} \alpha_i (s_{i,x} + s_{i,y} + s_{i,z})\)$ The ratio of pre-pruning to post-pruning mass, \(\rho_v = \mathcal{M}_v^{\mathrm{pre}} / \mathcal{M}_v^{\mathrm{post}}\), defines a local isotropic scale inflation factor for all surviving Gaussians within voxel \(v\). To prevent catastrophic primitive over-expansion in sparsely surviving voxels, each axis of the scaled Gaussian is strictly bounded by the maximum pre-pruning scale along that axis within the same voxel: \(\min(\rho_v s_{i,k}, \max_{g_j \in \mathcal{V}_v} s_{j,k})\). This conservative clamping restores local optical coverage while preventing blurry, floating artifacts.

3. Degree-wise SH Quantization and Spatial Reordering: High-Density Entropy Serialization

Higher-order spherical harmonic coefficients (AC SH) account for over 75% of raw 3DGS parameters. Because different frequency bands demonstrate distinct variance and energy distributions, coefficients are grouped by degree: \(L_1 \in \mathbb{R}^9\), \(L_2 \in \mathbb{R}^{15}\), and \(L_3 \in \mathbb{R}^{21}\). Under a total per-primitive bit budget \(B = b_1 + b_2 + b_3\), the global bit allocation across bands is resolved by minimizing the dimension-weighted reconstruction error: $\(\min_{b_1 + b_2 + b_3 = B} \sum_{i=1}^3 d_i \mathrm{MSE}_{L_i}(b_i), \quad d_i = 3(2\ell_i + 1)\)$ Codebooks of size \(K_i = 2^{b_i}\) are learned via K-Means clustering, and indices are stored in compact tables.

To maximize entropy coding efficiency, the primitives are serialized to minimize the total \(\ell_1\) path distance \(L(\pi) = \sum \|x_{\pi_i} - x_{\pi_{i+1}}\|_1\) in coordinate space. Because exact TSP optimization across millions of points is intractable, a two-stage heuristic is employed: first, a global spatial coarse sort using 3D Hilbert space-filling curve keys; second, a parallel windowed 2-opt local search within window \(W = |G'|^{1/3}\), where non-overlapping modular sub-regions \(i \equiv r \pmod{W+1}\) are swapped concurrently until relative path length improvement drops below \(\tau = 0.1\%\). The reordered attributes are uniformly quantized (16 bits for positions; 7 bits for scales, rotations, and opacities; 8 bits for SH codebook centroids) and compressed using lossless JPEG XL and ZIP, driving storage down to the theoretical entropy limit.

Loss & Training

RaPTGS is a strictly fine-tuning-free, optimization-free post-processing pipeline requiring zero rendering loss evaluations or gradient descent steps. Primary hyperparameters include pruning ratio \(p \in [0.35, 0.45]\), total SH bit budget \(B \in [28, 36]\) with candidate band allocations \(b_i \in \{8, \dots, 14\}\), voxel target count \(V^* = |G'|\), and 2-opt neighborhood window \(W = |G'|^{1/3}\).

Key Experimental Results

Main Results

The performance of RaPTGS was benchmarked against uncompressed 3DGS, fine-tuning-free methods (FlexGaussian, MesonGS), and fine-tuning-based methods (MesonGS-FT, PUP 3D-GS) across Mip-NeRF 360, Tanks & Temples, and Deep Blending:

Dataset Metric Ours (RaPTGS) 3D-GS (Uncompressed) FlexGaussian (No-FT) MesonGS (No-FT) MesonGS-FT (With FT)
Mip-NeRF 360 PSNR↑ / SSIM↑ / LPIPS↓ 26.41 / 0.782 / 0.252 27.09 / 0.805 / 0.226 26.31 / 0.775 / 0.254 26.30 / 0.785 / 0.250 26.98 / 0.799 / 0.240
Storage Size (MB)↓ 25.95 MB 795.26 MB 42.40 MB 28.29 MB 28.34 MB
Tanks & Temples PSNR↑ / SSIM↑ / LPIPS↓ 22.97 / 0.823 / 0.203 23.38 / 0.841 / 0.182 22.52 / 0.807 / 0.215 23.02 / 0.828 / 0.198 23.28 / 0.834 / 0.194
Storage Size (MB)↓ 13.64 MB 421.90 MB 16.31 MB 16.51 MB 16.56 MB
Deep Blending PSNR↑ / SSIM↑ / LPIPS↓ 29.36 / 0.899 / 0.248 29.52 / 0.902 / 0.243 28.65 / 0.887 / 0.265 29.22 / 0.898 / 0.250 29.45 / 0.902 / 0.247
Storage Size (MB)↓ 24.51 MB 703.77 MB 25.60 MB 27.62 MB 27.70 MB

In terms of encoding throughput (evaluated on an Intel Ultra 9 285K with RTX 5080), RaPTGS demonstrates dramatic speedups: - Mip-NeRF 360: FlexGaussian takes 42.44s, MesonGS takes 199.97s, while RaPTGS requires only 7.60s; - Tanks & Temples: FlexGaussian takes 30.25s, MesonGS takes 227.43s, while RaPTGS requires only 5.78s; - Deep Blending: FlexGaussian takes 44.20s, MesonGS takes 159.23s, while RaPTGS requires only 7.70s.

Against the feed-forward neural compression baseline FCGS, FCGS encounters out-of-memory (OOM) failures on all 30k models across all 13 evaluation scenes on a 16GB GPU. Even on 7k-iteration checkpoints, RaPTGS uses nearly 4.5x less memory (~2.8 GB vs. 10.5–13.5 GB) and yields models approximately 2x more compact (e.g., bonsai-7k: 8.61 MB vs. 17.78 MB).

Ablation Study

1. Cumulative Impact of Pipeline Stages (Mip-NeRF 360, \(p=0.45, B=32\)):

Stage PSNR (dB)↑ SSIM↑ LPIPS↓ Size (MB)↓ Time (s)↓ Note
Original 3D-GS 27.09 0.805 0.226 795.26 - Baseline uncompressed model
+ Pruning 26.57 0.793 0.238 437.39 0.18 Discards 45% of uninformative splats
+ VMC Refinement 26.74 0.795 0.237 437.39 0.21 Restores coverage (+0.17 dB) at zero byte cost
+ SH Quantization 26.40 0.788 0.249 119.93 3.31 Band-wise AC SH codebooks reduce size by 72%
+ Attribute Quantization 26.24 0.778 0.255 32.66 5.13 Fixed-point conversion of remaining attributes
+ Hilbert Sorting 26.24 0.778 0.255 25.39 5.19 Global spatial clustering boosts entropy coder
+ Windowed 2-Opt 26.24 0.778 0.255 24.55 6.08 Local trajectory smoothing cuts residual entropy
Full Pipeline 26.24 0.778 0.255 24.55 7.93 32.4x overall compression in under 8 seconds

2. Pruning Criteria and Aggregation Ablation (Mip-NeRF 360, \(p=0.45\)):

Variant PSNR (dB)↑ SSIM↑ LPIPS↓ Note
Only OS (Opacity) 25.77 0.781 0.241 Severe loss of high-frequency geometry
Only SAS (Anisotropy) 25.93 0.782 0.242 Fails to preserve flat and low-contrast regions
Only SAE (Appearance) 25.37 0.782 0.247 Unconstrained geometry and opacity noise
w/o OS 26.54 0.792 0.239 Slight degradation in compositing accuracy
w/o SAS 26.53 0.792 0.239 Degraded structural and edge preservation
w/o SAE 25.98 0.783 0.239 Severe drop (-0.59 dB) in view-dependent fidelity
SAE w/o clamp 26.46 0.791 0.240 Premature culling of tiny, detail-rich Gaussians
Sum Pooling 26.48 0.792 0.237 Specialized salient features diluted by averaging
Max Pooling (Ours) 26.57 0.793 0.238 Logical OR preserves complementary splat strengths

Key Findings

  • Multi-Criteria Synergy: Pruning via opacity alone achieves only 25.77 dB PSNR at \(p=0.45\), whereas combining opacity with shape anisotropy and appearance energy yields 26.57 dB (+0.80 dB). SAE alone accounts for a +0.59 dB boost over a model without appearance scoring. Max-pooling acts as an indispensable preservation filter, shielding any Gaussian that excels along at least one physical axis.
  • Analytical Coverage Recovery: VMC provides an instantaneous +0.17 dB PSNR recovery without modifying Gaussian counts or requiring iterative optimization. Its effectiveness is especially prominent under aggressive pruning (\(p > 0.4\)), preventing hole formation while conservative clamping prevents boundary dilation.
  • Spatial Serialization Compounding: Reordering primitives via Hilbert sorting followed by parallel 2-opt refinement reduces model footprint from 32.66 MB to 24.55 MB (a 25% relative reduction) without altering rendering metrics, demonstrating that minimizing \(\ell_1\) neighbor distances dramatically flattens residual entropy.

Highlights & Insights

  • Complete Emancipation from Viewpoint Dependencies: Shatters the convention that 3DGS compression must be supervised by multi-view camera rendering. By formulating significance purely from intrinsic geometry and SH energy, it unlocks rapid compression for uncalibrated, standalone 3D assets.
  • Closed-Form Physics-Based Mass Restoration: Voxel mass compensation proves that local optical density can be effectively repaired via analytical volumetric mass conservation, substituting multi-epoch fine-tuning loops with a fraction-of-a-second matrix calculation.
  • Production-Ready Efficiency: The parallelized windowed 2-opt and degree-wise K-Means execution compresses multi-million-Gaussian scenes in 6–8 seconds on consumer-grade hardware with less than 3 GB VRAM footprint, bypassing the massive memory exhaustion endemic to neural autoencoders.

Limitations & Future Work

  • Lack of Occlusion Awareness: Operating without rasterization renders the pipeline blind to inter-primitive line-of-sight occlusions. Interior or fully occluded Gaussians possessing high SH appearance energy may be mistakenly retained, which accounts for the performance gap compared to render-dependent methods like MesonGS at extreme pruning ratios (\(p > 0.6\)).
  • Upper Bound of Analytical Compensation: In severely sparse regimes where an entire voxel loses all constituent primitives, VMC cannot synthesize new primitives out of thin air, and its conservative clamping strictly limits the maximum expansion of adjacent Gaussians.
  • Future Directions: The authors suggest exploring "self-rendered synthetic views" where sparse viewpoints are rendered purely from the model itself to extract lightweight occlusion and visibility priors without violating the model-only operating constraint.
  • vs. LightGaussian & MesonGS: Both baselines require camera parameters and forward rendering passes over all training views to measure significance; MesonGS-FT further demands backpropagation fine-tuning. RaPTGS operates in a fraction of the time (7s vs. 200s) and functions entirely without camera poses or images.
  • vs. FCGS: FCGS applies a feed-forward autoencoder but requires massive GPU memory, triggering OOM errors on 16GB VRAM GPUs for 30k-iteration scenes. RaPTGS uses standard non-neural codecs, requiring only ~2.8 GB memory and achieving 2x smaller file sizes.
  • vs. Naive Opacity Pruning: Naive thresholding causes destructive blurring on thin geometries; RaPTGS successfully protects structural edges via shape anisotropy and fine textures via scale-clamped appearance energy.

Rating

  • Novelty: ⭐⭐⭐⭐☆ (Pioneering an effective render-agnostic, model-only compression formulation for 3DGS with analytical mass compensation)
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Comprehensive validation across three benchmark suites against fine-tuning-free, fine-tuned, video codec, and neural baselines)
  • Writing Quality: ⭐⭐⭐⭐⭐ (Exceptionally well-structured paper with crisp mathematical formulations, rigorous trade-off discussions, and clean visual figures)
  • Value: ⭐⭐⭐⭐⭐ (High practical utility for cloud 3D graphics, game asset pipelines, and edge device distribution)