EGGS: Explicitly Granular 3D Gaussian Splatting via Luma-Aware and Volume-Preserving Attribute Factorization¶
Conference: ECCV 2026
Paper: ECCV Official
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Neural Rendering, 3DGS Compression, Attribute Factorization, Neural Fields
TL;DR¶
Targeting the massive storage footprint caused by the explicit attribute parameters in 3DGS, EGGS introduces luma-based spherical harmonics DC compression and volume-preserving scale factorization to reduce per-primitive explicit attributes to as few as 9 dimensions, compressing gigabyte-scale scenes into a few megabytes while sustaining over 140 FPS real-time rendering.
Background & Motivation¶
3D Gaussian Splatting (3DGS) has established a new paradigm in novel view synthesis through its explicit representation of anisotropic Gaussian primitives and tile-based differentiable rasterization, delivering remarkable visual fidelity alongside real-time rendering frame rates. Nonetheless, photorealistically capturing real-world scenes requires millions of Gaussian primitives. Each primitive explicitly stores learnable parameters including 3D center positions, opacity, quaternion rotation, 3D scale, and high-order Spherical Harmonics (SH) coefficients. Consequently, uncompressed scenes easily demand hundreds of megabytes to multiple gigabytes of storage, severely obstructing practical transmission and deployment on resource-constrained platforms such as mobile devices and XR headsets.
Existing strategies to reduce 3DGS storage overhead broadly follow three trajectories: pruning primitive counts via importance metrics, quantizing explicit attributes with vector quantization and entropy coding, and implicitly predicting attributes using neural fields. However, existing hybrid neural field representations encounter critical dilemmas when compressing essential geometry and appearance attributes. On one hand, naively factorizing the 3D scale vector into a 1D scalar base and unconstrained implicit residuals frequently suffers from severe optimization failures due to the unbounded divergence of the scale factor. On the other hand, fully eliminating the explicit DC component of SH coefficients causes noticeable color shift and structural blur.
The key insight of this work stems from the perceptual asymmetry between luminance and chrominance, coupled with the observation that geometric scale optimization stability depends strictly on volumetric bounds rather than per-axis freedom. Core idea: compress per-primitive explicit attributes to as few as 9 dimensions by factorizing the DC component into an explicit luma scalar and factorizing the scale into a volume-preserving magnitude, reconstructing the remaining attributes via lightweight global MLPs for ultra-compact storage and configurable rate control.
Method¶
Overall Architecture¶
The core objective of EGGS is to replace the dozens of explicit parameters per Gaussian primitive with an ultra-compact explicit parameter set combined with global lightweight neural decoders. For each Gaussian primitive \(G_i\), EGGS explicitly stores only four types of information: 3D center position \(p \in \mathbb{R}^3\), 1D scale magnitude \(s_m \in \mathbb{R}^1\), 1D DC luminance \(k^0_Y \in \mathbb{R}^1\), and a configurable \(D_L\)-dimensional local feature \(f_L \in \mathbb{R}^{D_L}\) (yielding a minimal footprint of only 9 explicit dimensions when \(D_L=4\)).
During the reconstruction stage, EGGS recovers all remaining implicit attributes using six lightweight, single-hidden-layer global MLPs (each featuring 64 hidden neurons). The center coordinate \(p\) undergoes positional encoding (PE) and is processed by MLP\(_\text{Pos}\) to extract an implicit geometric feature, which is concatenated with the explicit local feature \(f_L\) to form a hybrid feature \(f_H\). This hybrid feature serves as the shared driving representation: it feeds directly into MLP\(_\text{SH(Non-DC)}\), MLP\(_\text{Rot}\), and MLP\(_\text{Opacity}\) to estimate non-DC SH coefficients \(k^{1+}\), rotation quaternion \(r\), and opacity \(o\). Concurrently, \(f_H\) is routed into specialized DC chrominance and scale factor branches, restoring the full suite of geometric and radiative attributes required for standard tile-based differentiable rasterization.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Explicit Gaussian Input<br/>p ∈ R³, sm ∈ R¹, k0Y ∈ R¹, fL ∈ R^DL"] --> Enc["Feature Concatenation<br/>PE(p) via MLP_Pos concatenated with fL to form fH"]
Enc --> D1["1. Luma-Based DC Representation<br/>fH feeds MLP_UV to predict k0_UV, joined with k0Y to RGB"]
Enc --> D2["2. Volume-Preserving Scale Factorization<br/>fH feeds MLP_Scale to predict sf, zero-mean exp modulating sm"]
Enc --> D3["3. Compact Hybrid Implicit Attribute Estimation<br/>fH decodes non-DC k1+, rotation r, and opacity o in parallel"]
D1 --> Out["High-Fidelity 3DGS Output & Real-Time Rasterization"]
D2 --> Out
D3 --> Out
Key Designs¶
1. Luma-Based DC Representation: eliminating color redundancy while anchoring structure
The 0-th order DC component \(k^0\) of Spherical Harmonics defines the base radiance of each Gaussian primitive, fundamentally governing the scene's ambient luminance and primary tones. Prior hybrid compression approaches either explicitly store full 3D RGB DC values—incurring non-trivial storage cost—or rely entirely on implicit estimation, which leads to color divergence and degraded textures. Drawing inspiration from classical video coding standards, EGGS exploits human visual perceptual characteristics: the human eye exhibits high sensitivity to luminance details while tolerating substantial loss in chrominance.
EGGS converts the 3D RGB DC component into the YUV color space, explicitly storing only the dominant luma scalar \(k^0_Y \in \mathbb{R}^1\), which encapsulates structural contours, fine textures, and lighting gradients. The perceptually secondary chrominance components \(k^0_{UV} \in \mathbb{R}^2\) are implicitly predicted by the lightweight network MLP\(_\text{UV}\) conditioned on the hybrid feature \(f_H\). The full RGB DC attribute is then reconstructed by concatenating the explicit luma with the predicted chroma and transforming back to RGB: $\(k^0 = \text{YUV-to-RGB}([k^0_Y, \text{MLP}_\text{UV}(f_H)])\)$ This formulation reduces the explicit storage overhead of the crucial DC component from 3 dimensions to 1, preserving high-frequency visual fidelity while slashing two-thirds of the DC data payload.
2. Volume-Preserving Scale Factorization: preventing optimization divergence in 1D scale compression
In 3DGS, the 3D scale vector \(s \in \mathbb{R}^3\) governs anisotropic spatial extents. Naively factorizing scale into an explicit scalar base \(\gamma\) and unconstrained implicit residual vectors frequently leads to optimization failures, because the per-axis factors easily diverge during backpropagation, causing Gaussians to blow up or degenerate into numerical singularities. EGGS proposes a volume-preserving factorization scheme that decomposes scale into an explicit scalar magnitude \(s_m \in \mathbb{R}^1\) and a 3D factor vector \(s_f \in \mathbb{R}^3\) predicted by MLP\(_\text{Scale}\).
To eliminate scale divergence, EGGS enforces a zero-mean constraint across the spatial axes of the predicted factor \(s_f\): $\(s_i = s_m \cdot \exp(s_{f,i} - \bar{s}_f), \quad \bar{s}_f = \frac{1}{3} \sum_{j \in \{x,y,z\}} s_{f,j}\)$ Because the zero-mean exponent terms sum to zero across the three orthogonal axes, their product cancels out exactly: $\(\prod_{i \in \{x,y,z\}} s_i = s_m^3 \cdot \exp\left(\sum_{i \in \{x,y,z\}} (s_{f,i} - \bar{s}_f)\right) = s_m^3\)$ This mathematical identity guarantees that the total physical volume of every Gaussian primitive is strictly constrained to \(s_m^3\). While the network can freely optimize anisotropic aspect ratios via \(s_f\), the total volume remains rigorously anchored to the explicit parameter \(s_m\), providing complete numerical stability throughout training.
3. Compact Hybrid Implicit Attribute Estimation and Configurable Rate Control: balancing minimal size with granular fidelity
Beyond scale and DC color, all remaining primitive attributes—non-DC higher-order SH coefficients \(k^{1+}\), rotation quaternion \(r\), and opacity \(o\)—are decoded in parallel by dedicated single-hidden-layer MLPs driven by the hybrid feature \(f_H = [f_L, \text{MLP}_\text{Pos}(\text{PE}(p))]\). All decoders maintain a compact footprint with only 64 hidden units. MLP\(_\text{Pos}\) employs ReLU activations, while all other MLPs use LeakyReLU. To further guarantee bounded behavior, MLP\(_\text{Scale}\) utilizes a scaled tanh activation (\(3 \cdot \tanh(\cdot)\)) on its output.
For storage serialization, EGGS encodes center coordinates \(p\) with MPEG G-PCC, applies Vector Quantization (VQ) to \(s_m\) and \(k^0_Y\), and compresses local features \(f_L\) with Sub-Vector Quantization (SVQ), followed by Huffman coding and LZMA archive packaging. Crucially, EGGS treats the local feature dimension \(D_L\) as a flexible rate-control hyperparameter varying from 4 to 12. Rate-distortion analyses reveal that in the high-fidelity regime, tuning \(D_L\) (e.g., from 12 to 8) provides significantly better reconstruction quality than aggressively pruning primitives, preserving spatial geometric density while enabling continuous rate adaptation.
Loss & Training¶
EGGS is built on top of Mini-Splatting2 and trained for 20,000 iterations. The first 8,000 iterations optimize Gaussian primitives using standard densification and simplification to construct a stable initial geometry. At iteration 8,000, explicit and implicit attribute factorizations are introduced. Cumulative importance-based primitive simplification occurs at iteration 13,000. Starting at iteration 19,000, VQ and SVQ codebooks are initialized and fine-tuned until convergence.
Key Experimental Results¶
Main Results¶
On the Mip-NeRF360, Tanks&Temples, and Deep Blending benchmarks, EGGS was extensively evaluated against state-of-the-art 3DGS compression baselines (measured on a single NVIDIA RTX 4000 Ada GPU):
| Dataset | Method | SSIM ↑ | PSNR (dB) ↑ | LPIPS ↓ | Size (MB) ↓ | FPS ↑ |
|---|---|---|---|---|---|---|
| Mip-NeRF 360 | 3DGS (Baseline) | 0.809 | 27.26 | 0.221 | 641.04 | 60 |
| Scaffold-GS | 0.811 | 27.73 | 0.227 | 170.03 | 67 | |
| Mini-Splatting2 | 0.814 | 27.21 | 0.224 | 146.04 | 162 | |
| HAC++ (highrate) | 0.810 | 27.79 | 0.231 | 18.05 | 63 | |
| LocoGS-L | 0.809 | 27.34 | 0.226 | 13.68 | 100 | |
| OMG-XL | 0.818 | 27.28 | 0.218 | 6.82 | 119 | |
| EGGS-XS (Ours) | 0.793 | 26.47 | 0.254 | 2.67 | 189 | |
| EGGS-L (Ours) | 0.817 | 27.45 | 0.221 | 6.21 | 140 | |
| Tanks & Temples | 3DGS (Baseline) | 0.852 | 23.75 | 0.169 | 371.32 | 92 |
| Mini-Splatting2 | 0.841 | 23.15 | 0.185 | 84.03 | 301 | |
| HAC++ (highrate) | 0.852 | 24.24 | 0.176 | 10.07 | 84 | |
| LocoGS-L | 0.843 | 23.75 | 0.192 | 10.68 | 144 | |
| OMG-XL | 0.847 | 23.63 | 0.178 | 4.26 | 223 | |
| EGGS-XS (Ours) | 0.812 | 22.46 | 0.231 | 1.33 | 389 | |
| EGGS-L (Ours) | 0.842 | 23.74 | 0.192 | 3.14 | 287 | |
| Deep Blending | 3DGS (Baseline) | 0.907 | 29.81 | 0.237 | 581.95 | 66 |
| Mini-Splatting2 | 0.912 | 30.00 | 0.241 | 153.53 | 255 | |
| HAC++ (highrate) | 0.906 | 30.13 | 0.257 | 6.43 | 103 | |
| LocoGS-L | 0.905 | 30.19 | 0.245 | 12.99 | 122 | |
| OMG-XL | 0.909 | 29.87 | 0.246 | 5.56 | 221 | |
| EGGS-XS (Ours) | 0.904 | 29.77 | 0.262 | 2.53 | 220 | |
| EGGS-L (Ours) | 0.911 | 30.21 | 0.242 | 5.45 | 197 |
Ablation Study¶
Table 1: Ablation on Explicit DC Representation Dimensionality¶
| Explicit DC Config | SSIM ↑ | PSNR (dB) ↑ | LPIPS ↓ | Size (MB) ↓ | Note |
|---|---|---|---|---|---|
| 0-dim (fully implicit estimation) | 0.965 | 26.59 | 0.242 | 3.87 | Lacks explicit baseline anchor; PSNR drops by 0.43 dB |
| 1-dim (Ours, Luma-based YUV) | 0.967 | 27.02 | 0.226 | 4.62 | Approaches 3-dim quality (within 0.08 dB) with minimal size gain |
| 3-dim (standard explicit RGB) | 0.967 | 27.10 | 0.223 | 5.17 | Adds 12% storage overhead for marginal visual improvement |
Table 2: Stability Analysis of Volume-Preserving Scale Factorization (10 trials across 13 scenes)¶
| Scale Factorization Strategy | Success Rate (%) ↑ | Mean PSNR (dB) ↑ | Mean SSIM ↑ | Size (MB) ↓ | Note |
|---|---|---|---|---|---|
| Naive 1D Factorization Baseline | 73.8% (30% on Train) | 27.20 | 0.827 | 4.86 | Unbounded implicit scale divergence causes frequent crash |
| Volume-Preserving Factorization (Ours) | 100.0% (all 13 scenes) | 27.18 | 0.833 | 4.62 | Zero-mean exponential cancellation ensures perfect training stability |
Key Findings¶
- Storage Footprint Record: EGGS-XS consistently achieves the smallest storage footprint across all three benchmark datasets (1.33 MB on Tanks&Temples, 2.67 MB on Mip-NeRF360), compressing original 3DGS models by 200× to 280× while rendering at 189 to 389 FPS.
- Asymmetric Payoff of Luma Separation: Retaining a single explicit luma dimension delivers an immediate 0.43 dB PSNR boost over pure implicit estimation, reaching parity with full 3D RGB DC storage while cutting two-thirds of the explicit DC parameters.
- Feature Tuning Outperforms Pruning: Rate-distortion sweeps across \(D_L \in [4, 12]\) confirm that in high-quality regimes, reducing feature dimensionality preserves visual coherence and geometric continuity much better than aggressive primitive pruning.
Highlights & Insights¶
- Cross-Disciplinary Perceptual Compression in 3DGS: EGGS translates the foundational YUV luminance/chrominance separation principle from classical video compression into 3D Gaussian radiance modeling, demonstrating that structural fidelity can be anchored by a 1D explicit scalar.
- Volume-Preserving Mathematical Invariant: By enforcing zero-mean centering on predicted anisotropic factors, the exponential components cancel out in product space, guaranteeing that Gaussian volume remains mathematically fixed to \(s_m^3\) without constraining orientation or aspect ratio.
- Flexible Rate Control Paradigm: The paper demonstrates that adjusting per-primitive local embedding capacity \(D_L\) provides a smoother and more fidelity-preserving rate-distortion trajectory than point pruning, offering a valuable insight for 3D streaming pipelines.
Limitations & Future Work¶
- Author-Acknowledged Limitations: Under the extreme compression profile (EGGS-XS), subtle smoothing can be observed in high-frequency specular reflections and fine specular details (such as the truck window reflections and intricate petals in Flowers).
- Noted Technical Trade-Offs: Restoring full 3DGS attributes requires forward inference through the 6 global MLPs before rasterization. While total inference latency is negligible, deployment on bare rasterizers without compute shader support necessitates an initial decompression baking step into system memory.
- Potential Improvement Directions: Future investigations could integrate the lightweight MLP decoders directly into custom WebGPU / CUDA rasterization compute kernels for on-the-fly streaming rendering without global unbaking.
Related Work & Insights¶
- vs LocoGS: LocoGS attempts 1D base scale factorization with hash grids but lacks volume constraints, resulting in severe optimization failures on complex scenes (e.g. 30% success rate on Train); EGGS resolves this instability via volume preservation (100% success rate) while requiring roughly half the file size.
- vs OMG: OMG explicitly stores multi-dimensional latent codes \(T\) and \(V\) for attribute reconstruction; EGGS pushes explicit attributes down to scalar luma and scalar scale magnitude, yielding a significantly smaller footprint (1.33–2.67 MB vs. 2.44–4.04 MB) and higher frame rates.
- vs HAC++: HAC++ uses complex anchor-based context entropy models and reaches high PSNR, but incurs high decoding latency (63–93 FPS); EGGS avoids iterative autoregressive entropy decoding, achieving 140–389 FPS real-time rendering.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Elegant adaptation of YUV luma separation and volume-preserving invariance to 3DGS attribute compression.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous multi-trial stability stress testing across 13 scenes alongside complete ablation studies.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear motivation, clean mathematical formulation, and well-structured rate-distortion analyses.
- Value: ⭐⭐⭐⭐⭐ Slashes gigabyte-scale 3DGS scenes to single-digit megabytes while sustaining over 140 FPS, delivering substantial real-world deployment value.