Skip to content

EGGS: Explicitly Granular 3D Gaussian Splatting via Luma-Aware and Volume-Preserving Attribute Factorization

Conference: ECCV 2026
Paper: ECCV Official
Area: 3D Vision
Keywords: 3D Gaussian Splatting, Neural Rendering, 3DGS Compression, Attribute Factorization, Neural Fields

TL;DR

Targeting the massive storage footprint caused by the explicit attribute parameters in 3DGS, EGGS introduces luma-based spherical harmonics DC compression and volume-preserving scale factorization to reduce per-primitive explicit attributes to as few as 9 dimensions, compressing gigabyte-scale scenes into a few megabytes while sustaining over 140 FPS real-time rendering.

Background & Motivation

3D Gaussian Splatting (3DGS) has established a new paradigm in novel view synthesis through its explicit representation of anisotropic Gaussian primitives and tile-based differentiable rasterization, delivering remarkable visual fidelity alongside real-time rendering frame rates. Nonetheless, photorealistically capturing real-world scenes requires millions of Gaussian primitives. Each primitive explicitly stores learnable parameters including 3D center positions, opacity, quaternion rotation, 3D scale, and high-order Spherical Harmonics (SH) coefficients. Consequently, uncompressed scenes easily demand hundreds of megabytes to multiple gigabytes of storage, severely obstructing practical transmission and deployment on resource-constrained platforms such as mobile devices and XR headsets.

Existing strategies to reduce 3DGS storage overhead broadly follow three trajectories: pruning primitive counts via importance metrics, quantizing explicit attributes with vector quantization and entropy coding, and implicitly predicting attributes using neural fields. However, existing hybrid neural field representations encounter critical dilemmas when compressing essential geometry and appearance attributes. On one hand, naively factorizing the 3D scale vector into a 1D scalar base and unconstrained implicit residuals frequently suffers from severe optimization failures due to the unbounded divergence of the scale factor. On the other hand, fully eliminating the explicit DC component of SH coefficients causes noticeable color shift and structural blur.

The key insight of this work stems from the perceptual asymmetry between luminance and chrominance, coupled with the observation that geometric scale optimization stability depends strictly on volumetric bounds rather than per-axis freedom. Core idea: compress per-primitive explicit attributes to as few as 9 dimensions by factorizing the DC component into an explicit luma scalar and factorizing the scale into a volume-preserving magnitude, reconstructing the remaining attributes via lightweight global MLPs for ultra-compact storage and configurable rate control.

Method

Overall Architecture

The core objective of EGGS is to replace the dozens of explicit parameters per Gaussian primitive with an ultra-compact explicit parameter set combined with global lightweight neural decoders. For each Gaussian primitive \(G_i\), EGGS explicitly stores only four types of information: 3D center position \(p \in \mathbb{R}^3\), 1D scale magnitude \(s_m \in \mathbb{R}^1\), 1D DC luminance \(k^0_Y \in \mathbb{R}^1\), and a configurable \(D_L\)-dimensional local feature \(f_L \in \mathbb{R}^{D_L}\) (yielding a minimal footprint of only 9 explicit dimensions when \(D_L=4\)).

During the reconstruction stage, EGGS recovers all remaining implicit attributes using six lightweight, single-hidden-layer global MLPs (each featuring 64 hidden neurons). The center coordinate \(p\) undergoes positional encoding (PE) and is processed by MLP\(_\text{Pos}\) to extract an implicit geometric feature, which is concatenated with the explicit local feature \(f_L\) to form a hybrid feature \(f_H\). This hybrid feature serves as the shared driving representation: it feeds directly into MLP\(_\text{SH(Non-DC)}\), MLP\(_\text{Rot}\), and MLP\(_\text{Opacity}\) to estimate non-DC SH coefficients \(k^{1+}\), rotation quaternion \(r\), and opacity \(o\). Concurrently, \(f_H\) is routed into specialized DC chrominance and scale factor branches, restoring the full suite of geometric and radiative attributes required for standard tile-based differentiable rasterization.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Explicit Gaussian Input<br/>p ∈ R³, sm ∈ R¹, k0Y ∈ R¹, fL ∈ R^DL"] --> Enc["Feature Concatenation<br/>PE(p) via MLP_Pos concatenated with fL to form fH"]
    Enc --> D1["1. Luma-Based DC Representation<br/>fH feeds MLP_UV to predict k0_UV, joined with k0Y to RGB"]
    Enc --> D2["2. Volume-Preserving Scale Factorization<br/>fH feeds MLP_Scale to predict sf, zero-mean exp modulating sm"]
    Enc --> D3["3. Compact Hybrid Implicit Attribute Estimation<br/>fH decodes non-DC k1+, rotation r, and opacity o in parallel"]
    D1 --> Out["High-Fidelity 3DGS Output & Real-Time Rasterization"]
    D2 --> Out
    D3 --> Out

Key Designs

1. Luma-Based DC Representation: eliminating color redundancy while anchoring structure

The 0-th order DC component \(k^0\) of Spherical Harmonics defines the base radiance of each Gaussian primitive, fundamentally governing the scene's ambient luminance and primary tones. Prior hybrid compression approaches either explicitly store full 3D RGB DC values—incurring non-trivial storage cost—or rely entirely on implicit estimation, which leads to color divergence and degraded textures. Drawing inspiration from classical video coding standards, EGGS exploits human visual perceptual characteristics: the human eye exhibits high sensitivity to luminance details while tolerating substantial loss in chrominance.

EGGS converts the 3D RGB DC component into the YUV color space, explicitly storing only the dominant luma scalar \(k^0_Y \in \mathbb{R}^1\), which encapsulates structural contours, fine textures, and lighting gradients. The perceptually secondary chrominance components \(k^0_{UV} \in \mathbb{R}^2\) are implicitly predicted by the lightweight network MLP\(_\text{UV}\) conditioned on the hybrid feature \(f_H\). The full RGB DC attribute is then reconstructed by concatenating the explicit luma with the predicted chroma and transforming back to RGB: $\(k^0 = \text{YUV-to-RGB}([k^0_Y, \text{MLP}_\text{UV}(f_H)])\)$ This formulation reduces the explicit storage overhead of the crucial DC component from 3 dimensions to 1, preserving high-frequency visual fidelity while slashing two-thirds of the DC data payload.

2. Volume-Preserving Scale Factorization: preventing optimization divergence in 1D scale compression

In 3DGS, the 3D scale vector \(s \in \mathbb{R}^3\) governs anisotropic spatial extents. Naively factorizing scale into an explicit scalar base \(\gamma\) and unconstrained implicit residual vectors frequently leads to optimization failures, because the per-axis factors easily diverge during backpropagation, causing Gaussians to blow up or degenerate into numerical singularities. EGGS proposes a volume-preserving factorization scheme that decomposes scale into an explicit scalar magnitude \(s_m \in \mathbb{R}^1\) and a 3D factor vector \(s_f \in \mathbb{R}^3\) predicted by MLP\(_\text{Scale}\).

To eliminate scale divergence, EGGS enforces a zero-mean constraint across the spatial axes of the predicted factor \(s_f\): $\(s_i = s_m \cdot \exp(s_{f,i} - \bar{s}_f), \quad \bar{s}_f = \frac{1}{3} \sum_{j \in \{x,y,z\}} s_{f,j}\)$ Because the zero-mean exponent terms sum to zero across the three orthogonal axes, their product cancels out exactly: $\(\prod_{i \in \{x,y,z\}} s_i = s_m^3 \cdot \exp\left(\sum_{i \in \{x,y,z\}} (s_{f,i} - \bar{s}_f)\right) = s_m^3\)$ This mathematical identity guarantees that the total physical volume of every Gaussian primitive is strictly constrained to \(s_m^3\). While the network can freely optimize anisotropic aspect ratios via \(s_f\), the total volume remains rigorously anchored to the explicit parameter \(s_m\), providing complete numerical stability throughout training.

3. Compact Hybrid Implicit Attribute Estimation and Configurable Rate Control: balancing minimal size with granular fidelity

Beyond scale and DC color, all remaining primitive attributes—non-DC higher-order SH coefficients \(k^{1+}\), rotation quaternion \(r\), and opacity \(o\)—are decoded in parallel by dedicated single-hidden-layer MLPs driven by the hybrid feature \(f_H = [f_L, \text{MLP}_\text{Pos}(\text{PE}(p))]\). All decoders maintain a compact footprint with only 64 hidden units. MLP\(_\text{Pos}\) employs ReLU activations, while all other MLPs use LeakyReLU. To further guarantee bounded behavior, MLP\(_\text{Scale}\) utilizes a scaled tanh activation (\(3 \cdot \tanh(\cdot)\)) on its output.

For storage serialization, EGGS encodes center coordinates \(p\) with MPEG G-PCC, applies Vector Quantization (VQ) to \(s_m\) and \(k^0_Y\), and compresses local features \(f_L\) with Sub-Vector Quantization (SVQ), followed by Huffman coding and LZMA archive packaging. Crucially, EGGS treats the local feature dimension \(D_L\) as a flexible rate-control hyperparameter varying from 4 to 12. Rate-distortion analyses reveal that in the high-fidelity regime, tuning \(D_L\) (e.g., from 12 to 8) provides significantly better reconstruction quality than aggressively pruning primitives, preserving spatial geometric density while enabling continuous rate adaptation.

Loss & Training

EGGS is built on top of Mini-Splatting2 and trained for 20,000 iterations. The first 8,000 iterations optimize Gaussian primitives using standard densification and simplification to construct a stable initial geometry. At iteration 8,000, explicit and implicit attribute factorizations are introduced. Cumulative importance-based primitive simplification occurs at iteration 13,000. Starting at iteration 19,000, VQ and SVQ codebooks are initialized and fine-tuned until convergence.

Key Experimental Results

Main Results

On the Mip-NeRF360, Tanks&Temples, and Deep Blending benchmarks, EGGS was extensively evaluated against state-of-the-art 3DGS compression baselines (measured on a single NVIDIA RTX 4000 Ada GPU):

Dataset Method SSIM ↑ PSNR (dB) ↑ LPIPS ↓ Size (MB) ↓ FPS ↑
Mip-NeRF 360 3DGS (Baseline) 0.809 27.26 0.221 641.04 60
Scaffold-GS 0.811 27.73 0.227 170.03 67
Mini-Splatting2 0.814 27.21 0.224 146.04 162
HAC++ (highrate) 0.810 27.79 0.231 18.05 63
LocoGS-L 0.809 27.34 0.226 13.68 100
OMG-XL 0.818 27.28 0.218 6.82 119
EGGS-XS (Ours) 0.793 26.47 0.254 2.67 189
EGGS-L (Ours) 0.817 27.45 0.221 6.21 140
Tanks & Temples 3DGS (Baseline) 0.852 23.75 0.169 371.32 92
Mini-Splatting2 0.841 23.15 0.185 84.03 301
HAC++ (highrate) 0.852 24.24 0.176 10.07 84
LocoGS-L 0.843 23.75 0.192 10.68 144
OMG-XL 0.847 23.63 0.178 4.26 223
EGGS-XS (Ours) 0.812 22.46 0.231 1.33 389
EGGS-L (Ours) 0.842 23.74 0.192 3.14 287
Deep Blending 3DGS (Baseline) 0.907 29.81 0.237 581.95 66
Mini-Splatting2 0.912 30.00 0.241 153.53 255
HAC++ (highrate) 0.906 30.13 0.257 6.43 103
LocoGS-L 0.905 30.19 0.245 12.99 122
OMG-XL 0.909 29.87 0.246 5.56 221
EGGS-XS (Ours) 0.904 29.77 0.262 2.53 220
EGGS-L (Ours) 0.911 30.21 0.242 5.45 197

Ablation Study

Table 1: Ablation on Explicit DC Representation Dimensionality

Explicit DC Config SSIM ↑ PSNR (dB) ↑ LPIPS ↓ Size (MB) ↓ Note
0-dim (fully implicit estimation) 0.965 26.59 0.242 3.87 Lacks explicit baseline anchor; PSNR drops by 0.43 dB
1-dim (Ours, Luma-based YUV) 0.967 27.02 0.226 4.62 Approaches 3-dim quality (within 0.08 dB) with minimal size gain
3-dim (standard explicit RGB) 0.967 27.10 0.223 5.17 Adds 12% storage overhead for marginal visual improvement

Table 2: Stability Analysis of Volume-Preserving Scale Factorization (10 trials across 13 scenes)

Scale Factorization Strategy Success Rate (%) ↑ Mean PSNR (dB) ↑ Mean SSIM ↑ Size (MB) ↓ Note
Naive 1D Factorization Baseline 73.8% (30% on Train) 27.20 0.827 4.86 Unbounded implicit scale divergence causes frequent crash
Volume-Preserving Factorization (Ours) 100.0% (all 13 scenes) 27.18 0.833 4.62 Zero-mean exponential cancellation ensures perfect training stability

Key Findings

  • Storage Footprint Record: EGGS-XS consistently achieves the smallest storage footprint across all three benchmark datasets (1.33 MB on Tanks&Temples, 2.67 MB on Mip-NeRF360), compressing original 3DGS models by 200× to 280× while rendering at 189 to 389 FPS.
  • Asymmetric Payoff of Luma Separation: Retaining a single explicit luma dimension delivers an immediate 0.43 dB PSNR boost over pure implicit estimation, reaching parity with full 3D RGB DC storage while cutting two-thirds of the explicit DC parameters.
  • Feature Tuning Outperforms Pruning: Rate-distortion sweeps across \(D_L \in [4, 12]\) confirm that in high-quality regimes, reducing feature dimensionality preserves visual coherence and geometric continuity much better than aggressive primitive pruning.

Highlights & Insights

  • Cross-Disciplinary Perceptual Compression in 3DGS: EGGS translates the foundational YUV luminance/chrominance separation principle from classical video compression into 3D Gaussian radiance modeling, demonstrating that structural fidelity can be anchored by a 1D explicit scalar.
  • Volume-Preserving Mathematical Invariant: By enforcing zero-mean centering on predicted anisotropic factors, the exponential components cancel out in product space, guaranteeing that Gaussian volume remains mathematically fixed to \(s_m^3\) without constraining orientation or aspect ratio.
  • Flexible Rate Control Paradigm: The paper demonstrates that adjusting per-primitive local embedding capacity \(D_L\) provides a smoother and more fidelity-preserving rate-distortion trajectory than point pruning, offering a valuable insight for 3D streaming pipelines.

Limitations & Future Work

  • Author-Acknowledged Limitations: Under the extreme compression profile (EGGS-XS), subtle smoothing can be observed in high-frequency specular reflections and fine specular details (such as the truck window reflections and intricate petals in Flowers).
  • Noted Technical Trade-Offs: Restoring full 3DGS attributes requires forward inference through the 6 global MLPs before rasterization. While total inference latency is negligible, deployment on bare rasterizers without compute shader support necessitates an initial decompression baking step into system memory.
  • Potential Improvement Directions: Future investigations could integrate the lightweight MLP decoders directly into custom WebGPU / CUDA rasterization compute kernels for on-the-fly streaming rendering without global unbaking.
  • vs LocoGS: LocoGS attempts 1D base scale factorization with hash grids but lacks volume constraints, resulting in severe optimization failures on complex scenes (e.g. 30% success rate on Train); EGGS resolves this instability via volume preservation (100% success rate) while requiring roughly half the file size.
  • vs OMG: OMG explicitly stores multi-dimensional latent codes \(T\) and \(V\) for attribute reconstruction; EGGS pushes explicit attributes down to scalar luma and scalar scale magnitude, yielding a significantly smaller footprint (1.33–2.67 MB vs. 2.44–4.04 MB) and higher frame rates.
  • vs HAC++: HAC++ uses complex anchor-based context entropy models and reaches high PSNR, but incurs high decoding latency (63–93 FPS); EGGS avoids iterative autoregressive entropy decoding, achieving 140–389 FPS real-time rendering.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Elegant adaptation of YUV luma separation and volume-preserving invariance to 3DGS attribute compression.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous multi-trial stability stress testing across 13 scenes alongside complete ablation studies.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear motivation, clean mathematical formulation, and well-structured rate-distortion analyses.
  • Value: ⭐⭐⭐⭐⭐ Slashes gigabyte-scale 3DGS scenes to single-digit megabytes while sustaining over 140 FPS, delivering substantial real-world deployment value.