NRF-GS: Neural Residual Fields for Expressive and Compact Gaussian Splatting¶
Conference: NeurIPS 2026 (Accepted, according to the assigned list)
arXiv: 2609.37115
Area: 3D Vision
Keywords: Gaussian Splatting, neural residual fields, view-dependent appearance, frequency decomposition, compact scene representation
TL;DR¶
NRF-GS replaces per-Gaussian spherical-harmonic appearance with shared low-/high-frequency neural residual branches, achieving higher average PSNR with fewer Gaussians under unchanged densification and pruning rules, but not uniformly better perceptual quality or rendering speed.
Background & Motivation¶
The efficiency of 3D Gaussian Splatting (3DGS) comes from explicit geometry and differentiable rasterization: each Gaussian represents a spatial region, its color varies with viewing direction, and contributions are composited in depth order. Standard appearance modeling typically uses degree-3 spherical harmonics (SH), storing 48 color coefficients per Gaussian. This is inexpensive to evaluate but struggles with narrow highlights and complex directional changes. When individual Gaussians cannot fit these changes, optimization may compensate by adding Gaussians, creating redundancy that is not determined solely by geometric detail.
Adding a neural network does not automatically resolve this issue. VDGS conditions prediction on attributes including Gaussian and camera positions; GSNB supplements SH with neural bases, increasing appearance capacity alongside storage and computation. NRF-GS makes a narrower intervention: it retains Gaussian positions and covariances while assigning a shared network only the directional residual beyond each Gaussian's base color. Compact latent features distinguish local appearances instead of giving every Gaussian an independent full directional function.
Directional residuals can still overfit, particularly in outdoor scenes with sparse view coverage. Rather than feeding all Fourier bands into one network, the method preserves a stable low-frequency path and adaptively modulates high-frequency contributions for each Gaussian and current view. Core idea: increase each Gaussian's directional appearance capacity while using shared, bounded, frequency-decomposed residuals to reduce the need for densification that compensates for inadequate appearance modeling.
Method¶
Overall Architecture¶
The input consists of multiview images with camera poses. The scene remains an explicit Gaussian representation, with each Gaussian storing its position, covariance, base color, base opacity, and a 32-dimensional latent feature. For the current camera, the method computes the cameraโGaussian direction and distance, then splits the directional encoding into low- and high-frequency groups.
The shared residual representation feeds the same latent feature into two branches: the low-frequency branch predicts a smooth color residual and an opacity residual; the high-frequency branch predicts a detailed color residual and a modulation weight. Bounded residual composition adds these predictions to the base appearance, after which the original Gaussian rasterizer produces the image. Training supervision comes from rendered versus observed images, not per-Gaussian material labels; inference requires only the optimized scene and target camera.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Gaussian attributes + target camera"] --> B["Shared residual representation"]
B --> C["Low-frequency branch"]
B --> D["High-frequency branch"]
C -->|Color and opacity residuals| E["Bounded residual composition"]
D -->|Color residual and modulation weight| E
E --> F["Standard rasterizer โ image"]
G["Training: observed images"] -.->|Pixel supervision and back-propagation| F
Key Designs¶
1. Shared residual representation: store base color separately from directional variation
NRF-GS does not regenerate the entire scene with a network or predict Gaussian geometry. Each Gaussian retains its explicit position and covariance, while base color represents the dominant view-independent diffuse component. The network learns only corrections relative to this base color, allowing information shared across views to settle into a stable base term and reserving neural capacity for directional changes. This is a division of appearance parameterization, not a physically constrained BRDF or a relightable material decomposition.
Each Gaussian's 32-dimensional latent feature is jointly optimized during training and decoded by a multilayer perceptron (MLP) shared across the scene. Function parameters are shared, not Gaussian colors or latent features; different features allow the same function to express different directional responses on different surfaces. The network does not directly receive the full covariance and other geometric attributes, but its outputs still affect geometry optimization through the rendering loss. Thus, โdecouplingโ refers to representation structure rather than complete gradient isolation.
The current direction points from the Gaussian center toward the camera center and is normalized; distance is separately supplied as a logarithmic scalar, with a small constant preventing numerical problems at zero distance. Distance provides a cue about observation scale and parallax, allowing the same directional encoding to receive different appearance corrections nearby and far away; it is not additional measured depth supervision. The paper reports reconstruction degradation when distance is removed but does not establish that this input precisely follows a physical reflection law.
The budget must include both per-Gaussian latent features and shared decoders. The main text counts 32 latent dimensions, 3 base-color parameters, and 1 base-opacity parameter per Gaussian, totaling 36 appearance-related parameters, plus approximately 6.7k network parameters across both branches. Against 48 SH coefficients, the authors report approximately 25% lower per-Gaussian appearance parameter count. This is not a complete scene-byte budget: geometry, precision, optimizer states, and the baseline opacity-counting convention require separate accounting. The โ6.7k networkโ must not be treated as the storage cost of the whole scene.
2. Low-frequency branch: jointly refine smooth directional appearance and opacity
The directional Fourier encoding uses 3 frequency bands. The low-frequency branch receives the raw direction and band 0, together with the latent feature, base color, base opacity, and logarithmic distance. It outputs an RGB residual and an intermediate quantity for updating opacity. This branch captures smoother, broader directional variation so that the high-frequency branch need not relearn the base trend.
Restricting opacity corrections to this branch is intentional. Color changes affect local appearance, whereas opacity changes also affect transmission and occlusion of Gaussians behind the current one. Allowing opacity to vary sharply with high-frequency directional signals can destabilize compositing and introduce flicker. The method retains view-dependent opacity adjustment but denies that capability to the high-frequency branch. Removing the opacity residual costs 0.35 dB in average PSNR in the ablation, although the paper does not independently quantify temporal flicker.
3. High-frequency branch: select detailed directional effects with learned gating
The high-frequency branch receives the remaining frequency bands and the same latent feature, base color, and distance. It neither receives base opacity nor predicts opacity. Alongside an RGB residual, it outputs a scalar modulation signal whose sigmoid controls the high-frequency color contribution. Modulation can therefore vary with the Gaussian and current input rather than imposing one fixed high-frequency ratio on the entire scene.
Why separate branches rather than simply enlarge the network? Low-frequency trends often have support from more views, while high-frequency directional effects may appear in only a few observations. Joint fitting can encourage the model to explain sparse signals as mandatory detail. Frequency separation with gating imposes a more appropriate structural bias: low-frequency variation remains represented, while local high-frequency variation contributes selectively. In the appendix, Bicycle and Stump show better training scores but worse test scores for the single branch, supporting the overfitting interpretation; this does not imply that high-frequency encoding is inherently harmful.
Both branches use 64 hidden units and ReLU, implemented with tiny-cuda-nn FullyFusedMLP. The unsplit ablation uses 128 hidden units and approximately 8.0k network parameters, slightly more than the full model's approximately 6.7k, so its degradation cannot simply be attributed to a smaller network. Conversely, this is not a matched-architecture experiment that removes only gating, and the paper does not separately isolate the gate's contribution.
4. Bounded residual composition: constrain corrections rather than regenerate appearance
The two color residuals are added after multiplying the high-frequency term by its modulation weight. A tanh nonlinearity and scale factor bound the combined residual before it is added to base color. The mechanism in the paper's Eq. (8) is:
The bound acts on the combined residual, not separately on each raw branch output. The color scale is 0.2, keeping network predictions near the base appearance. This reduces the risk of overwriting base color early in optimization, but also limits the magnitude of directional corrections.
Opacity is the base value plus a bounded low-frequency update, as specified by Eq. (9):
The opacity scale is also 0.2. A bounded update must be distinguished from valid final opacity: even if the base lies between 0 and 1, adding or subtracting a bounded residual does not automatically keep the result inside that interval. The cached main text does not specify final clamping or reparameterization, so no such implementation step is added here. Corrected color and opacity enter the original rasterization and front-to-back compositing pipeline, with geometry representation and densification/pruning rules unchanged.
A Worked Example¶
Consider a Gaussian representing a shiny surface in Kitchen. When the camera moves, the neural forward pass does not replace its position or covariance, but the relative direction and distance change, altering both branches' inputs. Base color supplies a stable foundation, the low-frequency branch corrects broad color variation and opacity, and the high-frequency branch strengthens detailed color changes at relevant views. Bounded residual composition combines these predictions before depth sorting and pixel compositing.
During training, the composited image must still match the observed view, and its error optimizes explicit scene parameters, latent features, and the shared network together. Better directional fitting may reduce the need triggered by densification rules, eventually producing fewer Gaussians without rewriting the densification criteria. Appendix Table 6 reports average Kitchen point counts of 1.818M for 3DGS versus 0.652M for NRF-GS, with PSNR of 31.46 versus 32.10. These are whole-scene optimization results, not measurements of the illustrative individual Gaussian.
Loss & Training¶
The paper jointly optimizes explicit Gaussians, latent features, and the shared residual network within the standard 3DGS training pipeline, retaining the original densification and pruning heuristics. Pixel-level reconstruction supervision propagates through the differentiable rasterizer. The cache does not state a separate new NRF-GS loss formula or complete loss weights, so customary 3DGS objectives are not presented here as explicitly specified author configurations.
A short warmup runs once per view before regular training, encouraging zero color residual and neutral opacity behavior. Its purpose is to start from view-independent base appearance rather than fit directional noise prematurely. The paper does not detail the warmup loss, optimizer learning rates, or training iteration count; reproduction still requires implementation documentation.
Key Experimental Results¶
Main Results¶
The main table below follows the paper's Table 3. Point counts are in millions and times are training minutes; higher PSNR and lower LPIPS are better. MipNeRF360 contains 9 scenes, DL3DV uses 9 randomly selected scenes from its 10k scenes, and Tanks and Temples (TnT) uses 19 scenes, with the reported 1:8 split. Experiments use an RTX 3090 with 24GB memory, a 24-core CPU, and 64GB RAM; MipNeRF360 is approximately 1400p and the others approximately 960p.
| Dataset | Method | Points M | Training min | PSNR | LPIPS |
|---|---|---|---|---|---|
| MipNeRF360 | 3DGS | 3.359 | 26.46 | 27.40 | 0.184 |
| MipNeRF360 | VDGS | 3.548 | 41.27 | 27.65 | 0.186 |
| MipNeRF360 | NRF-GS | 1.656 | 31.38 | 27.86 | 0.191 |
| MipNeRF360* | 3DGS | 2.605 | 24.22 | 27.74 | 0.199 |
| MipNeRF360* | VDGS | 2.828 | 36.70 | 28.03 | 0.200 |
| MipNeRF360* | GSNB | 2.667 | 59.06 | 27.71 | 0.201 |
| MipNeRF360* | NRF-GS | 1.361 | 28.83 | 28.27 | 0.205 |
| DL3DV | 3DGS | 1.158 | 12.74 | 29.48 | 0.088 |
| DL3DV | VDGS | 1.404 | 21.15 | 29.92 | 0.089 |
| DL3DV | GSNB | 1.032 | 25.28 | 30.60 | 0.086 |
| DL3DV | NRF-GS | 0.595 | 14.52 | 30.82 | 0.089 |
| TnT | 3DGS | 1.868 | 15.35 | 23.85 | 0.165 |
| TnT | VDGS | 2.113 | 26.37 | 24.12 | 0.163 |
| TnT | GSNB | 1.705 | 35.75 | 23.92 | 0.164 |
| TnT | NRF-GS | 1.040 | 19.91 | 24.68 | 0.164 |
MipNeRF360* excludes Bicycle and Garden because GSNB runs out of memory; its averages must not be mixed directly with those over all 9 scenes. Relative to 3DGS, point counts decrease by approximately 50.7%, 48.6%, and 44.3% on the three full datasets, while PSNR increases by 0.46, 1.34, and 0.83 dB. Average inference speeds are 105 FPS for NRF-GS, 115 FPS for 3DGS, 65 FPS for VDGS, and 92 FPS for GSNB: compactness does not imply faster rendering.
Ablation Study¶
The MipNeRF360 ablation follows the paper's Table 1. Its SSIM column label is retained rather than silently changed to the main table's MS-SSIM.
| Config | PSNR | SSIM | LPIPS | L1 |
|---|---|---|---|---|
| Full model | 27.86 | 0.814 | 0.191 | 0.029 |
| Without low-/high-frequency split | 27.38 | 0.796 | 0.211 | 0.032 |
| Without opacity residual | 27.51 | 0.801 | 0.195 | 0.031 |
| Without distance input | 27.65 | 0.809 | 0.201 | 0.031 |
Table 2 isolates appearance fitting from geometry, visibility, and rasterization. The first two rows approximately match learnable parameter budgets with 328 points/sampled pixels; the third is denser real-image appearance fitting, not full 3D-scene novel-view reconstruction.
| Appearance-fitting setup | Points | SH / NRF parameters | SH / NRF PSNR |
|---|---|---|---|
| Synthetic degree-5 SH target | 328 | 15,744 / 15,707 | 21.1 / 29.0 |
| Real BTF pixels | 328 | 15,744 / 15,707 | 27.8 / 28.8 |
| Real BTF image | 160k | 7.68M / 5.60M | 31.3 / 31.9 |
The text describes the real-image experiment as 200ร200, whereas Table 2 lists 160k points. The cache does not explain the counting conventions; both are retained without inferring the point-count origin. The table reports overall SH/NRF appearance-fitting budgets, not just shared MLP parameters, and its third row must not be described as a strictly equal-memory experiment.
Key Findings¶
- Frequency separation contributes the largest average PSNR gain: removing it costs 0.48 dB, versus 0.35 and 0.21 dB for removing opacity and distance. A larger single branch still performs worse, supporting structural bias over simply increasing capacity.
- Appendix Table 4 reports Bicycle single-branch train/test PSNR of 27.34/23.71 versus 26.11/25.06 for the full model; Stump gives 31.06/25.06 versus 30.10/26.39. Better training but worse testing is the central evidence that the split mitigates overfitting.
- PSNR advantages do not hold for every scene or metric. In the appendix, Bicycle NRF-GS scores 25.06 versus 25.17 for 3DGS; full MipNeRF360 LPIPS in the main table is 0.191 versus 0.184 for 3DGS, also worse.
- Source discrepancies require preservation: the ablation prose repeatedly points to Table 2 although the relevant results are in Table 1; the text summarizes PSNR gains as approximately 0.3โ0.8 dB, while the DL3DV main-table gain over 3DGS is 1.34 dB. Tables 3 and 5 also differ slightly, including subset NRF-GS PSNR of 28.27/28.28, MS-SSIM of 0.815/0.814, and full MipNeRF360 VDGS L1 of 0.029/0.030. This note consistently uses Table 3 for main results rather than silently merging values.
Highlights & Insights¶
- Appearance capacity affects geometric complexity. Different point counts under identical densification rules suggest that Gaussian count reflects not just geometry but also whether individual primitives can explain view changes; the evidence supports an optimization-mechanism interpretation rather than a universal causal theorem.
- Residual sharing is more constrained than sharing complete appearance. Base color retains a stable local anchor, letting the scene-level function specialize in directional changes without having to represent all diffuse content.
- Different frequency permissions for color and opacity are reusable. High-frequency prediction modifies color without directly perturbing transmission, suggesting that different attributes in explicit rendering should not automatically use the same high-capacity predictor.
Limitations & Future Work¶
- The authors acknowledge the cost of per-Gaussian neural evaluation: training and average FPS remain worse than analytic 3DGS. The method still depends on accurate camera poses and initialization and does not jointly optimize runtime or complete memory consumption.
- Synthetic fitting uses degree-5 SH targets, demonstrating degree-3 SH capacity limitations but not replacing complex real-material evaluation. DL3DV uses only 9 randomly selected scenes, so conclusions cannot be generalized to the entire 10k-scene dataset.
- Frequency splitting, gating, and the dual-branch architecture change together, preventing separate attribution of every component. Final opacity-range handling, warmup details, and budget conventions also need implementation clarification.
- Future studies could compare appearance capacity and point allocation under fixed total storage or FPS, then combine pruning, quantization, and faster neural evaluation. These are research suggestions, not results already validated by the paper.
Related Work & Insights¶
- vs 3DGS: Independent degree-3 SH appearance becomes latent features plus a shared residual field, while explicit geometry and rasterization remain. Benefits concern average PSNR and point count; costs include neural evaluation and some LPIPS degradation.
- vs VDGS: Both learn view-dependent color and opacity, but NRF-GS emphasizes appearance latent features, a residual anchor, and frequency decomposition rather than directly modeling full Gaussian parameters. Reported speed exceeds VDGS, but this is not a matched-architecture, single-factor comparison.
- vs GSNB / Latent-SpecGS: GSNB adds neural bases, whereas Latent-SpecGS decodes diffuse and specular components in image space; NRF-GS directly generates appearance residuals at the Gaussian level. Latent-SpecGS is absent from the benchmark, so superiority over it is not established.
Rating¶
- Novelty: 4/5. Shared appearance is not new, but residual anchoring, frequency-specific permissions, and the connection to point count provide clear value.
- Experimental Thoroughness: 4/5. Controlled appearance fitting, three datasets, and ablations are included, although budget conventions and some numerical records need clarification.
- Writing Quality: 4/5. The mechanism is generally clear, with incorrect table references and main-text/appendix numerical differences.
- Value: 4/5. A composable appearance-modeling direction for compact Gaussian representations, not cost-free acceleration or uniformly better perceptual quality.