Skip to content

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation

Conference: ECCV 2026
Paper: ECCV 2026
Project: Project Page
Area: 3D Vision
Keywords: implicit neural representation, sinusoidal recurrence, harmonic line spectrum, bipolar code space, Gray coding

TL;DR

The paper reveals that sinusoidal activations mathematically induce harmonic line-spectrum enrichment, proposing a weight-tied bias-free sinusoidal recurrence architecture with bipolar Gray-code alignment supervision that expands effective spectral support under a fixed parameter budget, achieving accelerated high-fidelity and exact discrete reconstruction across 2D and 3D tasks.

Background & Motivation

Implicit neural representations (INRs) parameterize continuous signals as coordinate-based neural networks, offering continuous sampling, resolution-independent evaluation, and analytical differentiability that have made them central to 3D reconstruction, novel view synthesis (NeRF), and continuous image representation. However, coordinate MLPs are fundamentally bottlenecked by spectral bias: networks tend to learn low-frequency structures early while struggling to capture fine high-frequency details. Existing solutions predominantly focus on activation function engineering (e.g., periodic activations in SIREN, variable-periodic activations in FINER, or complex Gabor wavelets in WIRE), multi-frequency input coordinate encodings (e.g., Random Fourier Features or multiresolution hash grids), or increasing model capacity by stacking deeper feed-forward hidden transformation layers.

Nonetheless, conventional deep feed-forward designs depend on stacking independently parameterized layers, which directly drives up model size, memory consumption, and training complexity. While equilibrium-style formulations (such as iSIREN) employ implicit fixed-point iterations to achieve constant-memory backpropagation, they are primarily motivated by amortized fixed-point solving; their latent representations remain stationary at equilibrium, leaving open the spectral mechanism by which repeated transformations alter feature representations. This exposes a core tension: can an INR's latent transformation layer continuously enhance its capacity to resolve fine-scale details under a strictly constrained parameter budget, while explicitly accounting for how repeated transformations enrich the feature spectrum?

This paper's angle of attack connects sinusoidal non-linearities with recurrent unrolling to formalize the mechanism of harmonic spectral enrichment. Through rigorous derivation using the Jacobi-Anger expansion, the authors demonstrate that passing through a shared sinusoidal block generates integer linear combinations of existing frequencies, thereby expanding effective spectral support across unrolled steps without adding independently parameterized depth. Core idea: employ weight-tied, bias-free sinusoidal recurrence to iteratively enrich harmonic spectral support, coupled with bipolar Gray-code alignment supervision, achieving high-fidelity and exact discrete signal reconstruction with fewer parameters and optimization steps.

Method

Overall Architecture

The overall architecture decomposes coordinate mapping into three stages: sinusoidal coordinate embedding, recurrent latent refinement, and code-space output alignment. Given spatial coordinates \(\mathbf{x} \in \mathbb{R}^d\), a learnable sinusoidal projection maps them to an initial latent feature \(\mathbf{h}^{(0)}\); this representation is subsequently passed to a weight-tied, bias-free sinusoidal recurrent block unrolled for \(R\) steps to generate higher-order harmonic interactions; finally, a linear output projection bounded by a smooth saturation function predicts a continuous codeword vector that is supervised against bipolar Gray targets via cosine similarity alignment. At test time, an exact discrete signal is recovered via sign thresholding and inverse Gray decoding.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Spatial Coordinate x"] --> B["Sinusoidal Coordinate Embedding<br/>learnable Fourier feature mapping"]
    B --> C["Bias-Free Sinusoidal Recurrence<br/>weight-tied unrolling for R steps + Jacobi-Anger harmonic expansion"]
    C --> D["Linear Output Projection & Smooth Saturation<br/>continuous code prediction c_hat"]
    D --> E["Bipolar Gray-Code Alignment Supervision<br/>Gray codes to {-1, +1}^B + constant-norm cosine loss"]
    E --> F["Distortion-Aware Bit-Plane Reweighting<br/>prioritize high-significance planes based on decoded distortion"]
    F --> G["Exact Discrete Signal Reconstruction & Rendering<br/>image fitting / super-resolution / NeRF / SDF"]

Key Designs

1. Bias-free sinusoidal recurrence: expanding harmonic line spectra while suppressing phase-shift drift To address parameter redundancy in deep feed-forward networks and cumulative phase distortion in recurrent architectures, this design reuses a single shared weight matrix \(W^{(\mathrm{rec})}\) across \(R\) unrolled steps: \(\mathbf{h}^{(r)} = \sin(\omega W^{(\mathrm{rec})} \mathbf{h}^{(r-1)})\). From a spectral viewpoint, the input sinusoidal embedding layer converts coordinates into a finite set of basis tones \([\mathbf{h}_0(\mathbf{x})]_i = \sin(\mathbf{\Omega}_i^\top \mathbf{x} + \phi_i)\). When the hidden pre-activation is a linear combination of these basis tones \(u_j(\mathbf{x}) = \sum_{i=1}^m \alpha_{j,i} \sin(\theta_i(\mathbf{x}))\), applying Euler's formula and the multi-tone Jacobi-Anger identity yields an exact Fourier expansion: $\(\sin(u_j(\mathbf{x})) = \sum_{\mathbf{k} \in \mathbb{Z}^m} \Big(\prod_{i=1}^m J_{k_i}(\alpha_{j,i})\Big) \sin\Big(\sum_{i=1}^m k_i \theta_i(\mathbf{x})\Big)\)$ This analytical result shows that the sinusoidal activation does not merely reweight existing tones, but explicitly synthesizes new spectral lines at integer linear combinations \(\mathbf{\Omega} = \sum_{i=1}^m k_i \mathbf{\Omega}_i\) of encoder frequencies. Consequently, each recurrent unrolling step induces a progressive spectral support growth \(\mathbf{\Omega}^{(r+1)} \subseteq \mathrm{span}_{\mathbb{Z}}(\mathbf{\Omega}^{(r)})\). Crucially, the recurrent block strictly omits the bias term \(b_{\mathrm{rec}}\). If a bias were present, the sine addition formula shows that an identical phase shift \(\boldsymbol{\beta} = \omega b_{\mathrm{rec}}\) is applied at every step, altering the Jacobian derivative gate to \(\operatorname{diag}(\cos(\mathbf{z}^{(r)} + \boldsymbol{\beta})) \omega W_{\mathrm{rec}}\). Over multiple unrolled iterations (\(R \ge 2\)), this static shift accumulates nonlinearly along the recurrent trajectory, perturbing derivative gating and severely destabilizing high-frequency gradient dynamics. Eliminating the bias ensures gate symmetry and training stability.

2. Bipolar Gray-code alignment supervision: regularizing discrete reconstruction via constant-norm cosine alignment Conventional continuous real-valued regression (e.g., MSE loss) frequently suffers from high-frequency ringing and edge blurring when reconstructing quantized signals, while remaining overly sensitive to output scale. This design maps quantized target values \(y_p \in \{0, \dots, 2^B-1\}\) to \(B\)-bit Gray codes \(\Gamma(y_p)\), which are transformed into zero-centered bipolar targets \(\mathbf{c}_p = 2\Gamma(y_p) - 1 \in \{-1, +1\}^B\). Because every bipolar codeword has an identical \(L_2\) norm \(\|\mathbf{c}_p\|_2^2 = B\), expanding the squared Euclidean distance reveals: $\(\|\hat{\mathbf{c}}_p - \mathbf{c}_p\|_2^2 = \|\hat{\mathbf{c}}_p\|_2^2 + B - 2\hat{\mathbf{c}}_p^\top \mathbf{c}_p\)$ This establishes that the target-dependent error is determined entirely by the angular alignment between \(\hat{\mathbf{c}}_p\) and \(\mathbf{c}_p\). Optimizing via cosine similarity: $\(\mathcal{L}_{\mathrm{align}} = \frac{1}{|\mathcal{P}|}\sum_{p \in \mathcal{P}} \frac{\hat{\mathbf{c}}_p^\top \mathbf{c}_p}{\|\hat{\mathbf{c}}_p\|_2 \|\mathbf{c}_p\|_2}\)$ directly penalizes code orientation errors while eliminating scale sensitivity. Furthermore, whereas natural binary coding (NBC) can flip up to \(n\) bits between adjacent quantization steps (yielding a steep local code-space sensitivity of \(L_{\mathrm{NBC}}^{\mathrm{local}} = n/\delta\)), Gray coding guarantees that adjacent levels differ by exactly 1 bit (\(L_{\mathrm{GC}}^{\mathrm{local}} = 1/\delta\)). This \(n\)-fold reduction in local sensitivity provides code-space continuity that dramatically smooths the optimization landscape.

3. Distortion-aware bit-plane reweighting: redistributing optimization budget according to decoded impact In capacity-limited regimes or early optimization stages where zero bit error cannot yet be attained, bit errors across different code planes exhibit vastly different impacts on decoded signal distortion. Errors in high-significance planes cause severe intensity shifts and prominent block artifacts, whereas errors in low-significance planes only introduce subtle noise. To optimize perceptual and quantitative fidelity when errors are inevitable, the loss reweights code planes according to decoded severity: $\(\mathcal{L}_{\mathrm{w}} = \sum_{i=0}^{B-1} w_i \mathcal{L}^{(i)}, \quad w_i = \begin{cases} 2^{-(8-i)}, & i \in \{0, 1, 2\} \\ 1, & \text{otherwise} \end{cases}\)$ where \(i=0\) corresponds to the least significant plane. Downweighting the three lowest planes directs the network's optimization capacity toward critical high-order planes, suppressing structural artifacts and yielding significant PSNR gains under constrained budgets.

Loss & Training

The network output is bounded by the smooth saturating activation \(G(z) = \sin(\arctan z) = z / \sqrt{1 + z^2}\), which maps into \([-1, 1]\) while remaining approximately linear near the origin. The training pipeline is fully end-to-end differentiable, optimized with AdamW at a learning rate of \(1.5 \times 10^{-4}\) without a learning rate scheduler. The thresholding operator \(T(\cdot) = \operatorname{sign}(\cdot)\) and inverse Gray decoding \(B^{-1}\) are evaluated only at test time.

Key Experimental Results

Main Results

On Set5, Kodak24, DIV2K-100, and FFHQ-600, the proposed model was evaluated against standard coordinate MLPs, band-limited networks, and learnable-frequency spectral INRs. Baselines were trained for 1,000 iterations with 791K parameters. In contrast, the proposed method with 609K parameters at only 100 iterations substantially outperforms all 1,000-iteration spectral baselines. When optimized with adaptive iterations (\(\le 1000\)), it consistently achieves zero bit error and exact quantized reconstruction (PSNR = \(\infty\)).

Method #Params #Iter Set5 (PSNR / SSIM / Bit err) Kodak24 (PSNR / SSIM / Bit err) DIV2K-100 (PSNR / SSIM / Bit err) FFHQ-600 (PSNR / SSIM / Bit err)
WIRE 791K 1000 26.91 / .6637 / 544.52K 26.83 / .6583 / 566.38K 26.65 / .7620 / 555.53K 26.43 / .5971 / 561.66K
Gauss 791K 1000 33.97 / .8802 / 453.77K 34.18 / .8867 / 460.84K 34.06 / .9211 / 456.13K 34.01 / .8647 / 458.33K
SIREN 791K 1000 38.54 / .9629 / 350.11K 35.01 / .9404 / 395.26K 32.81 / .9535 / 430.17K 37.68 / .9630 / 351.22K
FINER 791K 1000 41.78 / .9775 / 309.20K 42.15 / .9803 / 305.60K 40.81 / .9876 / 325.74K 41.08 / .9796 / 312.68K
BACON 791K 1000 33.71 / .9488 / 390.47K 31.93 / .9194 / 429.31K 29.73 / .9177 / 462.32K 33.15 / .9414 / 390.39K
FourierNet 791K 1000 45.67 / .9924 / 242.83K 41.77 / .9882 / 293.41K 39.31 / .9883 / 334.63K 42.64 / .9898 / 266.08K
GaborNet 791K 1000 50.31 / .9970 / 190.86K 51.38 / .9975 / 180.71K 49.51 / .9976 / 206.54K 49.54 / .9962 / 193.25K
iSIREN 791K 1000 51.80 / .9976 / 172.55K 52.67 / .9981 / 167.44K 51.43 / .9983 / 183.92K 50.90 / .9971 / 173.83K
Ours (#iter=100) 609K 100 58.16 / .9997 / 0.78K 53.20 / .9989 / 2.51K 46.28 / .9973 / 5.66K 61.39 / .9998 / 0.62K
Ours (adaptive) 609K \(\le\)1000 \(\infty\) (322 iters) \(\infty\) (447 iters) \(\infty\) (537 iters) \(\infty\) (440 iters)

Ablation Study

The ablation studies isolate the impact of unrolling steps \(R\), the presence of the recurrent bias term, and downstream novel view synthesis performance. Table 2A reports peak PSNR on Kodak24 across unrolling steps \(R \in \{1, \dots, 5\}\) comparing recurrent bias enabled (ON) versus disabled (OFF) under 792K parameters. Table 2B evaluates transferring the recurrent sinusoidal decoder into NeRF for novel view synthesis on LLFF scenes under a matched 600K parameter budget.

Table 2A: Effect of recurrent unrolling steps and recurrent bias on Kodak24 (PSNR dB)

Recurrent Bias Setting \(R=1\) (Feed-forward) \(R=2\) \(R=3\) \(R=4\) \(R=5\)
Bias ON 51.16 60.35 51.97 63.57 78.10
Bias OFF (Ours) 51.41 83.49 87.54 89.59 90.28
Gain (\(\Delta\)PSNR) +0.25 +23.14 +35.57 +26.02 +12.18

Table 2B: Novel view synthesis performance on LLFF dataset (600K parameter budget)

Scene Baseline NeRF (PSNR ↑ / SSIM ↑ / LPIPS ↓) Ours Decoder (PSNR ↑ / SSIM ↑ / LPIPS ↓) Gain (\(\Delta\)PSNR / \(\Delta\)LPIPS)
Flower 23.71 / 0.6316 / 0.4145 24.54 / 0.7138 / 0.2934 +0.83 dB / -0.1211
Orchids 17.85 / 0.4198 / 0.5172 18.79 / 0.5313 / 0.3797 +0.94 dB / -0.1375
Fern 21.06 / 0.5640 / 0.4943 21.60 / 0.6500 / 0.3704 +0.54 dB / -0.1239
Fortress 26.10 / 0.5987 / 0.4423 26.36 / 0.7189 / 0.3206 +0.26 dB / -0.1217

Key Findings

  • Crucial role of bias-free formulation: When unrolled for \(R \ge 2\) steps, retaining recurrent bias causes PSNR to degrade by 12 to 35 dB, corroborating that static phase shifts compound along the unrolled chain and distort derivative gating.
  • Sinusoid-specific spectral expansion: Under recurrent connections, non-sinusoidal baselines (Gaussian activations or positional encoding MLPs) experience complete spectral collapse with upper-band support dropping to 0 (Table 3c), leading to severe performance drops (Gaussian drops from 60.31 dB to 12.02 dB). In contrast, sinusoidal recurrence uniquely expands spectral support.
  • Disentangled contributions: According to Table 12, recurrence alone accounts for the primary jump in continuous fidelity (improving from 42.15 dB to 64.19 dB over feed-forward FINER), while bipolar Gray-code supervision eliminates residual quantization mismatch to reach exact reconstruction (\(\infty\) dB).

Highlights & Insights

  • Formalized sinusoidal non-linearities as harmonic line-spectrum generators: Utilizing the Jacobi-Anger expansion connects network depth and recurrence to structured Fourier harmonic expansion, providing an intuitive mathematical grounding for spectral bias mitigation.
  • Pinpointed phase-shift drift in weight-tied recurrence: Identified that recurrent bias accumulates static phase shifts that destabilize backpropagation gradients, demonstrating that a bias-free recurrent layer unlocks massive empirical gains.
  • Bridged discrete signal reconstruction with continuous optimization manifolds: Combined the single-bit adjacency of Gray coding with the constant-norm properties of bipolar space, delivering an optimization objective that avoids scale sensitivity while achieving lossless discrete reconstruction.

Limitations & Future Work

  • Author-acknowledged limitations: While the architecture yields substantial gains under moderate-to-large capacity regimes (\(\ge 400\text{K}\)), its advantages attenuate under severe capacity constraints (\(\le 100\text{K}\)), where base representation capacity limits the expressive power of weight sharing.
  • Computational trade-offs: Although training wall-clock time drops dramatically due to fast convergence, per-coordinate forward inference computation scales linearly with the recurrent unrolling depth \(R\).
  • Future directions: Developing spatially adaptive recurrent unrolling (dynamically allocating recurrent steps based on local geometric or frequency complexity), and extending the architecture to 3D Gaussian Splatting attribute modeling and neural audio compression.
  • vs SIREN / FINER: Conventional sinusoidal networks rely on stacking independent parameter layers to build depth, suffering from parameter bloat; this work shows that a single shared sinusoidal block achieves superior harmonic expansion with lower parameter overhead.
  • vs iSIREN (DEQ): iSIREN targets steady-state fixed-point representations with constant-memory backpropagation, but its spectrum remains stationary at equilibrium; this work leverages explicit finite unrolling to exploit dynamic spectral enrichment.
  • vs BACON / FourierNet: BACON uses analytical band-limiting and FourierNet employs explicit frequency filters; this approach demonstrates that a weight-tied sinusoidal block autonomously achieves rich harmonic line-spectrum generation.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Rigorous harmonic analysis using Jacobi-Anger expansion provides an elegant spectral foundation for recurrent INRs.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation spanning 2D image fitting, super-resolution, NeRF, and SDF reconstruction with detailed spectral support analyses.
  • Writing Quality: ⭐⭐⭐⭐⭐ Flawless mathematical formalization, clear logical progression, and well-integrated qualitative and quantitative evidence.
  • Value: ⭐⭐⭐⭐⭐ Delivers an efficient, theoretically grounded architectural component that overcomes spectral bias and parameter explosion in implicit representations.