M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals¶
Conference: NeurIPS 2026
arXiv: 2609.28684
Paper: Project page
Code: https://github.com/dsilvavinicius/m-plicits
Area: 3D Vision
Keywords: neural implicit surfaces, multiscale residuals, nested narrow bands, sphere tracing, attribute mapping
TL;DR¶
M-plicits reconstructs oriented point clouds with a full-domain coarse SIREN and progressively band-supervised residual SIRENs, reusing this hierarchy for tracing, mesh extraction, and attribute mapping to deliver real-time rendering and strong synthetic measurement-noise robustness with compact models, without winning every accuracy or speed metric.
Background & Motivation¶
Neural implicit surfaces map a three-dimensional position to a field value and represent the surface as its zero level set. When that field approximates a signed distance function (SDF), it supports normals, distance queries, and sphere tracing. SIREN's sinusoidal activations capture detail, but a single large MLP must be evaluated in full at every query: ray steps far from the surface and dense grid sampling still pay for the detail network. iNGP's hash grid improves speed but can fit input noise, while BACON's explicit frequency control supplies multiscale representations with its own spectral-truncation artifacts and capacity limitations.
Splitting a network into coarse and fine components does not by itself answer where the fine network should operate. Full-domain residual supervision still requires learning and querying regions far from the surface. Conversely, independently training a fine field only near the surface can leave additional zero level sets in unsupervised regions. This paper therefore co-designs the representation, supervision domains, and downstream evaluation: establish a full-domain distance prior, refine only near the preceding surface, and query expensive levels only within those neighborhoods at inference time.
Denoising here arises from the smoothing bias of a small coarse network and staged local training, not a proven frequency-cutoff filter. Core Idea: preserve coarse-to-fine geometric coupling through residual composition, concentrate fine-level supervision in nested narrow bands, and reuse those bands to reduce expensive rendering and extraction queries.
Method¶
Overall Architecture¶
The input is a point cloud with normals; texture tasks also require colors. The main outputs are coarse, medium, and fine neural distance fields and their surfaces. A full-domain base field first fits large-scale shape, and nested residual refinement progressively adds detail. Multiscale geometry evaluation handles sphere tracing and mesh extraction, while neural attribute mapping queries detailed normals or colors on existing geometry.
Training and inference are distinct: points, normals, and optional colors provide supervision, rather than being rendering inputs. Rendering uses camera rays, and mesh extraction uses grid vertices; both query the trained hierarchy. Dashed edges below indicate supervision, while solid edges indicate stage dependencies and inference queries.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Oriented point cloud<br/>Normals / optional colors"] -.->|Point and normal supervision| B["Full-domain base field"]
B --> C["Nested residual refinement"]
A -.->|Band-sampling supervision| C
C --> D["Multiscale geometry evaluation"]
R["Camera rays / grid vertices"] -->|Inference queries| D
D --> E["Neural attribute mapping"]
C -->|Fine-field gradients| E
A -.->|Optional color supervision| E
D -->|Sphere tracing / mesh extraction| O["Geometry / rendered output"]
E -->|Normals / colors| O
Key Designs¶
1. Full-domain base field: avoid querying the detail network far from the surface
The base field is a small SIREN covering the entire training domain. Input point positions, normal directions, and domain-wide Eikonal regularization supervise an initial approximation of overall geometry and distance. This is not a setup with a precomputed ground-truth SDF labeling every position: the main task starts from oriented point clouds and relies on regularization to constrain the field beyond the discrete data.
SIREN's sinusoidal frequency parameter influences its tendency and capacity to express high-frequency detail. The coarse, medium, and fine levels use 30, 45, and 100, respectively, together with different network sizes to implement progressive refinement. A smaller frequency parameter is not a strict band-limit proof for a composition of sinusoidal MLP layers, nor does it establish that the coarse network can never absorb noise regardless of training duration. The experiments support a smoothing prior and denoising behavior under the evaluated configurations.
2. Nested residual refinement: correct the preceding field rather than learn an independent local distance field
Each refinement forms a new distance field by adding a residual, rather than predicting an independent field disconnected from the coarse level. The default uses three SIRENs: one base network and two residual networks. Evaluating a fine field includes all preceding levels, not just the last residual.
After training a level, the next level receives supervision only in the intersection of the preceding training domain and its near-zero region. The band width is the maximum fitting error at input points multiplied by a safety margin. This retains the input points and leaves room to correct portions not yet accurately explained by the preceding field.
The maximum is over input surface points, not a full-domain field norm. Successive intersections guarantee nested supervision domains. However, each residual remains an ordinary continuous MLP without architecturally guaranteed compact support, and its values outside the supervised domain need not be exactly zero. The paper describes global zero-level-set nesting as guaranteed by construction, but local supervision and a data-error-based band width alone do not prove that stronger statement. Training results and band-culled inference provide empirical support, which should be distinguished from a full-space mathematical guarantee.
For Eikonal sampling, the authors dither each coordinate around input points and reject samples outside the valid band, concentrating evaluations near the surface rather than uniformly throughout the volume. Another sampling path offsets input points along their normals: within a region where the local distance relation remains valid, the offset supplies the SDF target and the input normal supplies the gradient-direction target. The authors choose a valid range by progressively shortening the offset and comparing it with the distance to the point cloud. This approximation depends on reliable normals and a usable local tubular neighborhood; arbitrary offsets do not have exact labels.
3. Multiscale geometry evaluation: approach through inexpensive offset surfaces before querying fine levels
Ordinary sphere tracing queries the target SDF at every step. This method first traces a positive offset level set of the coarse field, then a positive offset level set of the medium field, and finally the fine zero level set. Small networks therefore handle most early advancement, while high-capacity networks are queried near the surface. Subtracting the band width from coarse-level step sizes avoids terminating merely upon reaching the coarse surface.
This equation describes coarse and medium offset tracing; the final fine level uses ordinary steps without the offset subtraction. The experiments fix 20 iterations for the coarse level and 5 for each subsequent level to facilitate GPU batching. A band that is too narrow can miss the fine silhouette; an overly wide band can require more fine-level iterations than the budget allows, leaving holes. Training regularization does not provide a global Lipschitz certificate for the neural field or turn this fixed budget into a rigorous convergence guarantee. The paper's convergence wording should be read as an algorithmic description contingent on distance approximation and hierarchy assumptions.
Mesh extraction reuses the same structure. It first evaluates the coarse field at grid vertices, retains coarse values outside the band, adds residuals inside the band, and applies marching cubes after hierarchical evaluation. Culling skips expensive fine-level queries; it does not delete all outside-band vertices or assign them zero values. The supplementary ablation uses full-domain extraction for full-domain-supervised residuals, avoiding unverified band culling for a baseline trained under a different regime.
4. Neural attribute mapping: query detailed appearance on inexpensive geometry
The renderer can trace the coarse surface and directly query fine-field gradients at its intersections for shading normals. This avoids additional tracing of the fine surface while supplying finer apparent relief. Geometry remains at the coarse intersection positions, however, so normal mapping alone does not change the silhouette. The silhouette improvement in Figure 9 comes from additional medium-level sphere tracing, not from normals moving the surface.
Normals are computed by analytically propagating coordinate derivatives through the chain rule. The implementation batches derivatives for the three coordinate directions separately, alternating GEMMs with elementwise activation-derivative kernels; residual fields additionally sum the constituent network gradients. This bypasses inference-time autograd graphs and is not finite differencing. Smoothness comes from the continuously differentiable representation, while GEMM provides the efficient computation path. The speedup in supplementary Table S16 varies with model and resolution and should not universally be described as exactly twofold.
Colors are encoded by a network mapping three-dimensional coordinates to RGB. It fits surface colors and penalizes variation along the SDF gradient within a neighborhood, extending color information from a fine surface toward nearby coarse geometry without UV parameterization. This does not enforce a globally constant three-dimensional color field; it encourages consistency along normal directions.
This notation explicitly distinguishes prediction from ground-truth color, which appear with indistinguishable glyphs in the extracted source: the barred term denotes ground truth, and the second term is interpreted through the directional derivatives of the RGB channels. This clarifies the source's attribute-loss notation rather than introducing a new training term.
A Worked Example¶
Consider the Lucy point cloud with normals. The full-domain base field first produces smooth geometry lacking fine detail. Its fitting error at the input points determines a band, and dithered rejection samples plus normal-offset samples supervise the medium and fine residuals. The clean model's measured band widths are 2.67E-02 and 1.43E-02 in the unit-sphere domain, according to supplementary Table S15.
Rendering offers two distinct choices. For lower cost, trace only coarse geometry and query fine normals for improved shading. To improve the silhouette as well, continue tracing the medium and fine fields. The first choice improves appearance without replacing geometry; the second actually changes ray intersections. Mesh export instead queries fine levels inside the bands and uses coarse values outside them, rather than evaluating every fine network everywhere.
Loss & Training¶
Each level combines a zero-value constraint at input points, normal alignment, and Eikonal regularization over its current training domain. The following preserves the mechanism of main-paper Equation (4). The discrete implementation requires sampling and loss weights; the continuous integral should not be mistaken for an exactly evaluated quantity.
Residual stages fit the composed field rather than forcing the residual itself to vanish at input points. Normal-offset samples add nonzero distance supervision, while the Eikonal term constrains local distance behavior. When colors are available, the attribute field is trained separately; color errors are not treated as geometric supervision.
The default architecture notation is (128, 2) โท (256, 2) โท (400, 2), with two hidden layers of the corresponding width at each level. Adam training proceeds through the base field, medium residual, and fine residual in order. The band-width measurements in supplementary Table S15 use a safety margin of 0.3. Architectures and measured widths from different experiments should not be collapsed into a single fixed configuration.
Key Experimental Results¶
Main Results¶
The main evaluation uses Stanford and Thingi32, with centered point clouds normalized to the unit sphere and a single NVIDIA RTX 5090 with 32 GB. Mesh extraction uses a \(512^3\) grid, followed by 500K uniformly sampled mesh points for L2 Chamfer Distance (CD); IoU measures occupancy agreement. Rendering uses \(512^2\) images. These timings are not general claims across devices or resolutions.
Table 1: Clean-input results from main-paper Table 2. Lower CD is better; higher IoU/FPS is better. Sampling denotes grid-field evaluation.
| Method | mean CD | median CD | mean IoU | median IoU | Parameters | Training min | Sampling s | FPS |
|---|---|---|---|---|---|---|---|---|
| iNGP coarse | 6.87E-05 | 6.43E-05 | 4.12E-01 | 4.05E-01 | 2,040,864 | 8.22 | 0.40 | 86 |
| iNGP fine | 6.44E-05 | 6.04E-05 | 4.20E-01 | 3.90E-01 | 9,113,760 | 10.59 | 0.89 | 88 |
| IDF | 6.13E-02 | 7.74E-04 | 4.09E-01 | 9.20E-02 | 1,191,943 | 8.26 | 42.38 | N/A |
| BACON | 8.12E-03 | 2.18E-05 | 5.15E-01 | 7.32E-01 | 530,953 | 8.43 | 0.24 | N/A |
| Ours coarse | 5.36E-05 | 3.80E-05 | 4.48E-01 | 4.59E-01 | 17,153 | 4.07 | 1.51 | 180 |
| Ours fine | 6.57E-05 | 1.87E-05 | 5.86E-01 | 8.67E-01 | 246,627 | 8.53 | 5.29 | 43 |
The coarse model has the best mean CD. The fine model has the best median CD and both IoU aggregates, but its mean CD is slightly worse than iNGP fine. It is not faster than every baseline either: BACON samples fastest, iNGP fine has higher FPS, and Ours fine trains longer than IDF and iNGP coarse. Compact parameters, noise robustness, and continuous normals form the combined advantage; this table does not establish across-the-board dominance.
Table 2: Synthetic measurement-noise results from main-paper Table 3; random displacement along ground-truth normals is at most 1.0% of bounding-box scale.
| Method | mean CD | median CD | mean IoU | median IoU |
|---|---|---|---|---|
| Ours coarse | 7.42E-05 | 5.33E-05 | 3.81E-01 | 4.03E-01 |
| Ours medium | 4.62E-05 | 3.03E-05 | 5.37E-01 | 6.38E-01 |
| Ours fine | 3.36E-05 | 1.62E-05 | 5.82E-01 | 5.91E-01 |
| iNGP coarse | 6.70E-04 | 3.76E-04 | 1.90E-01 | 1.88E-01 |
| iNGP fine | 2.80E-04 | 1.09E-04 | 3.09E-01 | 3.05E-01 |
| IDF | 8.37E-02 | 2.49E-05 | 3.22E-01 | 3.08E-01 |
| BACON | 1.68E-03 | 1.51E-04 | 1.76E-01 | 1.65E-01 |
The fine level leads all four aggregates here, and the progressive CD improvement supports refinement after denoising. However, this setup retains ground-truth normals and is not equivalent to real sensor data, corrupted normals, or unoriented point clouds. Lower CD in the noisy table than in the clean table also does not establish that adding noise necessarily improves reconstruction: these are different inputs and fitted results, not a controlled causal conclusion from cross-table changes.
Ablation Study¶
Table 3: Supplementary Table S12 on a 10-shape subset. Variants share architectures, frequency schedules, and the coarse network; these results are not directly comparable with full-dataset Tables 1 and 2.
| Setting | Variant | mean CD | median CD | mean Hausdorff | mean IoU |
|---|---|---|---|---|---|
| 1% noise | (a) Residuals + bands | 6.39E-05 | 1.49E-05 | 1.45E-03 | 0.687 |
| 1% noise | (b) Residuals + full-domain supervision | 6.34E-05 | 2.26E-05 | 5.23E-03 | 0.555 |
| 1% noise | (c) Independent fields + bands | 1.04E-04 | 2.36E-05 | 5.72E-03 | 0.669 |
| clean | (a) Residuals + bands | 3.87E-05 | 1.38E-05 | 5.86E-03 | 0.653 |
| clean | (b) Residuals + full-domain supervision | 4.26E-05 | 2.36E-05 | 6.75E-03 | 0.655 |
| clean | (c) Independent fields + bands | 3.47E-05 | 1.62E-05 | 4.01E-03 | 0.749 |
Under noise, (b) has slightly better mean CD than (a), but substantially worse Hausdorff distance. Bands primarily improve the distant-error tail and occupancy structure, rather than every average point-error metric. On clean inputs, (c) has better mean CD, Hausdorff distance, and IoU. Residual composition is therefore not an unconditional accuracy improvement; its main justification is noise stability and behavior outside the bands.
Removing extraction culling for independent band-trained fields changes noisy-checkpoint mean CD to 2.8E-01 and produces 365 to 4,109 components. Removing culling for the residual field changes mean CD only from 6.39E-05 to 6.42E-05. This empirically contrasts behavior in unsupervised regions; it does not prove that residual fields can never have zero level sets outside the bands.
Key Findings¶
- Main-paper Table 4 reports 35 FPS for full geometry tracing of the three-level model and 43 FPS for the normal-mapping variant that omits final-level tracing; a single large SIREN runs at 19 FPS. Table 2 also labels the fine model's 43 FPS as geometry evaluation without attribute mapping. These descriptions do not fully align, so the original values are retained rather than silently corrected.
- Main-paper Table 5 reduces Armadillo extraction from a 4.87 s baseline to 0.99 s with culling, and Lucy from 4.87 s to 1.07 s. These gains concern a specific extraction configuration and do not establish faster sampling than BACON in the main comparison.
- The largest clean band width in supplementary Table S15 is 2.93E-02, below 3% of the unit-sphere radius. The noisy Thai widths are 6.9E-02 and 5.2E-02, however, so the claim that all bands are below 3% cannot be extended to noisy inputs.
- Supplementary Table S14 reports IoU 0.002 for this method on Lucy with 1% outlier contamination, versus 0.004 for untrimmed Poisson and 0.938 with density trimming. Outliers can inflate a maximum-error-based band; measurement-noise robustness does not imply outlier robustness.
Highlights & Insights¶
- The representation hierarchy also becomes a computation-budget hierarchy. Unlike a network that merely offers multiple output resolutions, the bands supply a spatial condition for deciding when a fine network is needed at both training and inference time.
- Residual coupling and supervision-domain contraction serve different purposes. The former inherits the full-domain base field, while the latter concentrates regularization and sampling near the surface; Hausdorff ablations explain their joint value better than CD alone.
- Detailed normals and detailed geometry can receive separate budgets. Attribute mapping is useful for expressing detail through shading, but silhouette-sensitive tasks still require tracing a finer surface.
Limitations & Future Work¶
- The smoothing bias sacrifices very sharp edges and microgeometry. The authors suggest local features, but the paper does not provide a complete solution that already resolves this limitation.
- Maximum-fitting-error band selection is sensitive to isolated outliers. Quantile-based or reliability-weighted widths are plausible alternatives, but replacing the maximum requires revisiting data coverage and inference safety rather than assuming all guarantees remain unchanged.
- Local supervision does not guarantee global zero-level-set nesting, and Eikonal regularization does not certify a global distance lower bound. More reliable directions include checking outside-band signs, bounding residual amplitudes, or introducing provably safe step sizes.
- The main evaluation assumes oriented point clouds and synthetic normal-direction noise. Non-watertight meshes, thin structures, self-intersections, and isolated components cause difficulties, but disconnected components alone do not mathematically preclude a continuous SDF.
- Supplementary image reconstruction, dynamic-shape, and large-scene demonstrations do not establish general scene robustness. Average DTU PSNR rises from 28.86 dB to 29.94 dB, but Scan 37 and Scan 40 fall from 27.10 and 28.13 to 26.50 and 27.78, respectively; the source's claim of consistent improvement is too strong.
Related Work & Insights¶
- vs SIREN: The method retains continuous sinusoidal MLPs but uses residuals and local supervision to reduce fine-level work. It does not obtain theoretical guarantees through a new strictly band-limited architecture.
- vs iNGP: Continuous networks avoid spatial hash grids and are more stable in the evaluated noise setting. iNGP fine still has advantages in clean mean CD, FPS, and sampling speed.
- vs BACON / IDF: This paper combines residual refinement with shrinking supervision domains and directly supplies multiscale tracing. Baseline artifacts and extreme errors in these experiments do not establish the same behavior for every implementation or input.
- vs BANF: The supplementary comparison uses the M-plicits authors' custom implementation, not an official BANF point-cloud-fitting baseline. They report an optimization gap against BANF's published fine-resolution results, so absolute rankings require caution.
- vs Screened Poisson: Supplementary Table S6 reports clean-subset mean IoU 0.910 for Poisson versus 0.653 for this method. The neural representation's advantages concern distance queries and hierarchical attributes, not an assertion that Poisson meshes cannot be rasterized in real time.
Rating¶
- Novelty: 4/5 โ Co-designing nested supervision and multiscale evaluation goes beyond a residual network alone.
- Experimental Thoroughness: 4/5 โ Noise tests, isolated ablations, and outlier failures are included, but evidence primarily comes from limited shapes and synthetic settings.
- Writing Quality: 3/5 โ The method is clear, while nesting guarantees, band-limit explanations, and some evaluation descriptions require qualification.
- Value: 4/5 โ Useful for compact implicit geometry and real-time shading, subject to limitations on sharp detail and anomalous data.