Skip to content

TriNLOS: Triplane Representations for Neural Non-Line-of-Sight Imaging

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/CIL-Research/TriNLOS
Area: LLM (Other)
Keywords: Non-Line-of-Sight Imaging, Triplane Neural Representation, Light-Cone Transform, Volumetric Rendering, Transient Light Transport

TL;DR

Addressing the cubic computational complexity \(O(N^3)\) and heavy diffuse background noise in dense 3D non-line-of-sight (NLOS) neural inversion, TriNLOS combines a differentiable Enhanced Light-Cone Transform (ELCT) initialization with an orthogonal triplane representation, achieving state-of-the-art volumetric and surface reconstruction in \(O(N^2)\) complexity with only 8.0 GB VRAM.

Background & Motivation

Non-line-of-sight (NLOS) imaging aims to recover the 3D geometry and reflectance of hidden scenes occluded from the direct line of sight by analyzing multiply scattered photons measured on an accessible diffuse relay wall. In confocal setups where pulsed laser illumination and single-photon avalanche diode (SPAD) detectors scan the relay wall coaxially, the system captures time-resolved transient histograms encoding distance through time-of-flight. Classical physics-based analytic inversion methodsโ€”such as the light-cone transform (LCT), phasor field diffraction, and \(f\text{-}k\) migrationโ€”offer training-free closed-form solutions based on wave propagation models. However, they assume idealized Lambertian reflection and uniform planar boundaries, leaving them vulnerable to severe shot noise, timing jitter, and diffuse clutter in real-world measurements.

To overcome model-mismatch artifacts, learning-based approaches have been introduced to either directly map transient measurements to 3D volumes or refine coarse physics-based reconstructions using deep neural networks. Nevertheless, directly learning a mapping from spatio-temporal transient cubes to 3D spatial grids suffers from weak geometric correspondence, while feeding coarse physics volumes into dense 3D CNN backbones introduces prohibitive cubic computational complexity \(O(N^3)\) and immense GPU memory consumption. Crucially, hidden NLOS targets are typically sparse, occupying only a small spatial volume, while the vast majority of voxels consist of diffuse background radiation and sensor noise, causing dense 3D backbones to squander capacity and memory on irrelevant regions.

The key insight is that practical NLOS targets predominantly present dominant-view or front-facing surface geometry, whose spatial constraints can be effectively captured through orthogonal 2D projections. Triplane factorizations reduce volumetric feature processing from \(O(N^3)\) to \(O(N^2)\), while axis-wise maximum projection naturally acts as a soft threshold that filters out low-magnitude background clutter and isolates salient foreground surfaces. Core Idea: combine a differentiable Enhanced Light-Cone Transform (ELCT) prior with an orthogonal triplane representation, filtering out diffuse volumetric clutter via axis-wise max-projection, restoring collapsed geometric cues through shared-axis cross-attention, and stabilizing multi-task intensity and depth learning via a gradient-blocked dual rendering head.

Method

Overall Architecture

The TriNLOS reconstruction framework operates through four sequential stages: first, the measured spatio-temporal transient tensor \(T \in \mathbb{R}^{\tau \times H \times W}\) is preprocessed by a shallow denoising block and transformed into a coarse, physically consistent 3D volume \(\tilde{V}\) via a differentiable Enhanced Light-Cone Transform (ELCT); second, learnable 3D positional embeddings are added before axis-wise maximum projection collapses the volume into three orthogonal 2D feature planes (\(xy\), \(xz\), \(yz\)), which are encoded independently by 2D ResConvNeXt backbones enhanced with multi-scale Restormer channel-attention; third, slice-wise cross-attention is executed between plane pairs along their shared physical axes to recover spatial cues lost during projection, followed by confidence-weighted back-projection and fusion into a unified 3D volume; finally, the fused volume is pooled and decoded by lightweight 2D rendering heads under a gradient-blocked depth branch to jointly predict hidden scene intensity and metric depth.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Transient Measurements T<br/>ฯ„ ร— H ร— W SPAD signals"] --> B["Differentiable Enhanced LCT ELCT<br/>depth remapping + frequency sharpening"]
    B --> C["Axis-Wise Max Triplane Projection and Feature Encoding<br/>ResConvNeXt + Restormer 2D extraction"]
    C --> D["Shared-Axis Cross-Attention and Confidence Volume Fusion<br/>cross-plane feature exchange and 3D lifting"]
    D --> E["Gradient-Blocked Depth Branch and Joint Rendering Heads<br/>block depth gradients to volume, render intensity and depth"]
    E --> F["Output Reconstruction<br/>high-fidelity intensity map and depth map"]

Key Designs

1. Differentiable Enhanced LCT (ELCT): Physics-Guided Coarse Initialization End-to-end deep networks struggle to infer non-local flight-time geometry without explicit physical guidance, whereas standard LCT suffers from non-linear depth distortion and high-frequency scattering noise. ELCT formulates a differentiable physical inversion operator governed by only three learnable parameters \((\alpha, s, p)\). It first applies a learnable power-law depth re-mapping \(t' = (t / \tau)^\alpha \tau\) along the temporal axis to correct non-linear travel-time distortions. Next, a Laplacian-of-Gaussian (LoG) filter sharpens fine structural edges in the frequency domain, while a soft cone mask attenuates off-cone diffuse scattering. A learnable depth gate \(\gamma(z) = \frac{1}{1 + (z/s)^p}\) dynamically balances masked and unmasked components before Wiener deconvolution with the system point spread function (PSF) yields the coarse volume \(V_0 = \text{ELCT}(T)\). Because temporal sampling is substantially denser than spatial dimensions (\(\tau=512 \gg 128\)), max-pooling and a lightweight learnable downsampler are fused via normalized convolution to produce a balanced coarse volume \(\tilde{V}\).

2. Axis-Wise Max Triplane Projection and Feature Encoding: Scalable \(O(N^2)\) Processing To bypass the memory bottlenecks of cubic 3D convolutions, the coarse volume \(\tilde{V}\) is augmented with a learnable 3D positional embedding \(E\) and factorized into three orthogonal 2D planes via axis-wise maximum projection: $\(M_{ab}(u, v) = \max_{w \in c} \left[ \tilde{V}(u, v, w) + E(u, v, w) \right], \quad (ab, c) \in \{(xy, z), (xz, y), (yz, x)\}\)$ The core geometric insight is that valid object surfaces manifest as high-amplitude backscattered peaks in the reconstructed volume, while ambient diffuse returns and sensor shot noise remain low in magnitude. Applying maximum projection along coordinate axes naturally suppresses ambient background clutter and concentrates network capacity onto geometrically salient foregrounds. Each plane is then processed by a 2D encoder featuring ResConvNeXt residual blocks with large receptive fields and multi-scale Restormer channel-attention blocks. Because Restormer computes attention across feature channels rather than spatial tokens, its complexity scales as \(O(HW C^2)\) rather than \(O(N^2)\), achieving robust multi-scale denoising with minimal GPU memory overhead.

3. Shared-Axis Cross-Attention and Confidence Volume Fusion: Restoring Collapsed Geometry While 2D triplane projection dramatically reduces computational complexity, collapsing each spatial axis eliminates depth ordering along that direction, leading to degraded separation in multi-object scenes. To counteract this loss, slice-wise cross-attention is deployed between plane pairs that share a common coordinate axis. For planes \(M_{xy}\) and \(M_{xz}\) sharing axis \(x\), feature slices \(\{F_i\}\) and \(\{G_i\}\) along \(x\) are updated via cross-attention restricted strictly to aligned slices: $\(F_i \leftarrow F_i + \text{softmax}\left(\frac{(F_i W_Q)(G_i W_K)^\top}{\sqrt{d}}\right) (G_i W_V)\)$ This constrains attention computations to physically co-located lines, recovering occluded and depth-displaced boundaries without expensive full 2D or 3D attention. Following cyclical cross-plane updates, features are lifted to 3D space via broadcasting along the collapsed dimension. A lightweight axis-attention network predicts voxel-wise confidence weights \(W_{xy}, W_{xz}, W_{yz}\) to blend the representations into \(V_{\text{tri}} = \text{Linear}(W_{xy} F_{bxy} + W_{xz} F_{bxz} + W_{yz} F_{byz})\), which is concatenated with the physical initial volume \(\tilde{V}\) and refined through shallow 3D convolutions into \(V_{\text{out}}\).

4. Gradient-Blocked Depth Branch and Joint Rendering Heads: Decoupled Multi-Task Optimization Rather than directly regressing a continuous 3D field of albedo and surface normalsโ€”which incurs parameter redundancy and misaligns with downstream 2D inspection tasksโ€”the network aggregates the volume into a compact 2D descriptor \(F(x, y) = \phi([\max_z V_{\text{out}}(x, y, z) \parallel F_{xy}^{2D}(x, y)])\) alongside the peak response depth index \(g(x, y) = \arg\max_z V_{\text{out}}(x, y, z)\). Two shallow 2D convolutional heads then predict the intensity map \(\hat{I}\) and depth map \(\hat{D}\). Crucially, backpropagating depth errors through non-smooth volumetric \(\arg\max\) and \(\max\) operators produces chaotic gradients that compromise feature contrast and severely blur reconstructed intensity. To prevent this destructive competition, gradients from the depth head are blocked at the volumetric interface: $\(\frac{\partial \hat{D}}{\partial (\arg\max_z V_{\text{out}})} = 0, \quad \frac{\partial \hat{D}}{\partial (\max_z V_{\text{out}})} = 0\)$ This unilateral gradient barrier isolates 3D volumetric learning for intensity and structural fidelity, while restricting depth error optimization to the parameters of the 2D depth rendering head.

Loss & Training

The network is supervised end-to-end using a composite \(L_1\) loss over intensity and depth: $\(\mathcal{L} = \| I - \hat{I} \|_1 + \alpha \| D - \hat{D} \|_1\)$ with balancing weight \(\alpha = 0.1\). Optimization is conducted with AdamW (weight decay \(1 \times 10^{-4}\), excluding biases and normalization layers) on a workstation with an Intel Core Ultra 9 285K CPU and a single NVIDIA RTX PRO 6000 GPU (96 GB VRAM) using batch size 4 for 35 epochs. The learning rate is initialized at \(1 \times 10^{-4}\) and decayed to \(1 \times 10^{-6}\) using a cosine annealing schedule. Realistic measurement artifacts are simulated during training through on-the-fly injection of additive Gaussian noise and Poisson shot noise, alongside planar rotations and horizontal flips.

Key Experimental Results

Main Results

Quantitative evaluations are performed on synthetic transient datasets generated with the LFE renderer [6] (evaluating both Seen motorcycle targets and Unseen classes including guns, aircraft, and cars) as well as public real-world SPAD measurements. Comparisons against state-of-the-art physics-based operators and learning-based architectures are summarized in Table 1.

Category Method PSNRโ†‘ (Seen) PSNRโ†‘ (Unseen) SSIMโ†‘ (Seen) SSIMโ†‘ (Unseen) RMSE_Depโ†“ (Seen) RMSE_Depโ†“ (Unseen) MAD_Depโ†“ (Seen) MAD_Depโ†“ (Unseen) Runtime (ms) Memory (GB)
Physics-based LCT [35] 25.34 22.72 0.644 0.437 0.664 0.576 0.691 0.606 14.35 7.5
Physics-based ELCT (Ours init) 26.82 23.63 0.618 0.496 0.680 0.593 0.700 0.614 14.80 8.6
Physics-based Phasor [26] 26.67 25.53 0.507 0.435 0.753 0.721 0.776 0.747 29.91 10.4
Physics-based FK [23] 26.24 24.96 0.874 0.816 0.555 0.539 0.571 0.558 22.63 9.3
Learning-based LFE [6] 27.42 25.43 0.894 0.827 0.066 0.086 0.019 0.034 48.80 12.2
Learning-based NLOST [20] 28.77 25.65 0.946 0.885 0.063 0.106 0.013 0.032 165.29 12.2
Learning-based TriNLOS (Ours) 30.26 26.13 0.958 0.911 0.054 0.087 0.011 0.025 103.22 8.0

Ablation Study

Systematic ablations evaluate the individual contributions of architectural modules, physics-based initialization, depth-gradient propagation, and planar projection components on seen test data.

Configuration PSNR_Int โ†‘ SSIM_Int โ†‘ RMSE_Dep โ†“ MAD_Dep โ†“ Observation & Impact
Full model (Proposed) 30.20 0.958 0.054 0.011 Optimal trade-off across intensity fidelity and metric depth precision
w/o Restormer attention 28.44 0.944 0.063 0.014 Largest performance drop (-1.76 dB PSNR); severe loss in multi-scale denoising
w/o Axis cross-attention 29.21 0.951 0.061 0.014 Decoupled planes fail to resolve depth ambiguities, degrading multi-object separation
w/o ResConvNeXt 29.31 0.952 0.062 0.014 Local high-frequency structural details become fragmented under high sensor noise
w/ depth-gradient propagation 28.47 0.946 0.058 0.012 Gradients from argmax destabilize volumetric contrast, causing severe blur (-1.73 dB PSNR)
Replace ELCT with LCT init 29.20 0.950 0.059 0.015 Uncalibrated depth scale and high diffuse clutter weaken backbone refinement
w/o yz projection 28.98 0.951 0.059 0.013 Eliminating lateral geometry causes largest drop among triplane factorizations

Key Findings

  • Synergy of Restormer and Cross-Axis Attention: Eliminating Restormer causes the steepest quantitative drop (PSNR drops by 1.76 dB), validating its channel-attention capability in restoring low-SNR transient structures. Concurrently, axis cross-attention successfully compensates for the dimensional collapse of triplane representations, restoring critical depth boundaries.
  • Superior Memory-Accuracy Efficiency: TriNLOS consumes only 8.0 GB VRAM, achieving a 34.4% reduction compared to NLOST (12.2 GB) and LFE (12.2 GB), while running 37.5% faster than NLOST (103.22 ms vs. 165.29 ms) and delivering superior PSNR and depth accuracy.
  • Essential Role of Depth Gradient Blocking: Allowing depth gradients to penetrate the shared volumetric feature space causes immediate optimization divergence between continuous depth interpolation and high-contrast intensity estimation, degrading PSNR by 1.73 dB.

Highlights & Insights

  • Physics Prior + Structured 2D Factorization: By using the differentiable ELCT operator to transform non-local temporal measurements into an approximate 3D volume, TriNLOS establishes a reliable geometric coordinate framework. Transitioning to 2D triplane processing then eliminates cubic \(O(N^3)\) computational explosion without sacrificing fine geometry.
  • Intrinsic Noise Suppression via Max-Projection: Selecting maximum values along projection axes inherently discards low-magnitude ambient noise and diffuse scattering clutter pervasive in NLOS time-of-flight measurements, focusing feature extraction directly on high-confidence reflecting surfaces.
  • Extensibility to Non-Confocal Setups: The framework generalizes seamlessly beyond confocal imaging. By substituting ELCT with a Relay-Surface Diffraction (RSD) operator, TriNLOS achieves 28.12 dB PSNR on non-confocal data, outperforming raw RSD (25.33 dB) by nearly 3 dB and demonstrating plug-and-play adaptability.

Limitations & Future Work

  • Severe Multi-Layer Occlusion Breakdown: While triplane representations effectively capture front-facing and convex geometries, scenes exhibiting deep concave pockets or complex multi-layer occlusions can still suffer from projection ambiguities that challenge cross-axis attention.
  • Sequential Triplane Inference Latency: Current processing executes triplane feature extraction sequentially, resulting in a 103.22 ms inference time. While substantially faster than dense 3D CNNs, parallel multi-stream implementations are needed to achieve 30+ FPS real-time imaging.
  • Future Directions: Integrating adaptive sparse voxel grids with triplane structures, and bridging this pipeline with 3D Gaussian Splatting (3DGS) for continuous novel-view synthesis of hidden environments.
  • vs. LCT [35] & Analytic Operators: Analytic physical operators assume ideal reflective surfaces and clean measurements, producing pervasive blur and noise artifacts. TriNLOS introduces learnable parameters into ELCT and pairs it with neural refinement, improving PSNR by approximately 5 dB.
  • vs. LFE [6]: LFE processes 3D volumes via dense volumetric embeddings, incurring 12.2 GB VRAM usage and producing diffuse, smoothed object boundaries. TriNLOS achieves sharper reconstructions with 34.4% lower memory footprint.
  • vs. NLOST [20]: NLOST relies on heavy 3D networks that suffer from 165.29 ms latency and severe boundary spreading artifacts. TriNLOS reduces runtime by 37.5% while preserving delicate peripheral details such as slender animal legs and fine mechanical contours.

Rating

  • Novelty: โญโญโญโญ [Pioneering integration of orthogonal triplane factorizations and differentiable physics-based light-cone operators for efficient NLOS imaging.]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive quantitative validation across synthetic and real-world datasets, seen/unseen categories, confocal/non-confocal setups, and thorough module ablations.]
  • Writing Quality: โญโญโญโญโญ [Exceptionally clear mathematical formulation, well-motivated architecture diagrams, and rigorous analysis of trade-offs.]
  • Value: โญโญโญโญโญ [Resolves the fundamental \(O(N^3)\) complexity barrier in learning-based NLOS reconstruction, paving the way for practical real-time hidden scene perception.]