mmIR: Frequency-Space Inverse Rendering for 3D Millimeter-Wave Radar ADC Synthesis¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Project Page: https://mmwave-inverse-rendering.github.io/
Area: Autonomous Driving
Keywords: mmWave Radar, Inverse Rendering, FMCW ADC Synthesis, Differentiable Rendering, 3D Occupancy
TL;DR¶
Addressing the severe scarcity of high-resolution 3D mmWave radar data and the ill-posed nature of raw ADC inversion, mmIR presents the first open-source, end-to-end differentiable FMCW radar inverse renderer that leverages LiDAR-derived meshes as geometric scaffolds to jointly optimize per-vertex ITU physics materials, vertex normals, and antenna beam patterns, enabling high-fidelity cross-sensor transfer and high-resolution 3D RAE occupancy synthesis via dense virtual apertures.
Background & Motivation¶
Millimeter-wave (mmWave) radar operates reliably through adverse weather conditions like dense fog, rain, and airborne dust, directly measures Doppler velocities, and produces intense specular returns from metallic surfaces, making it an indispensable sensing modality for autonomous systems. However, radar remains critically underutilized for 3D scene understanding. Commercial off-the-shelf single-chip automotive radars feature compact antenna arrays with only a dozen virtual elements, restricting angular resolution to several degrees and constraining outputs to 2D range–azimuth (RA) slices rather than full 3D range–azimuth–elevation (RAE) volumes. Compounding this hardware limitation, public autonomous driving benchmarks almost exclusively release heavily thresholded, sparse point clouds or post-processed 2D intensity maps rather than raw time-domain intermediate-frequency analog-to-digital converter (ADC) samples. Consequently, learning-based 3D object detection, occupancy prediction, and radar super-resolution are starved of high-dimensional, raw volumetric supervision.
Existing approaches to break this impasse face severe technical and physical barriers. Scaling physical antenna arrays drastically inflates packaging cost and RF complexity without expanding environmental diversity. Synthetic aperture radar (SAR) demands continuous, precise mechanical trajectory sweeps that cannot scale across commercial vehicle fleets. Emerging neural volumetric representations and Gaussian splatting methods are fundamentally bottlenecked by the low-resolution 2D observations they train on, unable to recover unobserved 3D elevation structure; moreover, directly fitting post-processed images forces neural models to memorize Fourier transform point spread function (PSF) sidelobes and windowing artifacts rather than true electromagnetic scattering. Conversely, classical physics-based shooting-and-bouncing ray (SBR) simulators are strictly forward-only and cannot calibrate material parameters against real measurements. Differentiable RF frameworks such as Sionna-RT detach path geometry and run solvers under gradient suspension, preventing gradients from flowing back to vertex positions or surface normals, while omitting wideband FMCW signal physics entirely.
Recognizing that recovering geometry from coarse, single-frame radar observations is mathematically ill-posed, this paper introduces a pragmatic paradigm shift: since modern autonomous platforms routinely capture high-density LiDAR data, mmIR employs a LiDAR-derived surface mesh as a rigid geometric scaffold, dedicating inverse rendering entirely to uncovering electromagnetic material properties and antenna characteristics. Core Idea: An end-to-end differentiable wideband FMCW radar inverse renderer that automatically differentiates through multi-bounce ray tracing, polarization-aware ITU BSDFs, and MIMO phasor accumulation to fit per-vertex material parameters, surface normals, and antenna beam patterns to real raw ADC captures, subsequently re-rendering frozen scenes through massive virtual apertures to synthesize dense, physics-consistent 3D radar ADC volumes.
Method¶
Overall Architecture¶
mmIR frames radar sensing as an end-to-end differentiable forward rendering and parameter inversion pipeline. The system takes as input a geometric triangle mesh reconstructed from multi-sweep LiDAR point clouds via screened Poisson surface reconstruction, along with real FMCW ADC signals recorded by a physical cascaded radar. A differentiable ray generator emits multi-bounce paths using a combination of Rx-centric reservoir sampling, specular manifold sampling, and free-space edge diffraction. At each surface interaction, a physically-based mmWave BSDF grounded in ITU-R P.2040 evaluates complex dielectric reflections and Jones matrix polarization transport. A topology-sharing mechanism coherently accumulates phase across all MIMO channels to synthesize complex-valued FMCW ADC data, from which 2D range–azimuth maps are computed via 2D spatial-temporal FFTs to compute a min–max normalized L1 loss against real measurements. Once optimized, all scene materials and normals are frozen, and a massive virtual aperture (e.g., \(100\times 100\) elements) re-renders the scene to generate high-resolution 3D RAE radar data and point clouds.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Data<br/>LiDAR Poisson Mesh + Real Cascaded Radar ADC"] --> B["Multi-Strategy Ray Generation<br/>Rx Reservoir Sampling + Specular Manifold + Free-Space Diffraction"]
B --> C["Physical mmWave BSDF<br/>ITU-R Materials + Rayleigh Coherent/Incoherent + Polarization"]
C --> D["Phase-Coherent MIMO ADC Synthesis<br/>Shared-Topology Complex Baseband Chirp Accumulation"]
D --> E["Differentiable Range-Azimuth Loss<br/>2D FFT + Min-Max Normalized L1"]
E -->|End-to-End DrJit AD Back-Propagation| F["Joint Parameter Optimization<br/>Per-Vertex ITU Dielectrics + Normals + Beam Patterns"]
F --> G["Dense Virtual Aperture Re-Rendering<br/>100×100 Virtual Array for Volumetric 3D RAE Occupancy"]
Key Designs¶
1. Physically-Based mmWave BSDF: Decoupling Coherent Specular, Roughness Scattering, and Polarization
Conventional radar simulation oversimplifies target scattering into scalar radar cross-sections (RCS) or simplistic diffuse approximations, completely failing to capture the complex angular and dielectric dependencies of real materials. mmIR formulates a per-vertex BSDF grounded in ITU-R P.2040 recommendations, equipping each mesh vertex with six physically meaningful parameters: complex relative permittivity \((\varepsilon'_r, \varepsilon''_r)\), RMS surface roughness \(\sigma_h\), correlation length \(l_c\), Kirchhoff-approximation / small-perturbation-method (KA/SPM) blend factor \(\tau\), and slab physical thickness \(d\). Triangle-interior properties are obtained via barycentric interpolation, guaranteeing \(C^0\) spatial continuity and gradient distribution across all three vertices. The BSDF decomposes into coherent and incoherent lobes modulated by the Rayleigh roughness factor:
where \(\theta_t\) denotes the local incidence angle. Modulated by an ITU-R slab Fresnel energy gate \(A(\boldsymbol{\omega}_t)\) accounting for internal reflections, the coherent component blends a GGX microfacet specular model with a specularly directed von Mises–Fisher (vMF) distribution. The incoherent component combines a wider vMF directional lobe with an isotropic Lambertian term. Throughout ray tracing, local \(s/p\) polarization bases are explicitly maintained, updating field amplitudes and phase via complex Fresnel coefficients at each reflection to guarantee wave polarization fidelity.
2. Hybrid Importance Ray Generation: Resolving Sub-Wavelength Specular Points and Boundary Diffraction
At 77 GHz, the radar wavelength (\(\lambda \approx 3.9\text{ mm}\)) is substantially smaller than typical triangle faces in LiDAR meshes. Standard image methods almost always fail because the specular reflection point falls outside the discrete boundary of the intersected triangle, while uniform spherical sampling incurs catastrophic Monte Carlo variance. mmIR establishes a three-part ray sampling architecture: First, Rx-centric reservoir sampling generates candidate rays from a cosine-weighted proposal centered along the receiver boresight, concentrating samples within high-gain antenna regions. Second, for coherent reflections, Specular Manifold Sampling (SMS) solves the half-vector constraint \(C(\mathbf{x}) = [\mathbf{s}\cdot\mathbf{h}, \mathbf{t}\cdot\mathbf{h}]^\top = \mathbf{0}\) through GPU-vectorized Newton iterations; by re-projecting rays onto the mesh between steps, the solver walks across triangle boundaries to pinpoint exact specular paths, with gradients propagated through the Implicit Function Theorem (IFT) Jacobian inverse. Third, to bypass expensive explicit edge detection required by Uniform Theory of Diffraction (UTD), mmIR integrates a Free-Space Diffraction (FSD) BSDF that projects local occluding boundaries onto virtual aperture screens, computes closed-form Fraunhofer integrals, and modulates diffracted paths with dielectric-aware Fresnel transmission and Jones polarization matrices.
3. Topology-Shared Phase-Coherent MIMO Signal Synthesis: Preserving Phase Consistency Across Channels
The spatial angular resolution of a MIMO radar relies upon precise sub-millimeter relative phase delays across disparate transmit-receive antenna baselines. Evaluating independent Monte Carlo path topologies for each Tx–Rx channel injects stochastic phase noise that destroys spatial beamforming coherence during azimuth Fourier transformation. mmIR introduces a path topology reuse strategy: shared random seeds generate a common multi-bounce path skeleton \(\mathbf{x}_1, \dots, \mathbf{x}_D\), after which only the initial transmitter \(\text{Tx}_i \to \mathbf{x}_1\) and final receiver \(\mathbf{x}_D \to \text{Rx}_j\) segments are vectorized across channels. This slashes computational complexity from \(\mathcal{O}(N_{\text{Tx}} N_{\text{Rx}} N)\) to \(\mathcal{O}(N)\) while enforcing deterministic multi-channel phase coherence. For each path with total travel distance \(R_{ij}^{(p)}\), intermediate-frequency complex baseband FMCW chirps are accumulated:
where \(f_c\) is carrier frequency, \(S\) is chirp slope, and \(t_k\) indexes fast-time ADC samples. A 2D FFT converts synthesized multi-channel ADC arrays into 2D range–azimuth maps. The objective function enforces an L1 loss over min–max normalized magnitudes:
Supervising in the RA magnitude domain implicitly penalizes relative phase errors across virtual elements (which would smear angular peaks) while remaining immune to unstable absolute oscillator phase drifts and hardware DC offsets.
Loss & Training¶
Optimization is executed entirely within a single DrJit JIT-compiled computational graph. Learnable variables comprise \(6 \times N_v\) per-vertex dielectric properties, per-vertex surface normals \(\mathbf{N} \in \mathbb{R}^{N_v \times 3}\), and directional antenna beam gains. Parameters are initialized with standard ITU-R reference values for common building materials (concrete, glass, metal). On a single NVIDIA RTX 4090 GPU, fitting converges in 500 iterations within approximately 11 minutes (1.32 s/iter), delivering a substantial speedup over Sionna-RT (3.28 s/iter). Once trained, all scene parameters are frozen for forward-only virtual aperture synthesis.
Key Experimental Results¶
Main Results¶
On seven outdoor scenes and six indoor scenes (laboratories, hallways, classrooms) from the ColoRadar benchmark, mmIR is evaluated against the differentiable RF baseline Sionna-RT using a physical 12Tx×16Rx cascaded radar (136 virtual elements, 86 uniform azimuth elements). Comprehensive quantitative results are summarized below:
| Evaluation Split | Model | RA Correlation ↑ | RA PSNR (dB) ↑ | RA SSIM ↑ | RA RMSE ↓ | ADC Log-Mag MSE ↓ |
|---|---|---|---|---|---|---|
| Outdoor Mean (7 scenes) | Sionna-RT mmIR (Ours) |
0.324 ± 0.076 0.919 ± 0.030 |
22.3 ± 3.0 37.6 ± 2.4 |
0.506 ± 0.115 0.902 ± 0.049 |
0.0808 ± 0.0229 0.0137 ± 0.0038 |
268.31 ± 20.37 28.49 ± 11.97 |
| Indoor Mean (6 scenes) | Sionna-RT mmIR (Ours) |
0.288 ± 0.107 0.909 ± 0.070 |
24.5 ± 3.7 42.9 ± 5.2 |
0.680 ± 0.143 0.918 ± 0.037 |
0.0657 ± 0.0311 0.0084 ± 0.0040 |
278.04 ± 15.56 24.24 ± 10.47 |
| Overall Mean (13 scenes) | Sionna-RT mmIR (Ours) |
0.307 ± 0.093 0.914 ± 0.053 |
23.3 ± 3.5 40.0 ± 4.8 |
0.586 ± 0.155 0.910 ± 0.044 |
0.0738 ± 0.0280 0.0112 ± 0.0047 |
272.80 ± 18.94 26.53 ± 11.50 |
Cross-Sensor Zero-Shot Transfer¶
To rigorously verify that mmIR recovers true physical scene properties rather than memorizing sensor-specific artifacts, scene parameters optimized on the cascaded radar are frozen and evaluated on a co-located single-chip radar (TI AWR1843, 3Tx×4Rx, 12 virtual elements, 8 azimuth elements) without re-training:
| Scene ID | Model | Transfer RA Corr ↑ | Transfer RA PSNR ↑ | Transfer RA SSIM ↑ | Transfer RA RMSE ↓ | Transfer ADC Log MSE ↓ |
|---|---|---|---|---|---|---|
| S0-F135 | Sionna-RT Ours |
0.002 0.533 |
17.8 25.6 |
0.464 0.505 |
0.1283 0.0524 |
301.43 10.96 |
| S1-F438 | Sionna-RT Ours |
0.079 0.684 |
13.5 23.2 |
0.202 0.674 |
0.2119 0.0694 |
348.02 21.87 |
| S2-F300 | Sionna-RT Ours |
0.008 0.745 |
23.6 27.2 |
0.412 0.468 |
0.0661 0.0434 |
287.94 7.72 |
| Outdoor Mean | Sionna-RT Ours |
0.114 ± 0.103 0.554 ± 0.143 |
20.0 ± 3.6 23.4 ± 2.5 |
0.468 ± 0.138 0.529 ± 0.080 |
0.1087 ± 0.0488 0.0707 ± 0.0206 |
317.35 ± 22.37 11.10 ± 7.39 |
3D Occupancy Inference & Ablations¶
Using frozen scenes, mmIR re-renders a massive virtual array of 10,000 elements (\(100\times 100\)), generating true 3D RAE tensors (\(256\times 127\times 127\)). Thresholding at the 99.9th percentile yields radar point clouds benchmarked against single-frame LiDAR points (\(\tau=0.5\text{ m}\)): - 3D Occupancy Gains: mmIR achieves an average Precision of 0.618 (vs. 0.319 for the cascaded radar baseline with heuristic elevation extrusion) and an F1-score of 0.352 (vs. 0.129 baseline), reducing Relative Chamfer Distance (R-CD) from 0.208 to 0.141. On structured scenes (e.g., S1-F185), Precision reaches 0.997 with an F1-score of 0.722. - Ablation Studies: Disabling antenna beam pattern optimization leads to a 3.7% drop in RA correlation; omitting free-space diffraction degrades correlation by 1.0%; and reverting from per-vertex properties to coarse per-object material assignment introduces severe misalignments across complex facades.
Key Findings¶
- End-to-End Differentiation Drives Multipath Convergence: Sionna-RT freezes ray paths, preventing gradient flow to normals and scattering geometry; mmIR unifies ray tracing, BSDF, and signal accumulation in a single graph, allowing normal optimization to correct LiDAR meshing defects.
- Transferable Physical Invariance: The 0.554 zero-shot transfer correlation and a \(28\times\) reduction in ADC log MSE on the single-chip radar confirm that mmIR optimizes intrinsic dielectric properties rather than sensor-coupled overfitting.
- Specular Scattering Geometry Governs Recall: In open scenes where ground planes lie nearly parallel to the radar line of sight, specular energy is redirected away from the receiver array, causing moderate recall degradation in 3D occupancy evaluation.
Highlights & Insights¶
- Bridging Differentiable Graphics and Wave-Domain RF: mmIR unifies modern Monte Carlo sampling (SMS, FSD, ReSTIR) with FMCW chirp phase physics inside DrJit, overcoming severe phase gradient instability at millimeter wavelengths.
- Resolution Lifting via Dense Virtual Apertures: By optimizing physics parameters under coarse sensor supervision and re-rendering via massive virtual apertures, mmIR establishes an effective workflow to generate high-resolution 3D volumetric radar data without expensive hardware.
- A Data Engine for Downstream Radar Perception: Frozen digital twins enable on-demand synthesis of realistic raw ADC signals across arbitrary array geometries, directly fueling downstream radar 3D occupancy and object detection models.
Limitations & Future Work¶
- Dependence on LiDAR Geometric Prior: mmIR requires an initial Poisson mesh scaffold; in areas where LiDAR returns are absent or degraded, radar inverse rendering cannot recover unknown geometry from single-view radar alone.
- Absence of Dynamic and Ego-Motion Doppler Modeling: The current model assumes static environments and omits Doppler phase modulation (\(4\pi v_{\text{ego}} t / \lambda\)), requiring extensions for dynamic highway scenarios.
- Phase Sensitivity Inhibits Direct Vertex Displacement: At 77 GHz, sub-millimeter geometric perturbations trigger complete phase inversions; optimizing vertex positions directly remains unstable without dense multi-view temporal regularization.
Related Work & Insights¶
- vs Sionna-RT: Sionna-RT detaches path geometry during ray tracing, lacks support for wideband FMCW ADC synthesis, and cannot optimize surface normals; mmIR achieves fully differentiable ADC rendering with multi-lobe ITU BSDFs.
- vs DART & RadarSplat: Neural volumetric and Gaussian splatting baselines fit post-processed 2D images directly, discarding raw phase and binding representations to specific sensors; mmIR recovers sensor-agnostic physical dielectrics capable of cross-hardware transfer.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneering end-to-end differentiable FMCW inverse renderer with virtual aperture 3D synthesis]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across 13 diverse indoor/outdoor scenes, zero-shot hardware transfer, and 3D occupancy benchmarking]
- Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous mathematical derivations linking electromagnetic wave theory with differentiable path tracing]
- Value: ⭐⭐⭐⭐⭐ [Provides a crucial open-source foundation for automotive radar data synthesis and physical digital twins]