Skip to content

Towards Practical Lossless Neural Compression for LiDAR Point Clouds

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/pengpeng-yu/FastPCC
Area: Autonomous Driving
Keywords: LiDAR point cloud compression, lossless neural compression, geometry re-densification, cross-scale feature propagation, integer-only inference

TL;DR

To tackle high-resolution contextual sparsity and entropy coding collapse caused by cross-platform floating-point discrepancies in high-precision LiDAR point cloud compression, this paper introduces a lightweight octree predictive coding framework integrating Geometry Re-Densification (GRED) and Cross-Scale Feature Propagation (XFP), paired with an end-to-end integer-only inference pipeline that achieves bit-exact cross-platform decoding and 14 FPS real-time compression.

Background & Motivation

With the rapid proliferation of autonomous driving systems and high-definition 3D mapping, massive volumes of onboard LiDAR point cloud data are collected daily, demanding efficient, practical, and real-time lossless compression solutions. Existing neural point cloud compression (PCC) frameworks typically structure raw coordinates into voxel grids or octrees and predict the conditional probability distributions of geometric occupancy states via autoregressive neural networks. However, these methods encounter a severe physical limitation known as High-Resolution Contextual Sparsity (HRCS) at high reconstruction precisions: as the quantization depth deepens, the local 3D neighborhood of each voxel or octree node becomes extremely sparse. Statistical profiling across KITTI reveals that at deeper octree levels, the average number of occupied neighbors within a \(3\times3\times3\) window drops sharply below one, depriving neural entropy models of meaningful local spatial context and stalling rate-distortion gains.

An even more critical barrier hindering commercial deployment is cross-platform bitstream decodability. Neural entropy coding (e.g., arithmetic coding or asymmetric numeral systems) requires perfect numerical alignment between the encoder and decoder; even an infinitesimal divergence in probability estimates will catastrophically derail the decoding pipeline, turning reconstructed geometry into unstructured uniform spatial noise. Because floating-point operations across heterogeneous accelerators (different GPU architectures, edge NPUs, or CPUs) and software runtimes are inherently non-deterministic, neural codecs have long struggled to operate reliably outside controlled homogeneous environments. Furthermore, existing integerization techniques developed for 2D images cannot handle the sparse convolutions and dynamic spatial execution flows ubiquitous in 3D LiDAR processing.

This paper's angle of attack is to decouple geometric sparsity from feature representation density by mapping shallow dense contexts into deep sparse domains, while enforcing strict deterministic computation across the entire model pipeline. Core idea: develop a Geometry Re-Densification (GRED) module that traces back to dense scales and re-sparsifies features via occupancy pruning, unify it with a thresholded Cross-Scale Feature Propagation (XFP) scheme to eliminate cross-scale redundancy, and engineer the first mixed-precision integer-only inference pipeline covering sparse convolutions and fixed-point lookup-table Softmax to achieve high compression ratios, 14 FPS real-time throughput, and bit-exact cross-platform decoding.

Method

Overall Architecture

The proposed framework performs progressive, dense-to-sparse predictive lossless coding on an \(L\)-level octree \(\mathbf{X} = \{\mathbf{X}^1, \dots, \mathbf{X}^L\}\). At each target level \(l\), the codec predicts the conditional probability distribution of the 8-bit occupancy codes \(\mathbf{X}^l \in \{1, \dots, 255\}^{N_l}\) conditioned on previously encoded levels. The architecture comprises four core stages: octree construction, prior construction, cross-scale feature propagation (incorporating GRED), and integer-only entropy coding. When processing level \(l\), the system evaluates the depth against a density threshold \(t\): shallow levels execute direct low-cost feature propagation, whereas deep levels trigger reverse re-densification to capture dense spatial priors. The resulting sparse feature map feeds a fixed-point MLP predictor to output a 255-way probability vector, driving an integer arithmetic coder to emit the final bitstream.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Raw LiDAR Point Cloud"] --> B["Octree Construction & Level Partition"]
    B --> C{"Check Current Level<br/>l โ‰ค t ?"}
    C -->|Yes: Shallow & Dense| D["Shallow-Level Propagation<br/>Direct feature transfer and concatenation"]
    C -->|No: Deep & Sparse| E["Deep-Level Propagation & GRED<br/>Dense-scale downsampling and re-sparsification"]
    D --> F["Integer-Only Inference Pipeline<br/>INT8 sparse conv & fixed-point math"]
    E --> F
    F --> G["Fixed-Point LUT Softmax<br/>Deterministic probability output"]
    G --> H["Cross-Platform Bit-Exact Entropy Coding"]

Key Designs

1. Geometry Re-Densification Module: tracing back to dense scales to overcome high-resolution contextual sparsity When encoding high octree levels where points are extremely scattered, attempting to expand receptive fields directly in the sparse domain causes prohibitive computational overhead and minimal contextual gain. GRED reverses the spatial flow: instead of operating solely within the sparse level \(l\), it downsamples the previous level's occupancy code \(\mathbf{X}^{l-1}\) into a pre-defined shallow dense octree level \(k\) (\(k < l\)) using sparse convolutions to form a densified feature map \(\mathbf{G}^k\): $$ \mathbf{G}^k = \mathrm{Downsampling}(\mathbf{X}^{l-1}) $$ In this dense domain, a lightweight ResBlock extracts rich spatial geometric features \(\mathbf{F}^k = \mathrm{ResBlock}(\mathbf{G}^k)\). To avoid the combinatorial cost of predicting sub-nodes directly within dense grids, GRED progressively re-sparsifies the features back to level \(l\) through recursive upsampling and occupancy-guided pruning: $$ \mathbf{F}^{m+1} = \mathrm{Pruning}(\mathrm{Upsampling}(\mathbf{F}^m), \mathbf{X}^m), \quad m = k, \dots, l-1 $$ where upsampling expands channels by \(8\times\) followed by PReLU, and pruning strictly eliminates features corresponding to empty child nodes based on true occupancy codes \(\mathbf{X}^m\). This allows the model to inject dense geometric semantics into high-resolution sparse nodes at negligible computational cost.

2. Cross-Scale Feature Propagation Module: dual-regime routing and inter-scale context reuse While GRED mitigates intra-level sparsity, running independent re-densification at every level overlooks strong structural redundancies across scales. XFP generalizes GRED into a hierarchical cross-scale network regulated by an inflection threshold \(t\). At shallow levels (\(l \le t\)), where native geometry remains relatively dense, XFP bypasses compute-heavy re-densification and directly extracts features \(\mathbf{S}^{l-1} = \mathrm{ResBlock}(\mathbf{F}^{l-1})\) from level \(l-1\), projecting them into level \(l\) through a single-step concatenation and pruning. At deep levels (\(l > t\)), where sparsity impairs direct propagation, XFP concatenates the densified feature \(\mathbf{G}^k\) with the existing multi-scale historical representation \(\mathbf{F}^k\), yielding a fused representation: $$ \mathbf{H}^k = \mathrm{ResBlock}(\mathrm{Concat}(\mathbf{F}^k, \mathbf{G}^k)) $$ Recursive upsampling and pruning then proceed from \(\mathbf{H}^k\). This dual-regime mechanism prevents redundant feature extraction at shallow depths while maximizing contextual awareness at fine resolutions.

3. Integer-Only Inference Pipeline: mixed-precision determinism to prevent entropy coding collapse To eliminate catastrophic decoding failure across heterogeneous hardware platforms, the entire inference path is re-architected into pure integer arithmetic. The pipeline employs mixed precision: computation-dominant operations (linear layers and sparse convolutions) run in 8-bit integer (INT8) arithmetic, while re-quantization and non-linear activations utilize 32-bit fixed-point (INT32) arithmetic. Weights are symmetrically quantized with zero-point \(z_w = 0\). The INT32 accumulated output is converted back to INT8 using a precomputed integer multiplier \(m\) and a right-shift factor \(r\): $$ \mathbf{q}y = \mathrm{clip}\left(\left\lfloor \mathbf{y}\right) $$ Sparse convolutions are systematically decomposed into deterministic coordinate mappings combined with indexed linear operations to remove library-level non-deterministic execution paths. Furthermore, to compute the 255-way occupancy probabilities without floating-point division or exponential functions, a 16-bit fixed-point exponential lookup table (LUT) defined over }} \cdot m 2^{-r} \right\rceil + z_y, \; q_{\min}, \; q_{\max\([-12, 0]\) with a step of \(1/512\) replaces standard Softmax: $$ p_i = \frac{\exp(\ell_i - \max_k \ell_k)}{\sum_j \exp(\ell_j - \max_k \ell_k)} $$ All accumulations and normalizations execute in INT32, producing bit-exact probability distributions across all compliant CPUs and GPUs.

Loss & Training

The network is optimized using multi-class cross-entropy loss over the 255 possible 8-bit occupancy states. Given the true conditional distribution \(P(\mathbf{X})\) and predicted distribution \(Q(\mathbf{X})\), the training objective minimizes negative log-likelihood: $$ \mathcal{L} = \mathbb{E}{P(\mathbf{X})} \left[ -\sum) \right] $$ Post-training quantization parameters are calibrated on a compact subset of training frames. All fixed-point and LUT operators operate without requiring additional retraining, facilitating rapid deployment.}^L \log Q(\mathbf{X}^l \mid \mathbf{X}^{1:l-1

Key Experimental Results

Main Results

Evaluations are conducted on the standard autonomous driving benchmarks KITTI (20,351 test frames) and Ford (3,000 test frames) across 11-bit to 16-bit quantization settings. Baselines include the MPEG standard G-PCC (octree mode), Transformer-based OctAttention and Light EHEM, voxel-based Unicorn, and real-time neural codec RENO.

Model Architecture Type KITTI BD-Rate (D1/D2) % KITTI BD-PSNR (D1/D2) dB Ford BD-Rate (D1/D2) % Ford BD-PSNR (D1/D2) dB
OctAttention (AAAI'22) Octree Transformer -7.294 / -7.332 +0.754 / +0.760 -6.205 / -6.193 +0.969 / +0.969
Light EHEM (CVPR'23) Octree Transformer +1.429 / +1.422 -0.152 / -0.152 +11.535 / +11.987 -1.899 / -1.960
Unicorn (TPAMI'25) Voxel Sparse Conv -10.862 / -10.889 +1.367 / +1.373 +3.922 / +3.989 -0.622 / -0.637
RENO (CVPR'25) Real-time Neural -15.582 / -15.579 +1.887 / +1.890 -11.049 / -11.040 +1.938 / +1.938
G-PCC octree (MPEG) Traditional Codec -21.931 / -21.954 +2.532 / +2.541 -19.150 / -19.143 +3.411 / +3.412
Ours (Float) Octree GRED+XFP Anchor Anchor Anchor Anchor
Ours (Integer) vs RENO Integer Pipeline -14.289 / -14.285 +1.728 / +1.730 -6.875 / -6.864 +1.185 / +1.185
Ours (Integer) vs G-PCC Integer Pipeline -20.705 / -20.728 +2.378 / +2.387 -15.355 / -15.348 +2.694 / +2.694

Cross-platform inference profiling verifies exceptional real-time speeds and bit-exact decoding consistency across hardware devices:

Codec System Device Platform KITTI Enc Time (s) KITTI Dec Time (s) Ford Enc Time (s) Ford Dec Time (s) Cross-Platform Consistency
G-PCC octree AMD EPYC 7R32 (CPU) 0.149 0.103 0.150 0.107 Bit-Exact (Hand-Crafted)
RENO NVIDIA RTX 4090 0.059 0.056 0.072 0.057 Non-Deterministic (Fails)
Ours (Float) NVIDIA RTX 4090 0.075 0.081 0.093 0.103 Non-Deterministic (Fails)
Ours (Integer) NVIDIA RTX 4090 0.060 0.070 0.067 0.082 Bit-Exact (Passed)
Ours (Integer) NVIDIA RTX 5880 0.047 0.056 0.054 0.067 Bit-Exact (Passed)
Ours (Integer) NVIDIA Tesla V100 0.166 0.175 0.187 0.199 Bit-Exact (Passed)

Ablation Study

Ablation experiments on KITTI isolate the individual rate-distortion contributions of GRED and XFP relative to the baseline and G-PCC:

Configuration vs Baseline BD-Rate (D1/D2) % vs Baseline BD-PSNR (D1/D2) dB vs G-PCC BD-Rate (D1/D2) % vs G-PCC BD-PSNR (D1/D2) dB Note
Baseline 0.00 / 0.00 0.00 / 0.00 -10.04 / -10.08 +1.13 / +1.14 No re-densification or XFP
+ GRED -3.78 / -3.78 +0.40 / +0.40 -13.44 / -13.47 +1.49 / +1.50 Incorporates dense-scale GRED
+ GRED + XFP (Full) -13.22 / -13.21 +1.46 / +1.46 -21.93 / -21.95 +2.53 / +2.54 Full dual-regime XFP framework

Key Findings

  • XFP provides the largest performance boost: Adding XFP onto GRED yields an additional 1.06 dB BD-PSNR gain and lowers BD-Rate by 9.44%, demonstrating that cross-scale contextual propagation is vital for eliminating spatial blind spots.
  • GRED prevents high-resolution quality decay: Under high-precision quantization (14โ€“16 bits), the baseline without GRED shows significant performance drop-off due to HRCS. GRED mitigates this drop and provides a 0.40 dB net gain.
  • Integerization achieves acceleration with minimal distortion penalty: The integer model incurs only a minor 0.16 dB degradation compared to the floating-point version on KITTI, while reducing RTX 4090 encoding time from 0.075s to 0.060s and enabling 14 FPS full-cycle execution at 12-bit precision.

Highlights & Insights

  • Dense-to-sparse contextual reconstruction: Rather than using computationally prohibitive wide-window attention on ultra-sparse voxels, GRED combines dense downsampled convolutions with occupancy pruning, restoring geometric context efficiently.
  • First bit-exact neural point cloud codec: By integerizing sparse convolutions and implementing a fixed-point LUT Softmax, this work eliminates cross-platform entropy coding divergence without transmitting side calibration data.
  • Thresholded inter-scale routing: By splitting propagation at density inflection threshold \(t\), the framework avoids wasting compute on re-densifying already-dense shallow levels while maintaining rich context in deep levels.

Limitations & Future Work

  • Limitations acknowledged by authors: On datasets with small training volumes like Ford (only 1,500 training frames), the model exhibits slightly lower compression gains compared to heavy Transformer models like Light EHEM.
  • Identified limitations: The integer pipeline has been validated across major server GPUs (RTX 4090, 5880, V100), but deployment on ultra-low-power automotive DSPs or heterogeneous edge NPUs requires further testing.
  • Future directions: Integrating temporal motion compensation across sequential LiDAR frames with integer-only cross-frame predictive coding.
  • vs OctAttention [AAAI'22] & Light EHEM [CVPR'23]: Heavy Transformer models incur substantial decoding latency and lack deterministic cross-platform execution; this paper matches their compression performance while running at 14 FPS with bit-exact guarantees.
  • vs RENO [CVPR'25]: While RENO focuses on speed via fast sampling, its RD performance lags at high resolutions; the proposed integer pipeline achieves +1.73 dB BD-PSNR gains over RENO on KITTI at comparable latency.
  • vs G-PCC octree [MPEG]: Traditional handcrafted codecs are deterministic but yield poor compression ratios; our integer model achieves over 20.7% bitstream savings over G-PCC while retaining bit-exact consistency.

Rating

  • Novelty: โญโญโญโญ [Addresses HRCS via geometry re-densification and introduces the first integer-only neural PCC pipeline]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive evaluations on KITTI/Ford, complete BD metrics, hardware latency, and multi-GPU consistency tests]
  • Writing Quality: โญโญโญโญโญ [Clear problem formulation, rigorous mathematical definitions, and well-structured empirical validation]
  • Value: โญโญโญโญโญ [Directly resolves the critical engineering bottleneck blocking neural point cloud codecs from automotive production deployment]