DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing¶
Conference: ECCV 2026
Paper: ECCV Official
Area: Autonomous Driving
Keywords: Autonomous Driving, 3D Gaussian Splatting, End-to-End Testing, Sensor Simulation, Closed-loop Evaluation
TL;DR¶
Addressing the core trilemma among fidelity, interactivity, and real-time throughput in closed-loop end-to-end autonomous driving simulation, DecoupleGS explicitly decouples driving scenes into static backgrounds and canonical dynamic assets, resolving efficiency, geometric, and photometric conflicts via asset compression, map-guided registration, and proxy-based relighting.
Background & Motivation¶
Autonomous driving (AD) is transitioning from classical modular pipelines toward End-to-End (E2E) paradigms that directly map raw sensor observations to vehicle trajectories or control commands. While E2E algorithms mitigate cascading error accumulation, their safety and policy robustness demand rigorous validation across diverse safety-critical corner cases. Because real-world on-road data collection for hazardous scenarios is prohibitively costly and risky, high-fidelity virtual simulation has become an indispensable data engine. However, existing sensor simulation approaches inherently compromise: game-engine simulators rely on hand-crafted assets and approximate rendering, leaving a non-negligible Sim-to-Real domain gap; 2D generative diffusion models lack strict multi-view 3D spatial-temporal geometric consistency and incur high inference latency incompatible with high-rate control loops; and static neural reconstruction methods (NeRF or 3D Gaussian Splatting) entangle scene geometry and illumination, making dynamic multi-agent manipulation challenging.
The fundamental tension in deploying neural representations for closed-loop E2E testing arises from representational conflicts when separating persistent infrastructure from dynamic traffic agents. Merging these distinct entities introduces three concrete bottlenecks: first, composing multiple detailed 3D vehicle assets rapidly exhausts GPU memory and causes severe frame-rate drops (efficiency conflict); second, unstructured neural coordinates fail to align with semantic metric maps, causing inserted vehicles to suffer lateral drift, floating, or ground penetration (geometric conflict); and third, lighting entangled in neural radiance fields prevents inserted dynamic vehicles from adapting to local ambient illumination and shadows, creating prominent visual incoherence (photometric conflict).
This paper's angle of attack is to decompose the scene fundamentally into a persistent static background in world coordinates and manipulable dynamic agents defined in a canonical coordinate frame, resolving all three conflicts without online neural network inference. Core idea: by establishing a decoupled canonical 3DGS architecture that combines semantic-aware asset compression, map-guided geometric registration and vertical grounding, and proxy-based local probe relighting with contact shadows, DecoupleGS achieves centimeter-level metric consistency, natural illumination harmony, and interactive 45 FPS multi-agent closed-loop simulation.
Method¶
Overall Architecture¶
DecoupleGS provides an interactive 3D Gaussian Splatting sensor simulation pipeline tailored for closed-loop E2E autonomous driving testing. The inputs are an offline-reconstructed static background, a pre-built canonical vehicle asset library, planner/behavior-driven trajectories, and HD map topology; the output is a multi-view video stream driving the E2E planner. Spatially, the global scene \(\Omega(t)\) at timestamp \(t\) is explicitly decomposed into a time-invariant static background field \(S_{bg}\) anchored in the World Coordinate System (WCS) and \(K\) dynamic agents \(V_k\) parameterized within a standardized Local Coordinate System (LCS). Given an agent's 6-DoF rigid transform \(T_k(t) = [R_k \mid t_k] \in SE(3)\), spatial attributes (means \(\mu_i^{(L)}\) and covariances \(\Sigma_i^{(L)}\)) are transformed to WCS:
Simultaneously, Spherical Harmonics (SH) coefficients \(c_i^{(L)}\) encoding view-dependent appearance are rotated via the Wigner D-matrix: \(c_i^{(W)} = \mathbf{D}(R_k) c_i^{(L)}\), preserving physical specularity under orientation changes. To resolve efficiency, geometric, and photometric conflicts, the framework sequentially chains asset compression, map-guided registration, and proxy-based relighting before passing all primitives to a unified rasterizer.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Static Background & Canonical Assets Input<br/>S_bg (WCS) & V_k (LCS) Coordinate Decoupling"] --> B["Semantic-Aware Asset Compression<br/>Explicit Importance Pruning + Dual-Codebook Vector Quantization"]
B --> C["Map-Guided Geometric Registration<br/>Lane Topology Constrained DTW + Opacity-Accumulated Grounding"]
C --> D["Proxy-Based Relighting & Contact Shadows<br/>Local Probe Linear SH Transfer + Parametric Dynamic Shadows"]
D --> E["Unified Depth Sorting & Rasterization<br/>Frustum Culling & Static Radix Cache Acceleration"]
E --> F["Closed-Loop Sensor Simulation Output<br/>Multi-View Sensor Stream Driving E2E Policy Feedback"]
Key Designs¶
1. Semantic-aware asset compression: breaking multi-agent memory and throughput bottlenecks
To resolve the efficiency conflict where uncompressed 3DGS vehicles overload memory and collapse rendering throughput during dense multi-agent simulation, the framework avoids computationally heavy neural pruning and adopts an explicit importance scoring and vector quantization scheme. First, an explicit scalar importance score \(s_i\) is computed for each primitive \(g_i\):
where \(V_i\) denotes visibility (expected opacity contribution over training views), \(D_i\) measures color contrast against the local background, and \(H_i\) captures local texture entropy. Primitives with \(s_i\) below a strict threshold \(\tau_s = 0.005\) are pruned, eliminating internal invisible or redundant floaters and focusing computation on salient boundaries (wheels, contours, and chassis). Second, retained high-dimensional attributes (covariances \(\Sigma_i\) and color SH coefficients \(c_i\)) undergo Vector Quantization (VQ) against compact learnable codebooks \(\mathcal{C}_{\text{shape}}\) (512 entries) and \(\mathcal{C}_{\text{color}}\) (1024 entries), updated via Exponential Moving Average (EMA). Floating-point tensors are replaced with compact integer indices, slashing asset memory footprints by roughly 80% and enabling smooth multi-vehicle rendering up to 50 concurrent agents without GPU out-of-memory errors.
2. Map-guided geometric registration: eliminating lateral drift and ground penetration
To address geometric conflicts between unconstrained neural coordinates and metric map space, where inserted vehicles drift across lanes or sink into asphalt, DecoupleGS combines topological trajectory alignment with opacity-accumulated 3D grounding. In the horizontal (2D) plane, vectorized lane centerlines are obtained from HD maps or topological extractors (e.g., MapTRv2). A cost matrix penalizing Euclidean distance and heading deviation is constructed between the planned trajectory and the candidate lane. Constrained Dynamic Time Warping (DTW) identifies the optimal temporal correspondence path, and Orthogonal Procrustes analysis solves for a global rigid 2D transform \(T_{\text{align}} \in SE(2)\), eliminating systematic lateral drift. In the vertical (3D) dimension, because 3DGS lacks a continuous surface mesh for standard ray casting, vertical grounding is achieved via opacity accumulation. At bottom anchor points \(p_j\) of the vehicle bounding box, intersecting background Gaussians along a vertical column \(\mathcal{N}\) are opacity-weighted:
A local ground plane is fitted across anchor heights \(\{z_j\}\) via least squares to extract the surface normal, deriving the asset's vertical position \(z_{\text{asset}}\) alongside pitch and roll to ensure wheels maintain physically plausible contact with the ground.
3. Proxy-based relighting and contact shadows: achieving seamless photometric integration without neural overhead
To resolve photometric conflicts where inserted assets retain foreign illumination and lack scene-consistent shadows, DecoupleGS introduces an analytical lighting transfer operator and dynamic contact shadows without per-frame neural network inference. Background Gaussians serve as dense spatial light probes. For an asset at position \(x\), an ambient descriptor \(\mathcal{L}(x)\) aggregates neighboring background SH coefficients \(c_g\) weighted by Gaussian distance and visibility \(V_g\):
Environmental lighting is transferred to the vehicle's canonical SH coefficients \(c_{\text{can}}\) via a pre-calibrated linear modulation operator: \(c_{\text{out}} = \mathcal{T}(\mathcal{L}(x)) \otimes c_{\text{can}} + b(\mathcal{L}(x))\). Transformation matrices \(\mathcal{T}\) and bias \(b\) are pre-computed offline via least-squares regression over canonical assets rendered under diverse light probes, replacing runtime MLP forward passes with lightweight linear operations. For contact shadows, a parametric footprint mask \(M_{\text{shadow}}(u)\) is projected onto the estimated ground plane, and its intensity is dynamically modulated by dominant intensity \(I_{\text{dom}}\) from the DC term of \(\mathcal{L}(x)\): \(c_{\text{final}}(u) = c_{\text{ground}}(u) (1 - I_{\text{dom}} M_{\text{shadow}}(u))\), producing sharp shadows in direct sun and soft diffusion in overcast weather.
Loss & Training¶
Static backgrounds are optimized over 30k iterations using standard 3DGS losses combining photometric \(L_1\) and D-SSIM: \(\mathcal{L}_{\text{rgb}} = (1 - \lambda) L_1 + \lambda (1 - \text{SSIM})\) with \(\lambda = 0.2\). Dynamic vehicles are segmented using SAM and optimized canonically. Codebook entries in the VQ compression module are updated during compression using an EMA decay of 0.99. During online simulation, all modules rely on analytical closed-form operations or cached lookups without online gradient updates. For maximal hardware throughput, the rasterizer integrates conservative frustum culling, cached static background sorting keys, and batched parallel spatial transformations for moving agents, achieving consistent 45 FPS multi-camera rendering on an NVIDIA RTX 4090.
Key Experimental Results¶
Main Results¶
The framework is evaluated on nuScenes and PandaSet sequences encompassing dusk, night, and overcast conditions, with 20 vehicle models sampled from 3DRealCar. Table 1 summarizes asset compression and rendering efficiency under varying traffic densities (Size S: 1–2 vehicles, Size M: 3–5 vehicles, Size L: 6–10+ vehicles).
Table 1: Asset compression and rendering efficiency under varying traffic densities
| Size | Methods | PSNR-Veh ↑ | PSNR-All ↑ | SSIM ↑ | LPIPS ↓ | FPS ↑ | VRAM ↓ | Top-2 Consist. |
|---|---|---|---|---|---|---|---|---|
| S (1-2 veh) | Plenoxels | 21.45 | 23.06 | 0.795 | 0.510 | 8.5 | 2.5 GB | No |
| Vanilla 3DGS | 28.52 | 29.41 | 0.903 | 0.243 | 55.4 | 1.2 GB | No | |
| LightGaussian | 26.21 | 27.50 | 0.865 | 0.281 | 75.2 | 180 MB | No | |
| DecoupleGS (Ours) | 28.10 | 29.25 | 0.898 | 0.252 | 68.5 | 850 MB | Yes | |
| M (3-5 veh) | Plenoxels | 20.12 | 23.08 | 0.626 | 0.463 | 6.2 | 3.1 GB | No |
| Vanilla 3DGS | 27.05 | 27.21 | 0.815 | 0.214 | 32.5 | 2.1 GB | No | |
| LightGaussian | 25.14 | 26.40 | 0.781 | 0.250 | 58.6 | 240 MB | No | |
| DecoupleGS (Ours) | 26.85 | 27.12 | 0.805 | 0.218 | 52.4 | 950 MB | Yes | |
| L (6-10+ veh) | Plenoxels | 18.55 | 21.08 | 0.719 | 0.379 | 3.5 | 4.8 GB | No |
| Vanilla 3DGS | 25.50 | 25.14 | 0.785 | 0.255 | 12.1 | 3.8 GB | No | |
| LightGaussian | 23.52 | 24.10 | 0.720 | 0.295 | 45.0 | 320 MB | No | |
| DecoupleGS (Ours) | 25.24 | 25.05 | 0.772 | 0.260 | 38.5 | 1.1 GB | Yes |
Table 2 evaluates open-loop sim-to-real consistency and closed-loop E2E simulation against state-of-the-art neural simulators using UniAD.
Table 2: Open-loop sim-to-real consistency and closed-loop E2E simulator comparison
| Setting | Metric | Real Input | DecoupleGS (Ours) | HUGSIM | RealEngine | OASim |
|---|---|---|---|---|---|---|
| Open-loop | mADE (m) ↓ | 0.76 | 0.82 | 0.94 | 1.15 | 1.02 |
| minTTC (s) ↑ | 3.5 | 3.3 | 2.8 | 2.4 | 2.9 | |
| Closed-loop | Driving Score (DS) ↑ | – | 0.884 | 0.765 | 0.682 | 0.748 |
| Route Completion (RC) ↑ | – | 0.956 | 0.814 | 0.795 | 0.851 | |
| minTTC (s) ↑ | – | 3.3 | 2.3 | 2.6 | 2.8 | |
| Rendering FPS ↑ | – | 45 | 12 | 32 | 18 |
Ablation Study¶
System-level ablations in Table 3 isolate the contribution of asset compression, map-guided registration, and proxy-based relighting on medium-density scenarios.
Table 3: System-level ablation study on medium-density scenarios
| Config | FPS ↑ | Trajectory ADE (m) ↓ | Peak Angular Error (°) ↓ | Note |
|---|---|---|---|---|
| w/o asset compression | 8.5 (↓26.9) | 0.05 (=) | 6.8 (=) | Primitive bloat drops throughput below interactive thresholds |
| w/o map-guided reg. | 36.2 (↑0.8) | 0.48 (↑0.43) | 6.8 (=) | Absence of map alignment leads to severe lateral drift and penetration |
| w/o relighting | 38.5 (↑3.1) | 0.05 (=) | 48.5 (↑41.7) | Unadjusted foreground illumination creates severe photometric error |
| Full Framework (Ours) | 35.4 | 0.05 | 6.8 | Balanced trade-off across throughput, geometry, and photometry |
In scenario stress testing across difficulty tiers (Easy, Medium, Hard, Extreme over 50 episodes), UniAD's Driving Score declines from \(0.725 \pm 0.04\) (Easy) to \(0.195 \pm 0.07\) (Extreme), while VAD drops from \(0.680 \pm 0.05\) to \(0.120 \pm 0.05\), validating the simulator's efficacy in generating controllable stress scenarios that uncover corner-case failure modes.
Key Findings¶
- Decoupled compression enables true interactive scaling: While Vanilla 3DGS achieves high visual fidelity, it suffers memory explosion and OOM failures when scaling toward 50 agents. DecoupleGS restricts memory growth to near-linear scaling, sustaining 38.5 FPS under large traffic densities.
- Superior open-loop policy fidelity: Under UniAD evaluation, DecoupleGS achieves an mADE of 0.82 m (closely approaching real images at 0.76 m), outperforming HUGSIM (0.94 m) and RealEngine (1.15 m).
- Map topology grounding is essential: Omitting Procrustes alignment and opacity-weighted vertical grounding increases trajectory displacement error by nearly an order of magnitude (from 0.05 m to 0.48 m).
Highlights & Insights¶
- Decoupled Canonical Gaussian Representation: Separates the scene into a time-invariant static background and standardized canonical dynamic volumes, synchronizing view-dependent SH reflections via the Wigner D-matrix to maintain physical validity under arbitrary rotations.
- Analytical Linear Probe Relighting: Avoids per-frame neural network inference by aggregating local Gaussian light probes and modulating vehicle SH coefficients via pre-calibrated linear matrices, achieving instant, natural illumination adaptation.
- Solver-Agnostic Plug-and-Play Modularity: Experiments demonstrate that the compression, registration, and relighting modules can be transplanted directly into third-party neural simulators (such as HUGSIM), delivering universal latency and photometric benefits.
Limitations & Future Work¶
- Approximated global illumination: The proxy-based linear probe transfer and parametric contact shadow projection do not model multi-bounce indirect illumination, concave vehicle self-shadowing, or specular ground reflections.
- Dependency on structured map topology: Map-guided registration relies heavily on accurate lane centerlines and planar ground assumptions, presenting challenges in multi-level interchanges, underpasses, or off-road settings with degraded lane markings.
- Sensor modality and non-rigid agent scope: The current platform is tailored for rigid vehicles and RGB camera observations; extending the framework to non-rigid pedestrians, cyclists, and multi-modal LiDAR/radar sensor simulation remains future work.
Related Work & Insights¶
- vs Vanilla 3DGS [Kerbl et al., 2023]: Vanilla 3DGS entangles static and dynamic elements and lacks runtime compression, causing severe frame drops and OOM when inserting dynamic vehicles. DecoupleGS introduces spatial decoupling and dual-codebook VQ compression, supporting 50+ vehicles in real time.
- vs HUGSIM [Zhou et al., 2024] / RealEngine [Li et al., 2024]: Existing neural driving simulators lack map-topology grounding and lightweight illumination transfer, leading to trajectory drift, ground penetration, and photometric mismatch (Driving Scores of 0.765 and 0.682). DecoupleGS achieves a superior Driving Score of 0.884 at 45 FPS.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Formulates a holistic decoupled 3DGS architecture targeting the efficiency, geometric, and photometric conflicts of multi-agent simulation.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluation across nuScenes and PandaSet, multi-density efficiency benchmarks, open-loop consistency, ablation studies, and closed-loop E2E planner evaluations.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear mathematical formulations, structured narrative, and clean alignment across architecture, diagram, and key designs.
- Value: ⭐⭐⭐⭐⭐ Provides an efficient, high-fidelity, and plug-and-play closed-loop sensor simulation engine for end-to-end autonomous driving safety validation.