Skip to content

TIDES: Time-Derivative Event Simulation via Deformable Reconstruction

Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: 3D Vision
Keywords: event camera, event simulation, 4D Gaussian splatting, time derivatives, occlusion dynamics

TL;DR

Addressing the pervasive timestamp batching and occlusion distortion caused by discrete frame differencing in conventional event simulators, TIDES leverages dynamic 4D Gaussian splatting with forward-mode time-derivative rasterization to analytically solve continuous-time threshold crossings directly from explicit 3D dynamics, combined with occlusion-risk adaptive time-stepping and a physical tile readout arbiter.

Background & Motivation

Event cameras asynchronously report per-pixel log-intensity changes with microsecond temporal resolution and high dynamic range, making them exceptionally well-suited for agile robot perception, high-speed visual odometry, and extreme-illumination imaging. However, annotated real-world event datasets remain scarce and expensive to capture at scale, rendering high-fidelity event simulation indispensable for training and evaluating neuromorphic vision algorithms. Mainstream simulators—such as ESIM, v2e, and DVS-Voltmeter—predominantly rely on discretely sampled frame sequences, estimating threshold-crossing timestamps by interpolating or differencing consecutive rendered frames. This reliance induces a severe failure mode termed "timestamp batching," where numerous continuous physical threshold crossings are unnaturally collapsed onto a tiny set of timestamps determined strictly by the rendering clock.

The fundamental physical tension lies in the nature of optical luminance changes: real-world brightness transitions arise from 3D objects with sharp, discrete geometric boundaries moving continuously through space, rather than smooth image-space interpolations. Approximating temporal intensity slopes via discrete finite differences inevitably blends completely distinct visibility regimes whenever occlusions or disocclusions occur between frames. While cranking up rendering rates or deploying video interpolation models can marginally reduce batching, doing so is computationally prohibitive and injects glaring visual artifacts near rapid motions. Dynamic 4D Gaussian splatting (4DGS) offers explicit 3D geometry, continuous motion, and strict front-to-back compositing order; however, merely sampling dense 2D images from 4DGS and differencing them leaves the root cause of timestamp batching unaddressed.

TIDES resolves this bottleneck by differentiating the 3D-to-2D rendering process directly with respect to continuous time, bypassing intermediate discrete frame differencing entirely. Core idea: propagate time derivatives analytically through dynamic 4D Gaussian splatting deformation and front-to-back rasterization via forward-mode automatic differentiation, producing visibility-consistent instantaneous log-luminance rates and occlusion dynamics to directly solve continuous-time threshold crossings while adaptively refining time steps in high-risk occlusion regions.

Method

Overall Architecture

TIDES operates directly on a continuous-time 4D dynamic Gaussian scene representation to infer per-pixel event timestamps from true 3D structure and motion. Within each simulation interval, time advances with an adaptive step \(\Delta t\). A dedicated forward-mode time-derivative rasterizer computes both primal appearance and instantaneous time derivatives under the identical front-to-back blending order. These signals feed a continuous-time residual-state comparator that analytically solves for threshold crossings. Simultaneously, occlusion risk indicators dynamically adapt the simulation step size, and an optional tile-level arbiter accurately reproduces hardware readout bottleneck effects.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Continuous-time 4DGS Scene<br/>Static Gaussians + Deformation MLP"] --> B["Forward-Mode Time-Derivative Rasterization<br/>JVP yields (L, ˙L, α, ˙α)"]
    B --> C["Continuous-Time Residual-State Comparator<br/>Multi-crossing Root Solving & State Update"]
    B --> D["Occlusion-Risk Adaptive Stepping<br/>Local Δt Refinement guided by (α, ˙α, b, c)"]
    D -.->|Adaptive Sub-step Query| B
    C --> E["Dynamics-Conditioned Tile Readout Arbiter<br/>Burst Queueing, Jitter & Event Drops"]
    E --> F["High-Fidelity Continuous-Time Event Stream"]

Key Designs

1. Forward-Mode Time-Derivative Gaussian Splatting: Eliminating Finite-Difference Errors with Visibility Consistency Conventional simulators approximate the rate of change in brightness using finite differences between rendered frames, which conflates distinct objects across occlusion boundaries. TIDES leverages forward-mode automatic differentiation (Jacobian-vector products, JVP) along the time axis through the deformation network, obtaining exact instantaneous derivatives of Gaussian parameters \((\dot{\boldsymbol{\mu}}_k, \dot{\mathbf{s}}_k, \dot{\mathbf{q}}_k, \dot{\sigma}_k, \dot{\mathbf{c}}_k)\) in a single forward evaluation. During projection and front-to-back volume accumulation, the time derivatives of transmittance \(T_k\) and blending weights \(w_k\) are propagated analytically via the chain rule under the active visibility ordering: $\(\dot{w}_k = \dot{T}_{k-1} \tau_k + T_{k-1} \dot{\tau}_k, \quad \dot{T}_k = \dot{T}_{k-1}(1 - \tau_k) - T_{k-1} \dot{\tau}_k\)$ This enables the rasterizer to concurrently render primal RGB \(\mathbf{I}\) and its true instantaneous derivative \(\dot{\mathbf{I}}\), yielding an unbiased instantaneous log-luminance derivative \(\dot{L}(\mathbf{u}, t) = \frac{\dot{Y}(\mathbf{u}, t)}{Y(\mathbf{u}, t) + \epsilon}\). By sharing the exact compositing sequence, early termination, and clamping criteria, the calculated brightness rate remains strictly consistent with 3D scene occlusions without selecting an empirical finite-difference interval.

2. Continuous-Time Residual-State Comparator: Closed-Form Root Finding for Multi-Crossing Events To eradicate timestamp batching caused by fast motion within simulation steps, TIDES maintains a per-pixel reference signal \(S_{\text{ref}}(\mathbf{u})\). Across a simulation step \((t_0, t_0 + \Delta t]\), the log-signal \(S\) is modeled using the analytically rendered instantaneous slope \(\dot{S}_0\): $\(S(\mathbf{u}, t_0 + \delta) \approx S_0(\mathbf{u}) + \dot{S}_0(\mathbf{u}) \, \delta, \quad \delta \in (0, \Delta t]\)$ Threshold-crossing candidates for positive (\(\dot{S}_0 > 0\)) and negative (\(\dot{S}_0 < 0\)) polarity departures are solved directly in closed form: $\(\delta_{+}(\mathbf{u}) = \frac{S_{\text{ref}}(\mathbf{u}) + C_{\mathbf{u}}^{+} - S_0(\mathbf{u})}{\dot{S}_0(\mathbf{u})}, \quad \delta_{-}(\mathbf{u}) = \frac{S_{\text{ref}}(\mathbf{u}) - C_{\mathbf{u}}^{-} - S_0(\mathbf{u})}{\dot{S}_0(\mathbf{u})}\)$ The earliest valid root \(\delta^* \in (0, \Delta t]\) generates an event with continuous timestamp \(t^* = t_0 + \delta^*\). The reference state updates as \(S_{\text{ref}} \leftarrow S_{\text{ref}} \pm C_{\mathbf{u}}^{\pm}\), and the root-finding process repeats iteratively over the remainder of the interval \((t^*, t_0 + \Delta t]\) for up to \(N_{\max}\) crossings per step. When a boundary render \(S_1\) is available, a quadratic model is alternatively fitted to capture signal curvature, completely eliminating the clustering of event timestamps on discrete grid boundaries.

3. Occlusion-Risk Adaptive Stepping: Concentrating Compute Where Local Linearity Fails Uniform temporal stepping squanders compute on static regions while under-sampling rapid disocclusions where linear brightness assumptions break down. During front-to-back compositing, TIDES extracts three lightweight geometric diagnostics at negligible computational overhead: the instantaneous transmittance rate \(|\dot{\alpha}(\mathbf{u}, t)|\) (indicating rapid visibility switching), the active contributor count \(c(\mathbf{u}, t)\) (indicating depth competition), and the accumulated coverage mass \(b(\mathbf{u}, t)\). These are combined into a per-pixel risk metric: $\(r(\mathbf{u}, t) = \mathrm{clip}\Big(w_b(1 - \mathrm{norm}(b)) + w_c \, \mathrm{norm}(c) + w_\alpha \, \mathrm{norm}(|\dot{\alpha}|)\Big)\)$ Where risk \(r\) is elevated, the time step contracts as \(\Delta t \leftarrow \Delta t \cdot \kappa(r)\) with \(\kappa(r) \in (0, 1]\), requesting local sub-pose rendering passes only where complex occlusion dynamics occur. Furthermore, spurious firings in poorly reconstructed background regions are effectively suppressed by masking pixels where \(b < b_{\min}\).

4. Dynamics-Conditioned Tile Readout Arbiter: Faithfully Reproducing Hardware Burst Saturation Physical event sensors serialize data through column/row arbiter buses with finite bandwidth; massive event bursts trigger queuing latencies, timestamp spreading, and event packet drops. TIDES incorporates a tile-scanning FIFO arbiter whose maximum throughput \(K_m\), drop probability \(p_{\text{drop}, m}\), and timestamp jitter \(\sigma_{t, m}\) are dynamically conditioned on a tile activity indicator \(\psi_m\) (aggregating normalized event rate \(|\dot{L}|/C\), transmittance rate \(|\dot{\alpha}|\), and contributor count \(c\)): $\(K_m = \lfloor K_0 (1 - \beta_K \psi_m) \rfloor, \quad p_{\text{drop}, m} = p_0 + \eta \psi_m, \quad \sigma_{t, m} = \sigma_0 + \gamma_\sigma \psi_m\)$ Under high scene activity and motion bursts, tile throughput is dynamically throttled while drop rates and jitter increase, accurately reproducing real-world neuromorphic sensor degradation under saturation without relying on ad-hoc empirical noise injections.

Loss & Training

The simulation pipeline is an analytical forward-mode process with no learnable weights. The underlying 4D Gaussian splatting representation is trained offline on multi-view or monocular RGB video sequences using standard photometric losses (\(\mathcal{L}_1\) and \(\mathcal{L}_{\text{SSIM}}\)) accompanied by temporal smoothness and rigidity regularizers. Uncalibrated camera trajectories are initialized via COLMAP for static backgrounds and Easi3R for moving camera setups in dynamic scenes.

Key Experimental Results

Main Results

Evaluation was conducted on four diverse benchmarks encompassing ego-motion (EDS, DSEC) and dynamic objects (HS-ERGB, BS-ERGB). Metrics include inter-spike interval negative log-likelihood (IG-NLL \(\downarrow\)) and motion-scaled spatiotemporal Chamfer distance (Chamfer \(\downarrow\)).

Simulator EDS: IG-NLL ↓ EDS: Chamfer ↓ HS-ERGB: IG-NLL ↓ HS-ERGB: Chamfer ↓ BS-ERGB: IG-NLL ↓ BS-ERGB: Chamfer ↓ DSEC: IG-NLL ↓ DSEC: Chamfer ↓
ESIM 0.004748 0.046932 0.006763 0.021335 0.007669 0.039233 0.096768 0.150414
V2E 0.004058 0.044124 0.005823 0.016149 0.006092 0.035944 0.099792 0.153248
ICNS/IEBCS 0.004575 0.046710 0.008341 0.028272 0.010608 0.045028 0.092576 0.142781
DVS-Voltmeter 0.006050 0.050914 0.006388 0.014871 0.005163 0.033895 0.080542 0.133779
TIDES (10x) 0.003840 0.043245 0.004984 0.014892 0.004392 0.033002 0.068641 0.121962
TIDES (adaptive) 0.003733 0.042622 0.002343 0.013772 0.004251 0.032260 0.068439 0.120611

Ablation Study

Ablations on EDS scenes dissect the contribution of each module under an identical rendering substrate and sensor parameter configuration.

Config IG-NLL ↓ Chamfer ↓ Note
Full TIDES 0.00373 0.04262 Full proposed model
Finite-diff slope (˙L by FD) 0.00521 0.04801 Drastic degradation highlighting the necessity of analytical time derivatives
No multi-crossing (1 event max / step) 0.00439 0.04487 Truncating firings to 1 per step under fast motion degrades temporal fidelity
Fixed step size (no controller) 0.00486 0.04672 Disabling adaptive time stepping harms complex occlusion timing
No risk-gated refinement (no (α, ˙α, b, c)) 0.00458 0.04593 Removing geometry/visibility guidance leaves step adaptation unguided
No mixed-visibility masking (\(b < b_{\min}\)) 0.00422 0.04431 Failing to suppress under-reconstructed backgrounds induces spurious events
No \(\dot{\alpha}\) cue 0.00461 0.04711 Omitting transmittance velocity notably degrades occlusion boundary accuracy
No readout model (ideal timestamps) 0.00388 0.04294 Slightly reduces fidelity by omitting realistic hardware queueing saturation
Compute-matched TIDES (fixed #queries) 0.00405 0.04371 Equal compute with fixed uniform queries still underperforms adaptive risk guidance

Batching Diagnostics and Downstream Transfer

  • Timestamp Batching Metrics (Table 2): ESIM exhibits a Same-timestamp event fraction (Same-ts) of 0.282 on EDS, whereas TIDES (adaptive) drops it to 0.025 (a >90% reduction). In burst dispersion error (Fano-factor error) and ISI spike ratios, TIDES consistently matches real physical sensor statistics.
  • Downstream Sim-to-Real Transfer (Table 3): Downstream perception networks trained purely on simulated events and tested directly on real hardware streams demonstrate marked superiority:
  • Video Reconstruction (E2VID): 18.16 dB PSNR / 0.1997 LPIPS (vs. ESIM's 12.23 dB / 0.3283).
  • Video Frame Interpolation (TimeLens-XL): 34.31 dB PSNR / 0.0335 LPIPS (surpassing all baselines by over 3 dB).
  • Monocular Depth Estimation (EDA on DSEC): RMSE log error drops to 0.2795 (vs. DVS-Voltmeter's 0.2865).
  • Semantic Segmentation (ESS on DSEC): mIoU achieves 17.69% (substantially exceeding V2E's 7.57% and DVS-Voltmeter's 16.92%).

Key Findings

  • Analytical time derivatives are the single most critical driver of fidelity: Substituting analytical \(\dot{L}\) with finite-difference slopes causes IG-NLL to jump from 0.00373 to 0.00521 and Chamfer distance to worsen by 12.6%, validating the danger of differencing across moving geometric boundaries.
  • Occlusion-guided adaptive stepping unlocks high gains in dynamic scenes: On HS-ERGB, where dynamic objects move across complex backgrounds and camera poses are sparse, risk-adaptive time stepping cuts IG-NLL from 0.00498 down to 0.00234 (>50% improvement).
  • Temporal continuity dictates synthetic-to-real transferability: Networks trained on batch-corrupted event streams overfit to artificial clock-aligned burst clusters, severely impairing their real-world generalization across diverse visual tasks.

Highlights & Insights

  • Unifying primitive dynamics and photometric derivatives in a single forward pass: By exploiting forward-mode JVP through both deformation and volume rendering, TIDES computes exact instantaneous log-luminance derivatives with zero temporal interpolation artifacts, retaining 3DGS's microsecond-level efficiency.
  • Geometry-aware, white-box adaptive time control: Instead of treating rendering step size as a blind hyperparameter, TIDES extracts transmittance rates \(|\dot{\alpha}|\) and contributor counts \(c\) directly from the alpha blending loop at zero marginal cost to steer compute where occlusion invalidates simple linear models.
  • Co-designing continuous physical generation with hardware readout arbitration: TIDES bridges the gap between ideal physical optical simulation and real-world circuit-level bottlenecks by modulating tile-based FIFO queue drop and jitter dynamics directly using renderer-native activity proxies.

Limitations & Future Work

  • Dependence on underlying 4DGS reconstruction fidelity: The quality of simulated events is intrinsically bounded by the 4D Gaussian representation; dynamic blur, floaters, or incomplete geometry in the reconstructed scene will produce corresponding event artifacts.
  • Order of local polynomial approximations: Currently, local trajectories are modeled with 1st-order linear or 2nd-order quadratic models; extreme angular accelerations or sub-pixel fluttering may require higher-order expansions or smaller base step sizes.
  • Differentiable event-to-geometry inverse optimization: TIDES focuses on forward event simulation; extending the pipeline to propagate continuous-time event loss gradients backwards into 4DGS parameters could enable end-to-end neuromorphic 3D dynamic reconstruction.
  • vs ESIM / v2e: Traditional simulators depend on discrete frame rendering and temporal interpolation, suffering from severe timestamp batching and boundary finite-difference slope errors; TIDES operates on continuous-time 4DGS derivatives to solve threshold crossings analytically.
  • vs DVS-Voltmeter / ICNS: Circuit-focused approaches excel at modeling analog noise and stochastic readout but still consume discretely sampled frame sequences; TIDES provides an accurate continuous-time driving signal and pairs it with a dynamics-conditioned arbiter.
  • vs Event-NeRF / Event-3DGS: While recent neural radiance and splatting works address the inverse task of recovering 3D scenes from event streams, TIDES is the first to pioneer the forward direction—leveraging learned 4DGS as a continuous-time substrate for high-fidelity physical event synthesis.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering forward-mode time-derivative 4D Gaussian rasterization for continuous-time event simulation.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Exhaustive evaluation across four real datasets, batching diagnostics, multiple downstream Sim-to-Real tasks, and comprehensive ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulations, clear pipeline presentation, and strong physical motivation.
  • Value: ⭐⭐⭐⭐⭐ Effectively solves the longstanding timestamp-batching dilemma, offering a powerful simulation foundation for neuromorphic vision.