Skip to content

Progressive Risk Estimation for Accident Anticipation

Conference: NeurIPS2026
arXiv: 2609.32811
Code: https://github.com/giddyyupp/PRE-ACT
Area: Autonomous Driving
Keywords: traffic risk prediction, continuous risk estimation, progress supervision, preference ranking, false-alarm evaluation

TL;DR

PRE-ACT uses continuous distance-to-event progress supervision and within-video clip ranking to improve traffic risk prediction from past-only sliding windows, raising CAP mAUC from TOP's 0.429 to 0.481 under FPR โ‰ค 0.1 and introducing a Separation Score to examine false-alarm tendencies across complete risk curves.

Background & Motivation

Dashcam warning systems must do more than recognize whether a scene is currently anomalous: they must increase risk scores as observable cues intensify without repeatedly warning during normal driving. DSTA and GSC strengthen video representations through spatio-temporal attention or graph structures, but their main supervision remains binary classification. TOP predicts event probabilities at multiple future timestamps, yet still treats those timestamps as discrete binary targets; two positive clips need not encode which one is closer to the event.

Simply raising scores earlier throughout every event video cannot solve this problem, because it may also increase false alarms during normal driving. Local AUC and mean time-to-accident (mTTA) examine selected temporal regions, potentially giving similar results to a curve that stays high prematurely and one that remains low during normal driving before rising progressively. The paper therefore changes both supervision and evaluation: event timestamps generate continuous risk targets during training, while evaluation examines average scores over the full regions before and after anomaly onset rather than only a few temporal windows.

Anomaly onset annotations are subjective and unavailable on Nexar. The authors define the training critical zone using a fixed 2-second pre-event window, avoid onset supervision during training, and sample later clips from the same video to establish ranking constraints. Core idea: turn distance-to-event progress into dense auxiliary supervision for shared video representations, so that the classifier learns the temporal structure of risk rather than only a positiveโ€“negative boundary.

Method

Overall Architecture

The input comprises 5 frames at and before the current time from an ego-view video, not the entire video or future frames. Videos are standardized to 10 FPS and 224ร—224; VideoMAE produces spatio-temporal tokens, which are average-pooled and passed to lightweight heads. The main warning branch outputs a classification score for the current clip, while additional continuous progress and preference ranking heads share its video features during training.

โ€œTemporal-Zone Supervision,โ€ โ€œContinuous Progress Supervision,โ€ and โ€œClip Preference Rankingโ€ respectively teach the model the positiveโ€“negative regions, risk variation within the critical region, and relative ordering of clips from the same video. Inference follows only the past-window โ†’ encoder โ†’ classification-head โ†’ current-score path. It requires neither event timestamps, anomaly onset, nor later paired clips: future information supplies offline training labels rather than online inputs.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Current and past 5 frames"] --> B["VideoMAE and average pooling"]
    B --> C["Classification head"]
    C --> D["Current warning score"]
    L["Training only: event time"] -.-> E["Temporal-Zone Supervision"]
    E -.-> F["Continuous Progress Supervision"]
    F -.-> G["Clip Preference Ranking"]
    E -.->|BCE| C
    B -.-> F
    B -.-> G
    H["Training only: shared encoder<br/>features of a later clip"] -.-> G

Solid arrows show online scoring; dashed arrows show training branches or supervision relationships. Paired clips in the ranking branch never enter the current online prediction.

Key Designs

1. Temporal-Zone Supervision: replace inconsistent onset annotations with a common temporal boundary

For a video containing an event, let its timestamp be \(t_{\text{acc}}\) and the critical-zone start be \(t_h\): earlier clips belong to the safe-driving zone, while clips from that boundary until the event belong to the critical zone. The implementation uses the 2 seconds preceding the event, so \(t_h\) denotes a temporal position, not an absolute frame index fixed at 2. For a video without an event, both boundaries are placed at the video end, leaving an empty critical zone and making all training clips negative.

The classification head learns a BCE target of 0 in the safe-driving zone and 1 in the critical zone. This retains an explicit event-discrimination task so that auxiliary progress learning remains relevant to warning generation. However, the boundary comes from dataset statistics rather than each video's actual visible cues: cues may emerge before the 2-second window or become observable considerably later. The training safe-driving zone therefore does not mean physically zero risk.

2. Continuous Progress Supervision: assign different targets to clips that are all positive

With BCE alone, the critical-zone entrance and the final pre-event moment both receive a positive target, providing no explicit instruction about how risk should grow. The authors instead construct a continuous target from relative time within the critical zone: it is strictly 0 in the safe-driving zone and rises from 0 toward 1 after entering the critical zone. An additional progress head fits this target with smooth L1 loss, encouraging shared representations to encode distance-to-event changes rather than only class membership.

\[ \tau=\frac{t-t_h}{t_{\text{acc}}-t_h},\qquad r_t=\begin{cases} 0,&t<t_h,\\ \dfrac{1-\exp(-\alpha\tau)}{1-\exp(-\alpha)},&t_h\leq t<t_{\text{acc}}. \end{cases} \]

The parameter \(\alpha\) controls curvature: positive values produce faster early growth, negative values shift pronounced growth toward later stages, and a linear target provides another comparison. The default is \(\alpha=3\), not a quantity predicted from video content. Safe videos have no nonempty critical zone, so the second branch is not evaluated with a zero denominator.

This target is a distance-to-event supervision proxy, not a calibrated hazard probability or direct regression of remaining seconds. A monotonic training target does not guarantee framewise monotonic predictions on test videos; visual evidence, occlusion, and domain shifts can still affect scores. The paper does not apply a hard monotonic projection during inference.

3. Clip Preference Ranking: identify states closer to the event within the same video

Absolute targets can be confounded by appearance differences: the same target value may correspond to very different roads or camera settings. The authors therefore select a later critical-zone clip from the same video, score both clips with an additional MLP, and use margin ranking loss to require a higher score for the later clip. The implementation samples paired clips 0.5โ€“1.5 seconds after the current clip; ordering supplies the constraint without human preference annotations.

Within-video comparison reduces scene differences, but its ranking assumption still comes from event proximity and does not imply that real risk always increases. This is an auxiliary training branch, not an online system looking 0.5โ€“1.5 seconds ahead. The source text does not specify the ranking margin value, so no unreported setting is supplied here.

A Worked Example

The following hypothetical timeline illustrates supervision, not a measured video result. If an event occurs at second 10, the training critical zone starts at second 8. A clip at second 7 belongs to the safe-driving zone, with both classification and progress targets equal to 0.

At second 8, the clip enters the critical zone: its classification target becomes 1, while its continuous progress target remains 0. At second 9, halfway through the zone, the classification target is still 1 but the progress target follows the equation above. Preserving this difference is the intended temporal structure; the progress and classification heads should not be interpreted as synonymous outputs.

If training pairs the clip at second 9 with one at second 9.5, the ranking head requires a higher score for the latter. During actual online operation at second 9, the model processes only the 5-frame window ending at second 9. It knows neither the actual event time nor the frames at second 9.5; those timestamps support offline optimization only.

Loss & Training

The three losses jointly update shared features: classification supplies the main task signal, progress supplies continuous temporal structure, and ranking supplies relative temporal structure. The source defines the combined objective as:

\[ \mathcal{L}=w_{\text{acc}}\mathcal{L}_{\text{BCE}}+w_{\text{prog}}\mathcal{L}_{1}+w_{\text{pref}}\mathcal{L}_{\text{rank}}. \]

The default weights are 1.0, 10.0, and 0.1, respectively. The encoder is VideoMAE-L pretrained on Kinetics-400; only its last 4 layers and the task heads are trained. Each task head has two linear layers with GELU. Encoder and head learning rates are \(10^{-5}\) and \(10^{-4}\), with batch size 32 and at most 50 negative clips per video to prevent normal clips from overwhelming critical clips.

Since VideoMAE normally accepts 16 frames, the 5-frame input is padded by repeating each of the first 4 frames 3 times and the final frame 4 times. This increases input length, not actual temporal context; the online model does not observe 16 distinct frames.

The main text reports 1 training epoch, whereas Appendix A.1 reports 2000 iterations. Both descriptions are retained because the cache does not explain their correspondence. The appendix reports 307M total parameters, 53M trainable parameters, and 14.44 inference FPS on the same NVIDIA A100. This does not establish real-time performance on automotive hardware.

Key Experimental Results

Main Results

MM-AU CAP contains 8,959 training and 800 test videos; DADA contains 1,771/198. The two subsets are trained and evaluated separately. Nexar contains 3,000/1,354 videos and uses CAP training followed by Nexar fine-tuning; its result is not a Nexar-only training comparison.

The following selection from source Table 1 uses FPR โ‰ค 0.1 throughout; mTTA is in seconds. Competitor results are taken from the TOP paper rather than uniformly retrained by the authors.

Dataset Method AUC 0.0 s AUC 0.5 s AUC 1.0 s AUC 1.5 s mAUC mTTA
CAP TOP 0.838 0.675 0.398 0.214 0.429 0.864
CAP PRE-ACT 0.887 0.775 0.465 0.201 0.481 1.125
DADA TOP 0.790 0.567 0.288 0.140 0.332 0.885
DADA PRE-ACT 0.861 0.751 0.398 0.170 0.440 1.069

CAP mAUC increases by 0.052, but its 1.5-second AUC falls from 0.214 to 0.201, so PRE-ACT does not outperform TOP at every warning horizon. Subtracting table values gives a CAP mTTA increase of 0.261 seconds, whereas the prose summarizes it as 0.25 seconds. DADA increases by 0.184 seconds, summarized as 0.18 seconds. The table values are preserved rather than adjusted to match those summaries.

Source Table 2 reports Nexar runner-up/winner/PRE-ACT mAP as 0.875/0.885/0.895. The claimed โ€œ+1 mAPโ€ means an absolute improvement of 0.010, or 1 percentage point when expressed as percentages, not an increase of 1.0 mAP.

Ablation Study

The following auxiliary-objective ablation comes from source Table 4 on CAP under FPR โ‰ค 0.1. BCE-only uses the paper's backbone and provides a closer controlled baseline.

Config AUC 1.5 s mAUC mTTA (seconds) Separation Score
BCE only 0.176 0.419 1.105 0.694
BCE + progress 0.197 0.477 1.128 0.723
BCE + progress + preference 0.201 0.481 1.125 0.724

Progress supervision supplies the main mAUC increase of 0.058, while preference ranking adds 0.004. Ranking slightly improves 1.5-second AUC but reduces mTTA from 1.128 to 1.125 seconds, so the full model is not best on every metric.

The Separation Score groups predictions by anomaly onset \(t_{\mathrm{ai}}\), not the training boundary: \(s_{\mathrm{pre}}\) is the mean prediction from video start until just before onset, and \(s_{\mathrm{post}}\) is the mean from onset through the event timestamp. Each region is normalized by its own number of predictions.

\[ s=\frac{1-s_{\mathrm{pre}}+s_{\mathrm{post}}}{2}. \]

The score lies in \([0,1]\); its ideal value of 1 means all scores are 0 before onset and 1 afterward. Any constant prediction yields 0.5 regardless of its level. It summarizes full-region means but does not measure post-onset monotonicity, thresholded false-alarm counts, or alarms per hour.

The following selection comes from source Table 6. Because pretrained competing models are unavailable, the authors report this metric only for their own BCE baseline and PRE-ACT.

Dataset Config Pre-onset mean (lower is better) Post-onset mean (higher is better) Separation Score
CAP BCE only 0.420 0.808 0.694
CAP PRE-ACT 0.344 0.792 0.724
DADA BCE only 0.372 0.763 0.696
DADA PRE-ACT 0.319 0.798 0.740

Key Findings

  • CAP's post-onset mean decreases slightly, yet its Separation Score improves mainly because the pre-onset mean falls from 0.420 to 0.344. The improvement does not come from raising all outputs together.
  • Source Table 5 reports mAUC/mTTA of 0.481/1.125 for \(\alpha=3\), 0.470/1.118 for a linear target, and 0.458/1.112 for \(\alpha=-3\). Positive curvature favors longer anticipation horizons, but \(\alpha=4\) also reaches 0.481/1.125, so the default is not uniquely optimal.
  • Tightening the FPR limit to 0.01 in source Table 7 reduces CAP/DADA mAUC to 0.210/0.226, with mTTA of 0.817/0.826 seconds. Timely warnings remain possible, but this does not establish that low-false-alarm operation is solved.
  • Source Table 3 gives mAUC 0.464 for direct CAP-to-DADA transfer, but only 0.361 for direct CAP-to-Nexar transfer, below the fine-tuned 0.510 in Appendix Table 9. In-domain performance cannot conceal cross-domain differences.

Highlights & Insights

  • The main intervention is supervision rather than a complex architectural expansion. Continuous targets turn binary-labeled videos into internally ordered learning signals, and the ablation shows that progress contributes more than ranking.
  • Classification is retained while progress and preference shape shared representations. Progress values need not be interpreted as real probabilities, making the approach closer to auxiliary representation learning for warnings.
  • Full-region evaluation exposes behavior missed by local metrics. Low scores during normal driving and high scores during the critical phase must both be checked; warning lead time alone can reward prematurely elevated risk.

Limitations & Future Work

  • The authors acknowledge that the work covers video monitoring only, excluding driver state, vehicle dynamics, and sensor reliability. It does not validate closed-loop control or actual reductions in driving risk.
  • A fixed 2-second boundary and monotonic target may disagree with actual cue onset. Variable horizons, risk-mitigation segments, and probability calibration should be tested rather than treating temporal distance as universally valid risk.
  • The Separation Score does not test monotonicity, and its ideal is an immediate high score after onset rather than gradual growth. Thresholded false-alarm frequency, detection latency, and repeated-warning statistics are still needed.
  • Nexar test videos omit event frames and lack anomaly onset. The appendix sets onset artificially to 2 seconds before the event, so its mTTA and separation results depend on a proxy onset and are not equivalent to human-annotated MM-AU evaluation.
  • Source Table 8 reports DADA mTTA of 1.011 for V-JEPAv2 and 1.069 for VideoMAE, whereas the appendix prose claims V-JEPA outperforms the default backbone on this metric. This tableโ€“text conflict does not support claiming that the replacement improves DADA warning lead time.
  • Multiple-seed uncertainty, automotive-device latency, and warning burden during extended normal driving are not reported. The authors identify snow and strong frontal sunlight as visibility-related failure causes; robustness still requires independent validation.
  • vs TOP: TOP predicts binary occurrence probabilities at multiple future timestamps; PRE-ACT constrains shared representations with continuous distance targets and clip ranking. It still trails TOP at CAP's earlier 1.5-second horizon, so this is not an improvement in every dimension.
  • vs DSTA / GSC: DSTA emphasizes spatio-temporal attention, while GSC emphasizes participant relationships and temporal continuity. PRE-ACT primarily changes supervision and uses a video foundation model; its entire gain should not be attributed to ranking.
  • vs robotic progress estimation: Both exploit relative trajectory position for dense feedback. In traffic prediction, the endpoint does not represent task success, and risk can be mitigated during a sequence, so monotonicity requires domain-specific validation.

Rating

  • Novelty: 4/5 โ€” Continuous progress and ranking supervision for traffic warnings, plus full-region separation evaluation, provide clearly defined contributions.
  • Experimental Thoroughness: 3/5 โ€” Multiple datasets, auxiliary objectives, curvature, FPR, and backbone analyses are included, but uncertainty and long-duration false alarms remain untested.
  • Writing Quality: 3/5 โ€” The method is intuitive, but training descriptions, some optimality claims, and backbone-result prose are inconsistent.
  • Value: 4/5 โ€” The supervision approach is reusable for temporal warnings, while evaluation and deployment gaps remain before real automotive safety use.