Skip to content

ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction

Conference: ECCV 2026
Paper: ECCV Official
OpenReview: ECCV 2026 #5593
Code: https://github.com/am3338/ECTraj
Area: Autonomous Driving
Keywords: trajectory prediction, consistency models, multi-agent interaction, diffusion acceleration, single-step inference

TL;DR

ECTraj introduces an enhanced conditional Consistency Model pipeline for multi-agent trajectory prediction that couples teacher waypoint ground-truth fusion with training-time top-K consistency matching, achieving single-step generation with a 10x inference speedup and superior accuracy over diffusion baselines.

Background & Motivation

Multi-agent trajectory prediction is an essential foundation for motion planning in autonomous driving and robotics, requiring models to forecast multi-modal, physically plausible, and socially compliant future paths conditioned on historical observations, high-definition HD maps, and social interactions. Diffusion-based trajectory predictors (such as MID, LED, MotionDiffuser, and OptTrajDiff) have achieved remarkable success in modeling complex multimodal distributions thanks to their flexible stochastic generative capabilities. However, standard diffusion models depend on iterative reverse denoising governed by stochastic differential equations, requiring dozens or hundreds of sequential evaluations. Even with fast-sampling variants like DDIM, more than 10 denoising steps remain mandatory, incurring prohibitive inference latencies that restrict practical deployment on real-time vehicle computing platforms.

To mitigate inference overhead, existing acceleration methods often replace standard Gaussian white noise with an informed initial distribution (for example, OptTrajDiff initializes denoising around the predictions of a pre-trained QCNet backbone). While this warm start reduces required diffusion steps, it binds the generation strictly to the vicinity of the prior: any flawed prediction or missed mode in the pre-trained prior misleads the entire denoising path, severely limiting free exploration across diverse plausible futures. Consistency Models (CMs) present a compelling alternative by learning self-consistency mappings along the Probability Flow ODE (PF-ODE) that directly project standard Gaussian noise to clean data in a single step. Yet, Consistency Distillation (CD) requires pre-existing, high-capacity diffusion teacher models that are virtually absent in the trajectory prediction domain, while training Consistency Models entirely from scratch (Consistency Training, CT) suffers from extreme training instability, weak supervision, and convergence difficulties.

To overcome this dilemma, this work designs a conditional consistency training framework trained from scratch without relying on diffusion distillation or restrictive prior-informed noise initialization. Core idea: build an enhanced consistency training pipeline incorporating deterministic teacher-side ground-truth waypoint fusion and training-time best-of-K multi-shot consistency matching, unifying single-step 10x inference acceleration with state-of-the-art multimodal trajectory prediction accuracy.

Method

Overall Architecture

ECTraj consists of a scene context encoder, a compact latent compression mapping, a composite attention conditional denoiser, and an enhanced consistency training schedule. The input comprises surrounding multi-agent historical states \(X \in \mathbb{R}^{N_a \times T_h \times 2}\) and polyline vector map features \(\mathcal{M} \in \mathbb{R}^{N_m \times D_p \times D_m}\). The model first extracts scene context representations \(c\) using a pre-trained backbone and derives marginal future trajectory proposals alongside confidence scores \((mm, p)\), forming the combined context \(C = [c, mm, p]\). To minimize computational dimensionality, future coordinate sequences are compressed into 10-dimensional latent vectors via learnable linear projections. During training, the pipeline samples standard Gaussian noise, executes progressive discretization time scheduling, generates \(K\) candidate trajectories in a single forward pass to identify best-matching modes, and deterministically injects ground-truth midpoint and endpoint waypoints into the teacher's latent prediction to provide strong, stabilized supervision for student updates.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input History & Vector Map<br/>(X, M)"] --> B["Scene Context & Marginal Prior<br/>Extract C=[c, mm, p]"]
    B --> C["Composite Attention Denoiser<br/>Cross-attends history/map/priors"]
    C --> D["Progressive Discretization Schedule<br/>Dynamic N steps with EMA teacher index"]
    D --> E["Training-Time Top-K Multi-Shot Matching<br/>Single-step K modes & min-ADE filtering"]
    E --> F["Teacher Ground-Truth Fusion<br/>Midpoint/endpoint replacement supervision"]
    F --> G["Single-Step Inference Output<br/>1-step latent decoded to future paths"]

Key Designs

1. Composite Attention Denoiser: lossless structural and prior integration under standard Gaussian noise
Prior trajectory diffusion methods relying on informed initial distributions suffer when the prior distribution is inaccurate. ECTraj enforces sampling from standard Gaussian white noise \(x_\sigma \sim \mathcal{N}(0, I)\) to preserve an unbiased generative space, but this requires an expressive architecture to align pure noise with complex spatial-temporal geometries. ECTraj introduces a 3-layer composite attention block \(F_\theta\). Within each layer, intermediate noisy latent states \(x_\sigma\) sequentially undergo: (1) cross-attention with target agent history embeddings; (2) cross-attention with map polyline embeddings; (3) cross-attention with neighboring agent context tokens; (4) self-attention across future representations of all participating agents to capture joint game-theoretic interactions; and (5) cross-attention with marginal trajectory priors \((mm, p)\) derived from a pre-trained QCNet. Following Karras formulation, the network output is parameterized with skip connections: $\(f_\theta(x, \sigma, C) = c_{\text{skip}}(\sigma) x + c_{\text{out}}(\sigma) F_\theta(x, \sigma, C)\)$ ensuring continuous boundary preservation where \(c_{\text{skip}}(\sigma_{\min}) = 1\) and \(c_{\text{out}}(\sigma_{\min}) = 0\).

2. Progressive Discretization Schedule: smooth transition from reconstruction to consistency
Training consistency models from scratch is prone to collapse due to volatile teacher predictions in early training stages. ECTraj implements an adaptive discretization curriculum across \(E\) total epochs. The number of ODE discretization steps \(N\) starts at \(N=10\) and doubles every \(E//3\) epochs, reaching \(N=40\) near convergence. After sampling a student index \(t \sim \text{DiscreteLogNormal}(\mu, \Sigma, N)\), the teacher index \(r\) is parameterized via continuous decay: $\(r = \left\lfloor t \left(1 - \frac{n(t')}{q^{\lceil e/d \rceil}}\right) \right\rfloor, \quad n(t') = 1 + \frac{k}{1 + e^{b t'}}\)$ with hyperparameters \(k=8\), \(b=1\), \(q=4\), current epoch \(e\), and phase threshold \(d = E//3\). In initial epochs, the penalty term drives \(r=0\), clamping the objective to an exact clean reconstruction target; as training progresses and the denominator expands, \(r\) smoothly approaches \(t\), transitioning seamlessly into the standard consistency regime. Combined with EDM variance scheduling \(\sigma_t \in [0.002, 1.04]\) (\(\rho=7\)), this schedule guarantees stable convergence without requiring a pre-trained diffusion teacher.

3. Training-Time Top-K Multi-Shot Matching: exploiting single-step nature for multimodal specialization
Conventional noise-predicting diffusion models (like OptTrajDiff) optimize an \(L_2\) error in the Gaussian noise space; because individual noise samples do not correspond to interpretable geometric trajectories before full reverse denoising, evaluating best-of-K loss during training is intractable. In contrast, consistency models map any noisy input directly to the clean trajectory manifold in one step. ECTraj leverages this capability by sampling \(K=6\) independent random Gaussian noise vectors \(\epsilon_k \sim \mathcal{N}(0, I)\) for both student and teacher, denoising all \(K\) modes simultaneously in a single forward pass, and picking the best-matching mode according to Average Displacement Error (ADE): $\(k_\theta = \arg\min_k \text{ADE}(f_\theta(x_{\sigma_t, k}, \sigma_t, C), x_0), \quad k_{\theta^-} = \arg\min_k \text{ADE}(f_{\theta^-}(x_{\sigma_r, k}, \sigma_r, C), x_0)\)$ This minimum-error selection aligns specific noise vectors with distinct future modes (e.g., left turn, cruising, lane change), successfully preventing mode averaging and dramatically enhancing multimodal coverage.

4. Teacher Ground-Truth Fusion: dynamic anchor guidance breaking self-consistent error propagation
In standard consistency training, the student is supervised entirely by the teacher's EMA output; when the teacher deviates, errors propagate and compound. However, replacing the teacher output entirely with ground truth forces an aggressive single-step reconstruction problem that fails to converge under compact step budgets. ECTraj establishes an optimal trade-off via deterministic waypoint replacement. The best teacher latent prediction is decoded into physical trajectory coordinates \(\hat{X}_{0, \theta^-} = V \hat{x}_{0, \theta^-}\). A fixed binary mask \(M\) replaces the temporal midpoint (\(t=30\)) and endpoint (\(t=60\)) with the actual ground-truth trajectory coordinates: $\(\hat{X}'_{0, \theta^-} = (1 - M) \odot \hat{X}_{0, \theta^-} + M \odot X_f\)$ The enhanced physical trajectory is projected back into latent space \(\hat{x}'_{0, \theta^-} = \hat{X}'_{0, \theta^-} U\). The midpoint captures the inflection vertex of complex maneuvers while the endpoint fixes overall navigational intent; injecting these two anchor points grounds teacher predictions without destroying intermediate model-predicted trajectory dynamics.

Loss & Training

The entire network is trained end-to-end. Teacher network weights \(\theta^-\) are updated as an exponential moving average (EMA) of student parameters \(\theta\), with stop-gradient (\(sg\)) applied to teacher paths. The ECTraj training objective is formulated as an adaptively weighted \(L_2\) distance in latent space: $\(\mathcal{L}_{\text{ECTraj}} = w(\sigma_t) \|\hat{x}_{0, \theta} - \hat{x}'_{0, \theta^-}\|_2^2, \quad w(\sigma_t) = \frac{1}{\sigma_t - \sigma_r}\)$ where \(w(\sigma_t)\) scales gradients inversely with variance step intervals. Optimization uses AdamW across 4 NVIDIA RTX A6000 GPUs over 60 epochs with a batch size of 16, initial learning rate of 0.002, and weight decay of 0.0001. During inference, ECTraj operates in strict single-step mode (NFE=1): sampling 6 standard Gaussian noise vectors followed by one forward pass and linear decoding yields full multi-agent joint predictions without iterative refinement or post-hoc clustering.

Key Experimental Results

Main Results

Quantitative evaluations are conducted on the Argoverse 2 Motion Forecasting Dataset (50 history frames, 60 future prediction frames). Comparisons include state-of-the-art interactive motion forecasting frameworks spanning GNN-based methods (FJMP, GNet), masked pretraining (Forecast-MAE), modern autoregressive / decoder architectures (QCNeXt, DONUT), and diffusion baselines (OptTrajDiff).

Method Year & Venue ADE6 (m) ↓ FDE6 (m) ↓ ADE1 (m) ↓ FDE1 (m) ↓ b-FDE6 (m) ↓ MR6 ↓ CR6 ↓
FJMP 2023 CVPR 0.81 1.89 1.52 4.00 2.59 0.19 0.01
GNet 2023 RA-L 0.69 1.46 1.23 3.05 2.12 0.19 0.01
QCNeXt 2023 CVPRW 0.50 1.02 0.94 2.29 1.65 0.13 0.01
Forecast-MAE 2023 ICCV 0.69 1.55 1.30 3.33 2.24 0.19 0.01
OptTrajDiff 2024 ECCV 0.60 1.31 1.08 2.71 1.95 0.17 0.01
DeMo 2024 NeurIPS 0.58 1.24 1.12 2.78 1.93 0.16 0.01
RealMotion 2024 NeurIPS 0.62 1.32 1.14 2.87 2.01 0.18 0.01
FutureNet-LOF 2025 ICRA 0.58 1.25 0.96 2.34 1.68 0.18 0.02
DONUT 2025 ICCV 0.55 1.13 1.07 2.62 1.79 0.15 0.01
ECTraj (Ours) 2026 ECCV 0.50 0.96 0.94 2.37 1.68 0.13 0.01

Inference runtime and computational overhead benchmarks demonstrate decisive efficiency gains over diffusion alternatives:

Method NFE (Steps) ↓ Compute (GFLOPs) ↓ Denoiser Latency (ms) ↓ End-to-End Latency (ms) ↓
QCNet (Prior backbone only) N/A N/A N/A 24.84
OptTrajDiff (Diffusion baseline) 10 \(\sim 1311\) 29.71 54.55
ECTraj (Ours) 1 \(\sim 132\) 2.97 27.81

Ablation Study

Ablations on Argoverse 2 systematically evaluate the individual contributions of the enhanced training pipeline (Table 4, Table 5, and Table 6):

Config / Variant ADE6 (m) ↓ FDE6 (m) ↓ ADE1 (m) ↓ FDE1 (m) ↓ b-FDE6 (m) ↓ MR6 ↓ CR6 ↓ Note
Full ECTraj Model 0.50 0.96 0.94 2.37 1.68 0.13 0.006 6-shot + teacher GT fusion + ECT schedule
w/o Multi-shot (One-shot training) 0.53 1.03 0.98 2.39 1.76 0.14 0.007 Single noise sample during training
w/o GT Fusion (Standard CM) 0.52 0.99 0.96 2.38 1.70 0.14 0.007 Standard EMA consistency without GT fusion
Full Reconstruction Objective 0.57 1.10 1.08 2.53 1.87 0.17 0.009 Teacher output completely replaced by \(X_f\)
OptTrajDiff + Best-of-K Variant 1 0.64 1.17 1.44 3.46 2.02 0.17 0.010 Select noisy mode closest to ground-truth data
OptTrajDiff + Best-of-K Variant 2 0.99 1.42 2.38 5.10 2.30 0.21 0.030 Select noisy mode closest to ground-truth noise
Replaced with ICT Scheduling 0.52 1.02 0.95 2.38 1.77 0.15 0.006 Improved Consistency Training schedule

Key Findings

  • Multi-shot consistency matching is vital for multimodal performance: Reverting training from six-shot to one-shot increases FDE6 from 0.96m to 1.03m and deteriorates Miss Rate to 0.14. Multi-shot matching enables specialized denoising trajectories for distinct maneuvers (sharp turns vs straight driving), resolving mode collapse and off-road excursions.
  • Teacher waypoint fusion strikes the optimal supervisory balance: Standard CM without GT fusion suffers from drifting teacher outputs (FDE6 0.99m), whereas complete replacement with ground truth overwhelms the single-step student network, degrading FDE6 to 1.10m. Fusing only the midpoint and endpoint provides indispensable physical anchors while retaining smooth learned dynamics.
  • Best-of-K is fundamentally incompatible with standard noise-predicting diffusion: Applying best-of-K strategies to OptTrajDiff severely degrades accuracy (Variant 2 yields FDE1 of 5.10m). Diffusion models optimize noise residuals where geometric proximity is poorly correlated; consistency models, by mapping directly to trajectory space, are uniquely suited to top-K evaluation.
  • Single-step inference outperforms multi-step sampling: Repeated noise injection across multiple steps corrupts trajectory coherence, making single-step generation (NFE=1) both the fastest and most accurate choice.

Highlights & Insights

  • First successful deployment of train-from-scratch consistency models in multi-agent trajectory prediction: Proves that consistency models can achieve state-of-the-art accuracy from scratch without pre-trained diffusion teachers, overcoming diffusion latency bottlenecks.
  • Kinematically-grounded deterministic teacher fusion: Replacing random masking with deterministic midpoint and endpoint injection provides stable, high-quality guidance tailored to vehicle motion constraints.
  • Dual utility of single-step generation: Highlights that single-step mapping is not merely an inference accelerator, but a transformative training inductive bias that enables effortless multimodal top-K loss optimization.

Limitations & Future Work

  • Dependency on pre-trained marginal priors: ECTraj still relies on pre-trained QCNet to provide anchor trajectory coordinates and scores as condition tokens, stopping short of an end-to-end prior-free framework.
  • Absence of joint probability scoring: The consistency network outputs joint geometric coordinates without evaluating joint probability confidence scores for each multi-agent configuration.
  • Vulnerability to complex U-turn maneuvers: When vehicles transition from a stationary state into sudden sharp U-turns, the model can struggle to identify the correct turning mode, a challenge shared with prior baselines.
  • vs OptTrajDiff (ECCV 2024): OptTrajDiff runs standard diffusion initialized from informed priors, requiring 10 denoising steps (29.7ms latency) and 128 sampled candidates with post-hoc clustering; ECTraj trains a consistency model from standard Gaussian noise, achieving superior FDE6 (0.96m vs 1.31m) in a single step (2.97ms latency, a 10x speedup).
  • vs QCNeXt (CVPRW 2023): QCNeXt represents a top-performing autoregressive/query-based baseline; ECTraj achieves superior FDE6 (0.96m vs 1.02m) with comparable ADE metrics, leveraging generative diversity to better capture long-tail trajectories.
  • vs TimeDiff (ICML 2023): TimeDiff employs random mask replacement in time-series diffusion; ECTraj introduces deterministic midpoint and endpoint substitution, preventing kinematic discontinuities while maximizing supervision efficiency.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Pioneers train-from-scratch consistency models in multi-agent trajectory prediction with novel waypoint teacher fusion and training-time best-of-K matching.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous comparisons on the Argoverse 2 benchmark, exhaustive computational latency analyses, and clear ablation breakdowns.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear motivation, well-structured mathematical formulation, and direct empirical validation.
  • Value: ⭐⭐⭐⭐⭐ Addresses the critical real-time barrier of diffusion models in autonomous driving, establishing an efficient generative trajectory prediction paradigm.