STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts¶
Conference: NeurIPS 2026
arXiv: 2609.34799
Code: https://github.com/sweetspot00/STRIDE-Bench
Area: Human Understanding / Pedestrian Trajectory Evaluation
Keywords: text-to-trajectory, behavioral decomposition, deterministic measurements, context alignment, crowd behavior
TL;DR¶
STRIDE precompiles textual crowd contexts into behavioral questions, deterministic measurement functions, and expected ranges, then checks generated trajectories against them; its benchmark contains 936 scenarios, Text-Crowd achieves an overall score of 0.645, and humans agree with benchmark answers on 80% of pairs, although the latter result comes from a limited annotation subset.
Background & Motivation¶
Text-conditioned pedestrian trajectory generation requires more than smooth paths: movement should match the described context. Ordinary transit, a guided visit, and event viewing can imply different speeds, directional consistency, lingering fractions, and spatial distributions even on the same map. Methods such as Text-Crowd and CrowdMoGen use language to control collective motion, but conventional ADE and FDE primarily measure geometric distance from a recorded future. Collision rates, density, and distribution distances do not automatically establish whether movement follows a textual description. Multiple trajectories can reasonably satisfy the same description, so a reference path should not be treated as the only correct behavior.
The difficulty concerns reference coverage as well as metric selection. Datasets such as ETH/UCY and SDD mainly record everyday walking; collecting real crowds for every combination of map and activity is impractical. Asking a large language model (LLM) to judge coordinate sequences or aggregate statistics directly conflates which quantities matter in a context with whether those quantities meet its requirements. Salient statistics can dominate the judgment without capturing the behavior actually requested. Human evaluation provides additional evidence but is difficult to sustain across many contexts and generators.
The paper organizes evaluation using behavioral dimensions from sociology, then uses an LLM to convert a context into executable checking specifications. Language-based inference occurs when specifications are curated, rather than at every scoring run. Accurate measurements and appropriate behavioral expectations remain distinct forms of reliability. Core idea: separate text-to-trajectory alignment evaluation into auditable semantic specification and reproducible numerical verification, using context-relevant behavioral questions instead of a fixed metric list or a holistic LLM judge.
Method¶
Overall Architecture¶
Inputs include context text, a map and metadata such as event centers, goals, and crowd initialization, together with the pedestrian trajectories to evaluate. STRIDE first organizes questions with a five-axis protocol, then uses DMT specifications and TrajFacts references to select functions, parameters, and expected ranges. Once these specifications are cached, scoring executes the functions, checks their outputs against the ranges, and aggregates the results hierarchically, producing an overall score and traceable checks.
Dashed edges in the diagram indicate offline specification curation, not supervision for training a trajectory model. Benchmark mode loads cached specifications and requires no LLM at scoring time. Open mode requires an LLM to curate questions and expected answers for new text, so only numerical verification after the specifications are frozen is deterministic.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Context text and metadata"] -.-> A["Five-Axis Behavioral Protocol"]
A -.-> B["Scenario-Adaptive Questions"]
B -.-> C["Reference Calibration and<br/>Specification Caching"]
F["DMT specifications and TrajFacts"] -.-> C
C -->|Cached specifications| D["Deterministic Verification and<br/>Hierarchical Scoring"]
T["Trajectories under evaluation"] --> D
D --> O["Overall score and check-level diagnostics"]
Key Designs¶
1. Five-Axis Behavioral Protocol: restrict judgments to behavior supported by coordinates
The protocol comprises Velocity, Realism, Direction, Spatial, and Temporal. Its sociological basis comes from McPhail and Wohlstein's account of collective behavior, supplemented by work on personal space, pedestrian dynamics, and group structure. Velocity captures activity intensity, Direction captures movement orientation, Spatial captures distribution and gathering, Realism concerns basic movement plausibility, and Temporal tracks changes in these properties. These are not five labels for the same speed statistic; they encourage specification authors to examine individual, group, and environmental aspects of behavior.
The appendix explicitly excludes substantive content such as speech and signage because coordinates alone cannot recover it. This distinction matters: trajectories can support a judgment about gathering near a region, but not about understanding an activity's meaning. The authors describe the axes as a complete evaluation space, yet actual detection coverage depends on the tool library. For example, Realism currently contains only collision_fraction and lingering_fraction, which does not establish comprehensive verification of physical constraints.
2. Scenario-Adaptive Questions: express relevance through question selection rather than uniform metric weighting
An LLM receives the context and five-axis protocol, generates at least five behavioral questions, and explains the decomposition. Ordinary transit can emphasize normal speed and interpretable directional structure, whereas a guided visit calls for gathering or cohesion checks. Standards are not revised after observing the trajectory under evaluation: the appendix prompt explicitly states that no trajectory data are available during benchmark specification generation, so expectations must come from the description.
A question can bind one or more measurement functions. This provides complementary evidence for a semantic requirement while avoiding duplicated weighting merely to increase the question count. For example, speed variation can overlap between human behavior and random motion; it needs support from mean speed or path linearity to diagnose ordinary transit. Scenario adaptation primarily selects questions and their measurements. The final formula still averages questions rather than learning weights over sociological dimensions.
3. Reference Calibration and Specification Caching: translate expectations into ranges consistent with function semantics
Each measurement specification stores a function name, parameters, an expected range, justification, and a small tolerance called delta_value. The LLM consults the function's computation, output type, and TrajFacts to choose two-sided or one-sided thresholds. For example, the appendix allows 0.37 to support an expectation whose lower bound is 0.4 and tolerance is 0.03. Thresholds are precompiled language-based expectations, not answers fitted to the trajectory being evaluated.
DMT contains 20 functions: 2 for Velocity, 2 for Realism, 3 for Direction, 5 for Spatial, and 8 for Temporal. They include mean speed, speed variation, flow alignment, path linearity, local density, convergence/dispersal scores, and trends. Outputs must be interpreted according to implementation: mean_local_density counts neighbors within 1.5 times the median nearest-neighbor distance, not people per square meter; the denominator of collision_fraction is the number of active pedestrian pairs, not the number of pedestrians involved in contact; trend functions compute normalized slopes over time bins. Terminology consistency cannot replace checking units and denominators.
TrajFacts is a reference knowledge base for calibrating expectations, not measured ground truth for each of the 936 scenarios. The main text mentions event references, while the appendix says current facts primarily cover ordinary gatherings and the Lyon light festival. Its displayed excerpt also includes anti-reference values for Random and Stop. It therefore combines empirical ranges with guidance for excluding trivial baselines; not every LLM-generated interval is a direct measurement of a real crowd.
The prompt also requires discriminative questions. Random motion can produce adaptive density in a plausible range and near-zero trends, while stationary output can satisfy some upper-bound checks. The instructions therefore favor discriminative combinations and discourage using near-zero trends alone to establish stability. This makes evaluation more sensitive to trivial outputs, but introduces baseline-dependent preferences into answer curation rather than replacing independent validation.
4. Deterministic Verification and Hierarchical Scoring: identical specifications and trajectories yield identical results
With specifications fixed, DMT calculates values from trajectory coordinates and the relevant states, then checks membership in the expected-answer space. For a scenario containing several questions and measurements, Appendix A.7 defines the scenario score as:
Here \(m_i\) is the number of questions, \(n_{ij}\) the number of measurements for question \(j\), \(C_{ijp}\) the function output, and \(A_{ijp}\) its expected-answer specification. Binary checks are averaged within a question, then questions within a scenario. The benchmark is a macro average over scenarios:
This hierarchy prevents a question from automatically receiving greater total weight merely because it contains more functions. It also means bundled measurements are not a logical AND: passing one of two checks yields a question score of 0.5, not an automatic failure of the whole question. The score measures specification satisfaction, not the probability that generated trajectories share the real human distribution.
Reproducibility requires fixed trajectories, specifications, and function implementations. Benchmark scoring does not call an LLM, but regenerating open-context specifications, updating TrajFacts, or changing the threshold-generation model can change the answer space. Means, fractions, and normalized trends accommodate different agent counts and durations computationally; this does not ensure equal semantic difficulty across time windows.
A Worked Example¶
Consider the ordinary description โpedestrians walk steadily along a campus path.โ The protocol guides questions about normal speed, reasonably coherent paths, and continued movement. The specification author can select mean_speed, path_linearity, and lingering_fraction without imposing temporal trends absent from the description.
To illustrate within-question averaging, suppose one question caches two checks: mean speed in 0.8โ1.3 m/s and path linearity in 0.85โ0.97. These ranges come from ordinary-transit references in the appendix, but their combination and evaluated values here are illustrative, not a specific released benchmark item. Outputs of 1.1 m/s and 0.90 pass both checks and yield a question score of 1. If only path linearity falls outside its range, the question receives 0.5. Other questions are scored separately before aggregation into the scenario score.
Measurement requires only trajectories and stored specifications, without asking the LLM again whether the result looks plausible. A failed check is traceable to its output and range, but deterministic computation cannot repair an inappropriate question or interval chosen for that campus context.
Loss & Training¶
The paper contributes an evaluation framework and benchmark, not a new training loss for trajectory generation. Benchmark construction uses GPT-5.1 for scenario descriptions and initialization metadata, and GPT-5.2 for questions, measurements, and expected answers. Maps undergo color-based semantic segmentation, obstacle polygon approximation, and mask rasterization. Filtering addresses obstacle conflicts, density, near-duplicates, and manual plausibility checks.
The final main-text counts are 936 scenarios, 6,633 questions, and 11,696 measurements across 11 crowd categories. The abstract states that the benchmark uses 30 maps, whereas the construction section reports collecting 113 maps. The cache does not clearly explain their selection correspondence, so it does not establish that all 113 maps appear in the final benchmark.
Key Experimental Results¶
Main Results¶
The following table separates the released benchmark from the real-event subset. Benchmark values are means and standard deviations from Table 1; Lyon values preserve the more precise means in Appendix Table 6. The columns involve different scenarios and specifications, so their differences are not degradation measurements on a single shared distribution.
| Model | STRIDE-Bench, overall mean ยฑ standard deviation | Lyon, mean over 12 recordings |
|---|---|---|
| Text-Crowd | 0.645 ยฑ 0.233 | 0.4115 |
| LLM-SFM | 0.434 ยฑ 0.231 | 0.5843 |
| SingularTrajectory | 0.421 ยฑ 0.177 | Not reported in this table |
| Random Walk | 0.311 ยฑ 0.117 | 0.3175 |
| Stop | 0.221 ยฑ 0.105 | Not reported in this table |
| Human | Not applicable | 0.9427 |
Lyon validation uses 12 trajectory recordings from three camera views, with 86 newly constructed questions and 154 measurements. The main text summarizes the human mean as 0.94. This supports high scores for these real-event trajectories, not real-data validation of every one of the 936 benchmark scenarios.
| Validation | Sample scope | Result | Evidence boundary |
|---|---|---|---|
| Real trajectories | Lyon, 12 recordings, 86 questions, 154 measurements | Mean 0.9427; appendix minimum 0.812 | This event and its generated specifications |
| Humanโbenchmark agreement | 50 annotators; 56 scenarios, 350 questions; 15,750 annotations | 80%; Cohen's \(\kappa=0.730\) | Annotation subset, not the entire benchmark |
| Inter-annotator agreement | Same human study | 66%; Krippendorff's \(\alpha=0.698\) | Fine-grained behavioral disagreement remains |
| Cross-LLM validation | 4 models; 588 items; comparison with human consensus | Main text reports scores within the leave-one-out human range | Exact model values are absent from the cache and are not reconstructed |
The cross-LLM study uses deepseek-v3, claude-sonnet-4.6, qwen3.6-plus, and gpt-5.5. Its main text describes comparison with human consensus, while the Figure 4 caption describes comparison with the benchmark. This note preserves the main-text sample and comparison setup while flagging that discrepancy; the finding does not establish equivalence of all backends on new contexts.
Ablation Study¶
The paper mainly provides stratified analyses rather than component-removal ablations. Five-axis column headers in Tables 2 and 3 are missing from the cache, so only the clearly identified final Mean column is reproduced below. The paper defines Mean as the average of protocol-axis scores; it should not automatically be equated with the hierarchical overall score in Table 1.
| Partition | LLM-SFM, Mean ยฑ standard deviation | SingularTrajectory, Mean ยฑ standard deviation | Text-Crowd, Mean ยฑ standard deviation |
|---|---|---|---|
| 0โ30 s | 0.574 ยฑ 0.190 | 0.454 ยฑ 0.228 | 0.547 ยฑ 0.156 |
| 30โ120 s | 0.632 ยฑ 0.228 | 0.464 ยฑ 0.244 | 0.619 ยฑ 0.153 |
| 120 s and above | 0.453 ยฑ 0.202 | 0.424 ยฑ 0.251 | 0.463 ยฑ 0.217 |
| 1โ25 agents | 0.445 ยฑ 0.138 | 0.372 ยฑ 0.228 | 0.638 ยฑ 0.185 |
| 26โ100 agents | 0.440 ยฑ 0.189 | 0.428 ยฑ 0.252 | 0.640 ยฑ 0.191 |
| 101โ300 agents | 0.453 ยฑ 0.243 | 0.417 ยฑ 0.260 | 0.638 ยฑ 0.172 |
| 301 agents and above | 0.635 ยฑ 0.324 | 0.507 ยฑ 0.292 | 0.645 ยฑ 0.161 |
LLM-SFM leads in the two shorter windows and Text-Crowd in the longest, but the long-window Mean difference is only 0.010 and does not establish statistical significance. Text-Crowd leads the final column across agent-count partitions. These results apply to the reported adaptations and partitions, not arbitrary deployment scales.
Key Findings¶
- Generator rankings differ between the released benchmark and Lyon. Text-Crowd supports only simple polygon obstacles, and the authors simplify the Lyon map to fit it, so ranking changes also involve map interfaces and adaptation.
- SingularTrajectory receives no direct text input, but its observation history comes from LLM-SFM. It uses an 8-in/12-out protocol and autoregressive rollout up to an 8-minute cap. It is not a context-free control, and its score near LLM-SFM does not show that language conditioning is irrelevant.
- Appendix Table 7 reports
speed_variation_coeffpass rates of 1.00 for Human, Text-Crowd, and Random on Lyon, whilespread_trendis 0.00 for every column. Check-level discrimination and threshold suitability are uneven, so aggregate scores still require individual inspection. - Human annotation combines text and visualizations. Spatial and directional judgments reach greater agreement, whereas fine-grained numerical questions produce more disagreement. The 80% agreement rate does not imply equal reliability across dimensions or crowd categories.
Highlights & Insights¶
- Freeze the judge into specifications: the LLM converts descriptions into auditable checking conditions, and numerical scoring does not call it again. Repeatability no longer depends on consistency of a fresh language judgment at every run.
- Prioritize measurement semantics over intuitive names: adaptive density is not physical areal density, and a near-zero trend does not imply plausible behavior. Pseudocode, reference ranges, and anti-reference values expose otherwise hidden assumptions in answer curation.
- Question structure determines weighting: binding more measurements to a behavior does not automatically increase its total question weight. This structure can transfer to other programmatically verifiable temporal-generation tasks, provided duplicate questions do not indirectly duplicate weight.
Limitations & Future Work¶
- TrajFacts has limited coverage, and precompiled LLM ranges can still express incomplete behavioral priors. Independent real-data calibration should expand, alongside tests of ranking stability under changed thresholds, tolerances, and specification-generation backends.
- The 20 tools do not cover all social relationships, group interactions, or individual differences; Realism is especially narrow. Theoretical dimension coverage and implemented tool coverage must remain separate claims.
- Human validation covers 56 scenarios and 350 questions, while cross-model validation uses 588 items. Neither subset establishes universal reliability for the whole benchmark or open mode.
- Anti-reference guidance explicitly uses Random and Stop statistics to design ranges. This may improve discrimination against those particular baselines rather than all plausible behavior, motivating independent generators and expert review of specification-selection preferences.
- Numerical and reference discrepancies remain: the main text gives a Lyon recording range of 0.87โ1.00, but Appendix Table 6 reports a minimum of 0.812. Validation paragraphs cite Figure 7(a/b), whereas the corresponding summary caption is Figure 4 and Figure 7 presents per-dimension results. These discrepancies are retained rather than resolved by guessing.
Related Work & Insights¶
- vs ADE/FDE: conventional prediction errors compare geometric positions with a recorded future; STRIDE checks behavioral statistics implied by text. The former suits observed-trajectory forecasting, while the latter adds context alignment; neither replaces the other.
- vs TIFA / GenEval: these text-to-image evaluations decompose prompts into verifiable claims. STRIDE follows the decomposition-and-verification idea but replaces visual-model judgments with deterministic trajectory functions. Semantic decomposition remains uncertain, even though verification becomes easier to inspect.
- vs Text-Crowd / LLM-SFM: these methods generate trajectories, whereas STRIDE checks whether generation satisfies descriptions. Adaptation, map representation, and initialization are comparison conditions, so evaluation rankings are not intrinsic model orderings independent of interfaces.
- Research direction: expose failed checks together with threshold justifications to experts, distinguish genuine generation failures from inappropriate specifications, and only then use the diagnosis to improve models. Open mode particularly needs this review rather than only an aggregate score.
Rating¶
- Novelty: 4/5 โ Combines a sociological protocol, contextual questions, and deterministic tools to address an evaluation gap in text-to-pedestrian trajectories.
- Experimental Thoroughness: 3/5 โ Includes real trajectories, human validation, and stratified model analyses, but calibration coverage and specification ablations remain limited.
- Writing Quality: 3/5 โ The method is clear, but figure numbering, validation comparison wording, and numerical discrepancies affect auditability.
- Value: 4/5 โ A useful foundation for traceable behavioral evaluation, not a final correctness certificate for real crowd behavior.