Summary
It is hard to evaluate how well text instructions match with predicted paths people take in crowds, because collecting real human movement for every scenario is costly. The authors created STRIDE, a new system that uses sociological ideas to break down complex scenarios into measurable questions, answered by tools rather than humans. This lets them check if pedestrian paths fit their text descriptions across many different settings without needing real human data. They tested STRIDE and found it agrees with human judgments about 80% of the time and revealed current models struggle with detailed context. This work takes an initial step toward fair and automated evaluation of how well text instructions align with pedestrian movement predictions.
What this means in practice
- •For autonomous vehicle teams: Evaluate how well pedestrian movement predictions align with varied traffic scenarios described in text without relying on costly real human data.
- •For simulation software developers: Automate verification that crowd movement simulations match descriptive scenarios using standardized, reproducible behavioral checks across diverse environments.
Authors
Wanchun Ni, Tao Qi, Leonel Aguilar, Jiugeng Sun, Marlene Wagner, Verena Zimmermann, Mennatallah El-Assady
Abstract
Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult. We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data. We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation.