Summary
Access to real human mobility data is often limited due to privacy and commercial reasons, which makes studying city movement patterns tricky. The authors review different ways to create fake, or synthetic, mobility data that can still help understand how people move in cities. They organize these methods by what type of movement data they produce and how useful they are for analyzing urban life. They also introduce a framework to compare these methods based on how realistic, large-scale, and automated the data generation is. Their work helps city planners and analysts know which synthetic data methods might suit their needs.
What this means in practice
- •For urban planning teams: Select synthetic mobility data generation methods that balance realism and scale for city movement analysis without relying on restricted real data.
- •For transportation modelers: Use the Meaning-Population-Autonomy framework to identify synthetic data methods that produce behaviorally meaningful and feasible travel trajectories.
A survey. It maps existing work.
Authors
Yanbo Pang, Chen Zhong, Song Gao, Yoshihide Sekimoto
Abstract
Human mobility data has become an increasingly important component of urban analytics. Although the range of available mobility data sources has expanded substantially, access remains highly constrained by commercial restrictions, privacy concerns, and institutional barriers. Data protection procedures also often reduce the analytical value of released datasets. Synthetic mobility data has emerged as a promising solution, but existing methods differ substantially in their underlying mechanisms, the information they preserve, the outputs they generate, and the analytical questions they can support. Their comparative strengths and trade-offs remain insufficiently understood for urban analytics. This paper presents a structured review of synthetic human mobility data generation from an urban analytics perspective. We review the literature by methodological family and index it by the mobility outputs each family generates natively and the analytical capabilities those outputs enable. We first provide a taxonomy of synthetic data products, including population and persona representations, activity schedules, trip and tour records, trajectories, and aggregate mobility patterns. We then review the major methodological families, spanning mechanistic models, survey-driven population synthesis, activity- and agent-based simulation, deep generative models, transformer-based mobility language models, and LLM-agentic systems. Building on this synthesis, we introduce a Meaning-Population-Autonomy framework that characterises these methods along three dimensions: behavioural meaning, population grounding and scale, and generation autonomy. We consider these dimensions the principal requirements for downstream urban analytics. Few methods deliver behavioural meaning, population grounding and autonomous generation at once, and fewer still with generation constrained to feasible trajectories.