Travel survey data generation improved with AI enhanced Bayesian networks
LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation
Artificial Intelligence
Summary
Gathering travel survey data is expensive and slow, so creating realistic fake data from small samples can save time and money. The authors found that using a combination of AI language models and a type of statistical model called Bayesian networks helps to overcome problems when only limited data is available. Their method improves how well the generated data reflects real people's travel habits and characteristics. This approach was tested on data from Hong Kong and showed clear improvements in accuracy compared to older methods.
Bayesian networksLarge language modelsSynthetic dataFew-shot learningTraveler personasTravel behavior analysisJensen-Shannon divergenceCramer's VData generationSurvey sampling
Authors
Zijian Shen, Bin Zhou, Jiguang Wang, Ya Zhao, Jintao Ke
Abstract
Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting large-scale representative samples is costly and time-consuming. A practical alternative is to generate synthetic survey records from a few-shot sample. However, such samples provide incomplete coverage of heterogeneous traveler groups and insufficient evidence for recovering the complex dependencies between demographic characteristics and travel behavior. Existing approaches have complementary limitations. Probabilistic generative models such as Bayesian networks (BNs) offer explicit distributional control, but structures learned from few-shot samples may omit meaningful dependencies or retain spurious ones. Large language models (LLMs) can help address these difficulties in BN structure learning by providing behavioral knowledge that complements the limited statistical evidence. We therefore propose LEBGen, an LLM-enhanced BN framework that uses this knowledge to refine network structure for few-shot travel survey data generation. Specifically, the LLM first identifies traveler personas from demographic attribute and travel behavior statistics, then recovers dependencies missed by the persona-augmented BN structure and prune spurious ones. The refined BN is parameterized exclusively from the observed data to generate synthetic records. Under a 2% few-shot setting on the 2022 Hong Kong Travel Characteristics Survey, LEBGen reduces the mean marginal Jensen-Shannon divergence from 0.0671 to 0.0091 and the mean absolute Cramer's V error by 14.3% over the best-performing baseline, substantially improving both distributional and dependency fidelity.