Extended Field of View Analysis for VideoGAN-based Trajectory Generation
2026-08-03 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionMachine Learning
AI summaryⓘ
The authors developed a method using generative adversarial networks (GANs) to create realistic and varied traffic movement patterns from a bird's-eye view. They improved the way traffic scenes are represented, used a graph-based method to track vehicle paths, and tested their approach on larger areas. They also created a system to measure how well the generated videos avoid unrealistic changes and keep objects consistent over time. Their experiments show that their method can efficiently produce believable traffic scenarios useful for tasks like predicting and planning in self-driving cars.
Generative Adversarial Network (GAN)Trajectory GenerationBird's-eye-viewTraffic SimulationGraph-based AssociationObject PermanenceAutomated DrivingPrediction and PlanningVideo GenerationField of View
Authors
Annajoyce Mariani, Kira Maag, Hanno Gottschalk
Abstract
Realistic and diverse trajectory generation is central to enabling higher levels of vehicle automation. While rule-based and classical learning-based methods may struggle to capture the complexity of traffic behavior, generative models have already demonstrated in other fields that they can handle a comparable level of complexity. In this paper, we build upon previous work on generative adversarial network (GAN)-based semantic bird's-eye-view traffic generation and extend the proposed framework in several key aspects. We improve the semantic representation, replace the trajectory extraction procedure with a graph-based association method, and systematically investigate increasingly larger fields of view. In addition, we introduce a quantitative evaluation framework to assess hallucinations and object permanence in generated videos. Our experiments demonstrate that the framework generalizes to larger and more complex traffic scenes while maintaining statistically realistic trajectories and coherent spatial relationships between traffic participants. Within 150GPU hours of training and with inference times below 20ms for scenes of up to 20s, our results demonstrate that video-based GANs remain an efficient and scalable approach for realistic trajectory generation, even in substantially larger traffic scenes, making them well suited for downstream tasks such as prediction, planning, and simulation in automated driving.