Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
2026-07-01 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionArtificial IntelligenceRobotics
AI summaryⓘ
The authors explain that creating good autonomous driving datasets should start by figuring out if the problem lies with the data itself or how the data is evaluated. They suggest using the simplest method to fix the problem before collecting new data, which helps save resources. They studied past driving datasets and created a step-by-step guide for designing datasets, including choosing sensors and labeling data. They also tested their ideas with their own KITScenes datasets.
autonomous drivingdataset designdata evaluationdata annotationsensor suitedata operatorsKITScenesresource allocation
Authors
Richard Schwarzkopf, Jonas Merkert, Frank Bieder, Annika Bätz, Alexander Blumberg, Carlos Fernandez, Felix Hauser, Fabian Immel, Christian Kinzig, Hendrik Königshof, Fabian Konstantinidis, Martin Lauer, Willi Poh, Nils Rack, Kevin Rösch, Yinzhe Shen, Marlon Steiner, Gleb Stepanov, Dominik Strutz, Ömer Şahin Taş, Julian Truetsch, Kaiwen Wang, Royden Wagner, Jan-Hendrik Pauls, Christoph Stiller
Abstract
Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and medium-sized labs and startups that cannot afford to misallocate scarce resources. We argue that impactful dataset creation begins with a diagnosis: whether a research question is blocked by a data problem or an evaluation problem, and proceeds by selecting the minimal data operator(s) that closes the resulting gap, recording new data only when no cheaper operator(s) suffices. We analyze the evolution of major autonomous driving (AD) datasets through this lens and distill a strategic framework spanning gap identification, operator choice, sensor suite design, and annotation strategy. We ground the framework in a running case study of our KITScenes dataset family. The datasets are available at: https://kitscenes.com/