Driving world model system generates realistic multi-view and LiDAR simulation

HelloWorld: Towards Practical Applications of Generative Driving World Models

Computer Vision and Pattern Recognition

Summary

Simulating driving scenes is hard because the system must create realistic views from many cameras and sensors while following control inputs over time. The authors built HelloWorld, a model that learns from lots of driving videos to generate scenes that match given car paths and map data. HelloWorld can produce images from seven cameras and LiDAR data that stay consistent over time, allowing for efficient and interactive simulation. This helps create more realistic virtual driving data beyond just replaying recorded videos.

What this means in practice

  • For autonomous vehicle developers: Generate realistic driving scenarios with controllable multi-view and LiDAR data for testing and training autonomous driving systems.
  • For robotics simulation teams: Create interactive multi-sensor driving environments that update based on user controls for robot navigation experiments.

Authors

Fan Lu, Hanshi Wang, Zijing Wang, Quan Feng, Zhi Wang, Shijie Chen, Xianming Zeng, Yujian Zhang, Jiazhe Wang, Xin Zha, Kai Wang, Zhijie Zhao, Lin Zhu, Tianyi Yang, Yucheng Xu, Tao Ji, Haodong Zhang, Zhipeng Zhang, Peixi Peng, Guang Chen, Xingliang Liu, Lei Yang, Jianyun Xu

Abstract

Driving world models provide a promising route toward scalable counterfactual data generation and interactive simulation beyond recorded driving logs. Realizing this potential requires a system that can generalize across diverse scenes, respond faithfully to prescribed controls, generate coherent multi-sensor observations, and operate efficiently under repeated inference. We present \textbf{HelloWorld}, a 2B driving world model system designed around these requirements. HelloWorld progressively specializes broad visual and motion priors from heterogeneous video data into controllable driving generation using ego pose, HD maps, and 3D boxes. A block-causal generation interface, together with adaptation to self-generated context, aligns the model with sequential simulation. The system further supports synchronized seven-camera RGB generation and conditional LiDAR synthesis, and is distilled toward few-step inference for efficient deployment. Experiments evaluate visual quality, control fidelity, cross-view consistency, robustness under repeated generation, inference efficiency, and LiDAR synthesis. Together, HelloWorld provides a unified framework for scalable driving data generation and interactive simulation.