Visual parkour benchmark suite helps test robot locomotion in realistic scenes

The Neverwhere Visual Parkour Benchmark Suite

RoboticsComputer Vision and Pattern RecognitionMachine Learning

Summary

It is hard to test how well robots move in complicated real-world places before sending them out. The authors created a set of super-realistic 3D scenes called the Neverwhere Benchmark Suite to help test robots’ movement controllers in virtual environments. These 3D scenes combine indoor and outdoor city spots made from a method called Gaussian Splatting. The authors also found that training robots using only these types of scenes can limit their ability to work well in new places, showing the need for varied data.

What this means in practice

  • For robot developers: Test and improve robot movement controllers in hyper-realistic urban scenes before real-world deployment.
  • For video game designers: Integrate realistic urban scene reconstructions to create challenging navigation environments for characters.

Authors

Ziyu Chen, Henghui Bao, Haoran Chang, Alan Yu, Ran Choi, Kai McClennen, Gio Huh, Kevin Yang, Ri-Zhao Qiu, Yajvan Ravan, John J. Leonard, Xiaolong Wang, Phillip Isola, Ge Yang, Yue Wang

Abstract

State-of-the-art visual locomotion controllers are increasingly capable at handling complex visual environments, making evaluating their real-world performance before deployment increasingly difficult. This work intends to narrow this train/evaluation gap by developing a collection of hyper-photo-realistic, closed-loop evaluation environments - The Neverwhere Benchmark Suite - comprised of over sixty 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes. Our goal is to encourage large-scale and reproducible robot evaluation by making it easier to create and integrate Gaussian splats-based reconstructions into simulated continuous testing setups. We also underscore the potential pitfalls of relying exclusively on 3D Gaussian-generated data for training, by providing policy checkpoints trained over multiple Neverwhere scenes and their performance when evaluated in novel scenes. Our analysis illustrates the necessity of sourcing diverse data to ensure performance. Code and data are available on the project page: https://ziyc.github.io/neverwhere-bench/.