Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation

2026-07-20Robotics

Robotics
AI summary

The authors address the challenge of controlling robots to handle objects consistently over time, even when the objects move, get covered, or the robot recovers from errors. They introduce Persistent Object Tokenization (POT), a method that keeps track of 3D object details from camera data to guide and verify the robot's whole-body actions. Their system, POT-VLA, improves the success rate of real robot tasks compared to previous methods, particularly for tasks needing the robot to maintain precise 3D relationships with objects. This shows that keeping a continuous, detailed understanding of objects helps robots perform complex actions more reliably.

vision-language-action policieshumanoid loco-manipulationobject-state divergencePersistent Object Tokenization (POT)RGB-D observationswhole-body action expertgeometric predicate checksclosed-loop executionUnitree G1 robot3D object representation
Authors
Peng Ren, Haoyang Ge, Jiang Zhao, Cong Huang, Yukun Shi, Pei Chi, Kai Chen
Abstract
Vision-language-action policies are a promising foundation for general robot control, but long-horizon humanoid loco-manipulation requires the robot to treat task objects as persistent physical entities across movement, contact, occlusion, and recovery. We study this problem as object-state divergence: the object state used to condition a whole-body action can differ from the state used to decide whether the action achieved the intended physical relation. We propose \emph{Persistent Object Tokenization} (POT), which maintains role-indexed 3D object records from RGB-D observations and converts them into object tokens for a whole-body action expert. Instantiated as \emph{POT-VLA}, the same object records condition action generation and support geometric predicate checks, yielding a closed-loop execution system in which object state is both actionable and verifiable. On a Unitree G1, POT-VLA improves a matched direct GR00T-N1.7 baseline from 39/80 to 71/80 successes over eight real-world task families. In an external Being-0-aligned reference, POT-VLA achieves 44/50 successes on aligned service tasks, compared with the 37/50 success reported by the Being-0 paper. The largest gains occur on tasks requiring maintained 3D relations, suggesting that persistent object-centered state is a useful abstraction for verifiable humanoid VLA execution.