Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids
2026-07-22 • Robotics
RoboticsArtificial Intelligence
AI summaryⓘ
The authors address the challenge of making humanoid robots perform well in real-world tasks despite errors and changing conditions. They introduce DEED, a method that improves robot performance by carefully refining the robot's training with better data and experience-based learning. They tested this on a supermarket restocking task and found that success depends more on smart system design and data use than on new robot architectures. Their work shows that with targeted improvements, existing robot models can work reliably using modest computing resources.
Vision-Language-Action (VLA)post-trainingexperience-driven learningdistribution shiftdata curationcontrol-frequency alignmentlatent-space analysispolicy refinementhumanoid robotsfoundation model
Authors
Roger Sala Sisó, Tiago Silvério, Jakob Sand, Tran Nguyen Le
Abstract
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability. This paper presents DEED (Data-Efficient Post-Training and Experience-Driven Learning), a systems-level approach evaluated on a supermarket chip-restocking task using a Unitree G1-Edu humanoid robot and the GR00T N1.6 foundation model. DEED comprises three key components: (1) a data-efficient post-training pipeline with control-frequency alignment, data curation, task-relevant visual highlighting, and reduced VLA dependence; (2) a real-world study of experience-driven refinement, adapted from RECAP via a text-based advantage prefix and a vision-language value function; and (3) a latent-space analysis tool for studying in- and out-of-distribution behavior. Our results suggest that bridging the lab-to-store gap is primarily a systems integration challenge rather than an architectural one: careful data design and targeted post-training can transform a policy that fails under naive fine-tuning into a competent real-world system using only a single GPU.