RoboICL improves robot control using GPT-6 Astra with fewer mistakes

RoboICL: Embodied In-Context Learning with GPT-6 Astra

RoboticsMachine Learning

Summary

Controlling robots precisely and for long tasks using general AI models is hard because they struggle with detailed or extended actions. The authors present RoboICL, a method that helps the AI learn from examples and its own actions during a task without changing the AI itself. RoboICL stores previous steps and outcomes to improve decision-making and correct errors as they happen. This approach improves robot performance on many tasks and even helps robots do better with just a few demonstrations.

What this means in practice

  • For robotics engineers: Control robots for complex tasks with improved accuracy using a framework that integrates AI actions and past experiences during execution.
  • For automation system developers: Build systems that adapt to new manipulation tasks quickly by using AI models enhanced with example-based learning and memory of interactions.
  • For industrial robot integrators: Increase robot task success rates on factory floors by applying an AI-driven control method that reduces errors with minimal demonstrations.$Commercial implications: This method enables selling improved AI-based robot control solutions that enhance precision and efficiency in industrial automation.

Authors

Fangcheng Liu, Yeqing Shen, Anda Cheng, Weishi Mi, Chao Tang, Chenyuan Liu, Yushun Xiang, Tingguang Li, Yong-Lu Li, Yehui Tang

Abstract

General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that narrows these gaps without robot-specific parameter updates or a learned VLA. RoboICL separates \emph{demonstration context}, which provides recorded examples when available, from \emph{interaction memory}, which accumulates the model's own actions and observed outcomes. Both use a shared observation--action--receipt--observation grammar. To preserve experience across task stages, RoboICL combines sampled demonstration blocks with bounded anchored memory. Fixed anchors keep earlier rollout interactions available for in-context learning, while the latest interaction supports immediate error correction. Across 30 RoboDojo tasks, using zero shot for Open and one demonstration elsewhere, RoboICL improves on official zero-shot \gptastra{} by 20--27 progress-score points in every category. It leads the leaderboard baselines on Memory and Open, achieves comparable performance to the strongest Precision baseline, and remains competitive on Long-Horizon. Its 30-task Overall score is 50.64, versus 33.68 for the strongest baseline. On a separate ten-task subset, RoboICL scores 60.60, within 2.00 points of the $π_{0.5}$ + \gptastra{} hybrid approach. On three real-robot tasks, mean progress rises from 14.45 at zero shot to 63.33 at one shot and 78.89 at three shots. On two development tasks, optional Jev-gated action reuse reduces \gptastra{} calls by 33--48\%. Code is available at \href{https://github.com/Mosi-AI/RoboICL}{https://github.com/Mosi-AI/RoboICL}.