RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

2026-07-20Robotics

Robotics
AI summary

The authors present RynnBrain 1.1, a set of large AI models designed to help robots understand and interact with the physical world better. These models learn to perceive objects, reason about space, and plan actions all within a shared framework, improving capabilities like predicting contact points and 3D grounding for smaller models. They also developed RynnBrain-VLA, a version that works across different robot types and was tested on several real robots. Their largest model outperformed existing ones on multiple benchmarks, and real-robot tests showed improved robotic task success when using their models compared to others. Training across multiple tasks and robots led to better overall performance than training on single tasks alone.

embodied foundation modelsspatio-temporal framework3D groundingcontact-point predictionrobot manipulationcross-embodiment action spacelocalizationrobot planningmulti-task trainingvisual-language-agent (VLA)
Authors
Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang, Jianwei Gao, Shijian Lu, Chengju Liu, Jianfei Yang, Mingxiu Chen, Deli Zhao
Abstract
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.