Robot and human team up to paint complex scenes over time
CoBrush: A Hierarchical Planning Framework for Human-Robot Co-Painting
Robotics
Summary
Painting together with a robot is hard because the robot needs to understand what the human wants as the painting changes. Existing robot painters usually only finish one small part or sketch at a time, which makes it tough to work on a big picture step by step. The authors created CoBrush, a system that helps the robot plan in stages: it figures out the human's intent, decides where to paint, and controls each brush stroke. This way, the robot and human can build a detailed painting over multiple rounds. Tests showed CoBrush works better than simpler methods for matching what the human wants and making the painting progress smoothly.
What this means in practice
- •For robotics system developers: Develop robots that progressively co-create paintings with humans by coordinating intent understanding and brushstroke control over multiple interactions.
- •For interactive art installation creators: Build collaborative art exhibits where visitors and robots create complex paintings together across multiple sessions.
Authors
Dantong Qin, Yike Guo, Qinlin Liu, Alessandro Bozzon, Pan Wang
Abstract
Embodied co-painting requires a robot to repeatedly update a shared physical canvas while human intent evolves over interaction. Existing reference-driven painters or reactive assistants are typically optimized for single-shot rendering or sketch completion, limiting their ability to sustain coherent multi-round collaboration or to construct complex, content-rich scenes over time. We present CoBrush, a hierarchical framework that formulates multi-round co-painting as a coordinated semantic, spatial, and execution process. By separating high-level intent inference from spatial grounding and stroke-level control, the system supports progressive scene development on real acrylic canvases. We evaluate the framework through real human-robot painting sessions, stress tests, and user studies. Compared to single-turn baselines, our approach achieves stronger semantic alignment, more stable spatial progression, and higher perceived plausibility of robot actions. These results demonstrate that structured multi-stage reasoning improves the coherence and robustness of interactive painting and supports the progressive development of content-rich physical artworks.