General agents assemble 3D objects using only visual interaction

AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents

Computer Vision and Pattern RecognitionRobotics

Summary

Putting together 3D objects from parts usually needs special training for robots or software. This work introduces AssemblyWorld, a computer environment where general-purpose agents try to build objects just by looking at pictures and moving parts, without extra special training. The researchers tested eight different agents on lots of tasks like furniture and machine parts, finding big differences in how well they assembled everything. This helps understand what current agents can do and how far they are from perfect 3D assembly by vision alone.

What this means in practice

Authors

Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, Suhas Lohit, Bernhard Egger, Tim K. Marks, Anoop Cherian, Stephen Gould

Abstract

The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning? To investigate this question, we introduce AssemblyWorld, an interactive 3D environment in which agents inspect rendered views and manipulate supplied rigid parts, guided by images or assembly manuals when available. Agents perceive part geometry through 2D views rather than direct access to mesh vertices or faces, while their resulting assemblies are evaluated geometrically. Building on this environment, we construct AssemblyWorldBench, comprising 100 assembly tasks across 80 objects spanning furniture, industrial assembly, and fracture reassembly. Evaluating eight agent systems reveals substantial differences in their capabilities. The strongest system achieves 80.9% part accuracy but 59.4% complete-assembly success. The evaluated open-source systems lag substantially behind their stronger closed-source peers in both execution reliability and assembly accuracy. Analyses of visual references, interaction trajectories, and failures show how agents revise assemblies while leaving residual positioning errors. AssemblyWorld provides a common setting for both assessing the capabilities of interactive assembly agents and characterizing the gap between approximate structure recovery and precise reconstruction.