Papers for

robotic assembly engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

General agents assemble 3D objects using only visual interaction

AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents

Abstract: The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning? To investigate this question, we introduce AssemblyWorld, an interactive 3D environment in which agents inspect rendered views and manipulate supplied rigid parts, guided by images or assembly manuals when available. Agents perceive part geometry through 2D views rather than direct access to mesh vertices or faces, while their resulting assemblies are evaluated geometrically. Building on this environment, we construct AssemblyWorldBench, comprising 100 assembly tasks across 80 objects spanning furniture, industrial assembly, and fracture reassembly. Evaluating eight agent systems reveals substantial differences in their capabilities. The strongest system achieves 80.9% part accuracy but 59.4% complete-assembly success. The evaluated open-source systems lag substantially behind their stronger closed-source peers in both execution reliability and assembly accuracy. Analyses of visual references, interaction trajectories, and failures show how agents revise assemblies while leaving residual positioning errors. AssemblyWorld provides a common setting for both assessing the capabilities of interactive assembly agents and characterizing the gap between approximate structure recovery and precise reconstruction.

Wed 30 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Putting together 3D objects from parts usually needs special training for robots or software. This work introduces AssemblyWorld, a computer environment where general-purpose agents try to build objects just by looking at pictures and moving parts, without extra special training. The researchers tested eight different agents on lots of tasks like furniture and machine parts, finding big differences in how well they assembled everything. This helps understand what current agents can do and how far they are from perfect 3D assembly by vision alone.
Open → 2609.40353v1

Visual geometry model improves robot precision in delicate assembly tasks

VGM-VS: Rethinking Visual Geometry Model for High-Precision Visual Servoing

Abstract: We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the current view and a reference image captured at the target configuration, we estimate the relative camera pose with a visual geometry model and apply it iteratively as the pose increment of a closed-loop pose-based visual servoing (PBVS) scheme. The geometry-aware representation acquired from large-scale pretraining keeps this estimate reliable when the target is occluded, weakly textured, or covers only a small part of the image. However, the scale ambiguity inherent to these models leaves the predicted translation defined up to an unknown scale, while the pose increment must be metric for robot control. We close this gap with a scene-specific metric adaptation: the robot autonomously records image--pose pairs along a predefined motion starting from the target pose, and we fine-tune the camera head on these data, jointly learning the hand--eye transform and thus removing the need for a dedicated calibration process. We evaluate our method on three real-world assembly tasks with demanding tolerances: USB-C cable picking, cable insertion, and RAM insertion. Running in real time at 30Hz, VGM-VS converges to submillimeter terminal accuracy on the cable tasks, and reaches success rates of 90--100\% when the target is moved during servoing. It converges in all trials under initial displacements of up to 30cm from the reference pose and with 50\% of the target object occluded, outperforming the compared visual servoing baselines.

Wed 23 SeptRoboticsComputer Vision and Pattern Recognition
The gist
Robots need to know exactly where their tools and targets are to do precise jobs like plugging in cables or installing computer parts. The authors developed a method that lets a robot use pictures to figure out its position more accurately, even if parts of the target are hidden or hard to see. The robot learns from its own movements to understand the real scale of what it sees, so it can adjust precisely. Their method works fast and can handle tricky situations, helping robots do delicate assembly tasks with very high accuracy.
Open → 2609.28312v1