Summary
Robots often receive instructions in natural language that describe relations like “next to” or “above,” but these are vague and don’t specify exact positions needed for precise control. The authors created a system that translates these language descriptions into detailed geometric maps for robot controllers while keeping the flexibility that the original instruction allows. Their method ensures the robot understands which parts of the task are fixed and which parts can vary, helping it perform tasks more accurately. They tested this by comparing different interfaces and running simulations with a robot arm, showing their approach preserves important freedoms in the task.
What this means in practice
- •For robotics engineers: Develop robot controllers that convert natural language spatial instructions into precise, flexible task objectives for better manipulation.
- •For motion planning developers: Integrate semantic-to-geometric compilers to create task maps that preserve task-specific freedoms and improve continuous policy composition.
Authors
Jaegyun Park, Jingwang Lee, Jungsoo Lee, Soonwoong Hwang, Wansoo Kim
Abstract
Natural-language manipulation instructions specify qualitative relations, whereas continuous controllers require state-evaluable task quantities, differentials, and completion conditions. Because a qualitative relation generally leaves part of the relative configuration unspecified, expanding it into a complete pose can introduce unintended constraints. We present a typed semantic-to-geometric interface in which language specifies entities, relations, and phases, while each relation indexes a registered specification of its task-relevant distinctions and preserved freedoms. A robot-side compiler grounds these specifications, constructs relation-specific task maps and consistent differentials using conformal geometric algebra, and composes the resulting policies through RMPflow. To evaluate the division of responsibility between the language model and the compiler, we compared a Semantic Topology interface with one that additionally requires relation-specific geometric specifications over 60 instructions. Both produced correct shared semantic content in 41/60 cases, but critical errors under their respective interface requirements occurred in 19/60 and 58/60 cases. Across 64 grounded evaluations spanning eight geometric relation forms, the task maps preserved registered null directions and responded to relation-relevant perturbations; analytic directional derivatives agreed with finite differences, and Jacobian ranks matched the registered dimensions. In three closed-loop ablations using a simulated Franka Emika Panda in MuJoCo, fixing a relation-preserved coordinate increased median terminal progress error by 20.24--71.00~mm while the retained relation errors remained within their evaluation bounds. These results support compiling relation-visible geometry and preserved freedom together into composable continuous objectives.