Foundation model framework improves robot task planning with incomplete scene knowledge

Search, Ground, Plan: Functional Sufficiency for Task and Motion Planning under Incomplete Scene Knowledge

Robotics

Summary

When robots try to complete tasks using language and vision, they often don’t have full knowledge of their surroundings. This makes it hard to know if the task can actually be done. The authors created a system called GRAB-TAMP that carefully searches and checks objects in a scene to make sure they can fulfill the roles needed for the task before planning actions. This method improved success rates in robot tasks across different home and workshop environments, beating existing similar methods.

What this means in practice

  • For robotics engineers: Build robots that more reliably plan tasks by verifying object roles in partially known scenes.
  • For smart home device developers: Improve smart assistants’ ability to recognize and use household objects for instructions despite incomplete visual data.

Authors

Narendhiran Vijayakumar, Nav Singhal, Girish Varma, Antony Thomas

Abstract

Foundation models (FMs) have expanded task and motion planning (TAMP) to manipulation problems specified through language and visual observations. However, incomplete scene knowledge leaves a critical gap between understanding what the task requires and knowing whether the physical scene can actually realize it. We introduce GRAB-TAMP, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency. We represent the task through functional roles, relations, and assignment constraints, and incrementally inspect the scene while requirements remain unresolved, verifying candidate objects through semantic, geometric, and relational checks. We evaluate GRAB-TAMP across 32 scene variants spanning Kitchen, Living Room, and Workshop domains. Across 200 feasible trials, our approach achieves 54.0% end-to-end success with 67.3% plan goal coverage. Compared with three FM-based TAMP frameworks under the same execution setting, GRAB-TAMP improves end-to-end success by 25.7 percentage points over the mean baseline. Implementation and evaluation code: https://github.com/Narendhiranv04/GRAB-TAMP