Visual symbolic agent improves language guided actions in virtual humans

A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans

Artificial IntelligenceGraphics

Summary

Virtual humans need to see, think, and act based on language instructions in a 3D world. This paper presents A.D.A.M.O., a system that combines visual input and symbolic understanding to help virtual humans perform tasks from natural language prompts. The authors show A.D.A.M.O. can better understand tasks by using labeled visual information, which reduces confusion but shifts some errors to action execution rather than reasoning. They also introduce a way to test these tasks with increasing complexity.

What this means in practice

Authors

Alessandro Emmanuel Pecora, Stefano Calzolari, Francesco Strada, Andrea Bottino

Abstract

Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations), a visual-symbolic framework for language-driven vh that leverages a pretrained vlm with tool calling to unify perception, reasoning, and action within a single control loop. A.D.A.M.O. maintains a dual visual-symbolic world model that combines egocentric visual input and synchronized symbolic state to support grounded task-oriented behavior from natural language prompts. To support diagnostic evaluation, we introduce a controlled task suite organized by a cd taxonomy that breaks down spatial tasks into procedural and linguistic complexity. Experiments in controlled scenes show that semantic labeling strongly influences task completion and failure modes, reducing perceptual ambiguity while shifting failures toward downstream execution, whereas reasoning errors remain comparatively rare.