Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

2026-08-31Computation and Language

Computation and LanguageArtificial IntelligenceComputer Vision and Pattern RecognitionMachine Learning
AI summary

The authors focus on how AI agents that use both text and visuals can lie or trick others, which is important for making AI safe and reliable. They created a 3D game environment called MineAmongUs, similar to the social game Among Us, where AI agents communicate using words and body language to deceive others. They also developed ARIA, a tool to study how different parts of these AI agents work. Their findings show that non-verbal cues like actions or gestures play a big role in successful deception. This work helps advance research on making AI agents that interact with the world in more human-like ways.

Large Language Models (LLM)Vision-Language Models (VLM)AI AlignmentSocial Deduction GamesAmong UsMultimodal CommunicationDeception TaxonomiesAgent HarnessAblation StudyEmbodied AI
Authors
Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim
Abstract
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.