Large audio language models struggle to infer high level human actions
Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions
Sound
Summary
This paper looks at how well large audio-language models can understand complex human activities from sounds. While these models are good at recognizing simple sound events, the authors found they cannot reliably piece these sounds together to identify higher-level actions like cleaning or preparing breakfast. They created a detailed test to check this and showed that current models fall short in combining small sound clues into broader activity understanding. All their data and tests are available for others to use.
What this means in practice
- •For smart home device developers: Improve audio-based activity recognition systems by assessing their limits in inferring complex human actions from sound alone.
- •For human-computer interaction designers: Design better assistive technologies that understand everyday tasks through sound by knowing current models' weaknesses in compositional audio reasoning.
Authors
Michel Olvera, Paraskevas Stamatiadis, Changhong Wang, Ga{ë}l Richard
Abstract
Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a breakfast emerge compositionally from temporally distributed sound events, requiring abstraction beyond the event-centric granularity that dominates current training and evaluation paradigms. Our benchmark evaluates a wide set of LALMs under a principled framework that tests how language-based reasoning, grounded in acoustic perception, structures sound abstractions into higher-level understanding. By systematically varying exemplar typicality and distractor similarity, our evaluation exposes \added{that current models do not reliably perform compositional inference from atomic acoustic events to higher-level human activities solely from audio.} All data, taxonomies, and evaluation scripts are publicly available on our companion website: https://alm-sounding-actions.onrender.com/