Large audio language models struggle to infer high level human actions

Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions

Sound

Summary

This paper looks at how well large audio-language models can understand complex human activities from sounds. While these models are good at recognizing simple sound events, the authors found they cannot reliably piece these sounds together to identify higher-level actions like cleaning or preparing breakfast. They created a detailed test to check this and showed that current models fall short in combining small sound clues into broader activity understanding. All their data and tests are available for others to use.

What this means in practice

Authors

Michel Olvera, Paraskevas Stamatiadis, Changhong Wang, Ga{ë}l Richard

Abstract

Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a breakfast emerge compositionally from temporally distributed sound events, requiring abstraction beyond the event-centric granularity that dominates current training and evaluation paradigms. Our benchmark evaluates a wide set of LALMs under a principled framework that tests how language-based reasoning, grounded in acoustic perception, structures sound abstractions into higher-level understanding. By systematically varying exemplar typicality and distractor similarity, our evaluation exposes \added{that current models do not reliably perform compositional inference from atomic acoustic events to higher-level human activities solely from audio.} All data, taxonomies, and evaluation scripts are publicly available on our companion website: https://alm-sounding-actions.onrender.com/