Multimodal dataset improves attention tracking with brain and eye signals
MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus
Human-Computer InteractionMachine LearningSound
Summary
People use where they look, head movements, and seeing cues to focus on someone talking in noisy places, but most studies only use brain waves to follow attention. The authors created a new dataset called MAESTRO that records brain activity, eye gaze, pupil size, video from a wearer’s perspective, and head motion all at once with multiple people talking. They found that combining these different signals helps track who a person is listening to better than just using brain waves alone. This dataset and code are publicly available for others to build better listening technology.
What this means in practice
- •For hearing aid designers: Improve devices that can detect which speaker a user is focusing on in noisy environments by using combined brain, eye, and head movement signals.$Commercial implications: Enables advanced hearing aids that better isolate attended speech, increasing user comprehension in complex auditory scenes.
- •For human-computer interaction developers: Build systems that adapt in real time to a user's intended speaker by leveraging combined physiological and behavioral data from the MAESTRO dataset.
Authors
K M Naimul Hassan, Ali Alavi, Donald S. Williamson
Abstract
Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise across multiple signal-to-noise ratio (SNR) conditions, enabling attention decoding under realistic listening scenarios. Through a four-speaker attention decoding benchmark, we show that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding. These findings open the door to new applications, analyses, and methodological advances in multimodal AAD. The complete dataset is publicly available at https://huggingface.co/datasets/aspire-osu/maestro-eeg-dataset . The official code repository is available at https://github.com/ASPIRE-OSU/MAESTRO .