Hierarchical model improves word retrieval from non-invasive brain signals

HDND: Hierarchical Dynamic Neural Decoding for Multilingual Word/Character Retrieval from Non-Invasive Brain Recordings

Computer Vision and Pattern Recognition

Summary

Decoding individual words from non-invasive brain recordings is hard because the brain signals are weak and mixed with many other factors. The authors developed a new method called HDND that breaks down the decoding process into smaller steps, refining the word guess over time instead of guessing all at once. They tested HDND on brain signals recorded while people listened to, read, or spoke words in four different languages. This method consistently outperformed previous approaches in identifying the correct words from brain activity. The results suggest that using hierarchical, stepwise models can better capture the complex brain signals involved in language.

What this means in practice

  • For brain computer interface developers: Use hierarchical decoding methods to improve word-level communication from non-invasive brain recordings across multiple languages.$Commercial implications: Enables advanced brain-computer interfaces that decode words from EEG or MEG signals for communication aids and assistive devices.
  • For speech technology engineers: Incorporate structured refinement in models to enhance speech and reading-related brain signal decoding accuracy in diverse language conditions.

Authors

Yueyang Li, Shuran Chen, Wai Ting Siok, Nizhuan Wang

Abstract

While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an unresolved challenge. Decoding individual words from non-invasive brain recordings is particularly difficult, as word-level neural evidence is weak, temporally distributed, and entangled with acoustic, lexical, and semantic structure. Existing retrieval pipelines often collapse these factors into a single representation, potentially discarding information available at intermediate temporal scales. Here, we introduce Hierarchical Dynamic Neural Decoding (HDND), a hierarchical dynamic decoding framework that treats word decoding as structured refinement rather than flat label retrieval. HDND combines intermediate neural representations, contextual semantic predictions, and, for selected reading conditions, an auxiliary character-form objective. We evaluate HDND across seven electroencephalography (EEG) and magnetoencephalography (MEG) datasets spanning English, Dutch, Mandarin, and Cantonese listening, reading, and reading-aloud conditions. Across the nine-condition word-retrieval benchmark, the proposed HDND yields a higher participant-averaged balanced Top-10 point estimate than the matched contextual word-decoding baseline in every condition and achieves the highest mean among all compared methods in eight of nine conditions. Across the same nine matched conditions, HDND also yields higher token-micro and pooled word-macro Top-10 point estimates in every setting. Sentence retrieval favors HDND in eight of nine conditions, while auditory speech-segment retrieval is mixed across the six listening conditions. These results show that hierarchical residual refinement can improve multilingual word retrieval from heterogeneous non-invasive brain recordings.