Masking strategies improve EEG model training and efficiency
What masking geometry works best for EEG foundation models?
Machine Learning
Summary
Training brain signal decoding models involves hiding parts of the data and asking the model to guess them, but no one had tested which way of hiding data works best. The authors tried different ways of masking brain wave data while training models and found some masking patterns work better than others. They also discovered a new problem specific to one training method that standard checks miss. Using the best masking approach, their training method matched top performance while using much less computing power.
What this means in practice
- •For clinical neuroscientists: Train efficient EEG models for diagnosing neurological conditions using optimized masking strategies that reduce compute costs.
- •For cognitive neuroscience teams: Build scalable models for analyzing cognitive EEG data with improved training stability and performance.
Authors
Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi, Arnaud Delorme, Thomas Moreau, Michael Tangermann
Abstract
EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it determines what the network must predict and from which context. Yet it has never been ablated in isolation, as each new model bundles a new masking strategy with a new backbone and objective. In this paper, we formalize the design choices for spatio-temporal masking strategies and train various models with a single pipeline under varying masking configurations across two SSL frameworks (MAE and JEPA). We then systematically evaluate the resulting 58 pre-trained models on the 12 datasets of OpenEEGBench under a linear probe. Both frameworks agree on an optimal masking configuration and on shared failure modes. Outside these, performance is robust: 11 MAE and 9 JEPA configurations are statistically indistinguishable from the best. We further identify a novel JEPA-specific failure mode, tagged bias-inflation collapse, invisible to standard detectors. With a well-chosen mask, our pipeline reaches REVE-level downstream performance at a fraction of REVE's pre-training compute.