Personalizing EEG models boosts accuracy beyond general brain data gains

Separating personal from population gains when calibrating EEG foundation models for new users

Machine Learning

Summary

Calibrating brain-computer interface (BCI) models for each new user can make them work better, but sometimes improvements are just because the overall model is better, not the personal tuning. The authors tested three advanced EEG models on hundreds of new users and showed that tuning for each person really does improve performance over general models or using other people's tuning. However, the size of this personal benefit depends on how much population data was used to train the base model, and using few labels or unlabeled data isn’t always reliable for personalization. They suggest comparing personal tuning results against both the base model and tuning from other users to truly measure gains.

What this means in practice

  • For bci system developers: Improve calibration protocols by measuring real personalization gains over both population and exchanged user models to enhance EEG decoding accuracy.
  • For medical device engineers: Design EEG-based assistive technologies with informed expectations on calibration benefits depending on available population training data and label budgets.

Authors

Xilin Tao, Kani Chen

Abstract

Foundation models are increasingly adapted to individual users, but an apparent personalization gain can simply reflect a stronger population model. This distinction matters for brain-computer interfaces, where every new user must be calibrated. We evaluated personal adaptation of three frozen EEG foundation models (CBraMod, REVE and LaBraM) in 235 held-out subjects from three motor-imagery datasets, comparing each subject's adapter with the population model and with adapters fitted to other subjects. Using all first-half session labels, personal adapters improved mean balanced accuracy over the population model by 1.5-5.4 percentage points and outperformed exchanged adapters by 2.3-7.3 points in all nine model-dataset combinations. The size of this benefit depended on population training: with four times the original budget, median gains remained positive (1.0-2.0 points) but were smaller for every model, and no population model reached a confirmed plateau. Acquiring the benefit cheaply was unreliable: few-label calibration was consistently non-negative on only one dataset, and in CBraMod neither unlabeled context nor meta-learned initialization outperformed matched controls. Personalization should therefore be evaluated against both a population reference and exchanged parameters, across population-training budgets.