MEL: Coordinate-Preserving EEG Tokenization for fMRI Translation
Machine Learning
Summary
The authors address the challenge of using EEG data, which shows brain activity over short times and many channels, to predict fMRI signals that represent slower, spread-out brain responses. They propose a new method called MEL that organizes EEG data into tokens reflecting time delays, brain regions (channels), and frequency bands to better match how fMRI measures brain activity. Their method improves prediction accuracy compared to existing models without relying on bigger or more complex decoders. Tests show that the improvement is due to the way EEG data is structured rather than other factors.
Authors
Xiangyu Liu, Zeting Yan, Zhitong Yin, Boyang Li, Xi Zhang
Abstract
Translating electroencephalography (EEG) into functional magnetic resonance imaging (fMRI) is important for medical neuroimaging, clinical brain-state monitoring, and multimodal neural decoding, because it aims to infer spatially organized hemodynamic activity from fast and accessible electrophysiological recordings. Existing EEG-to-fMRI studies mainly pursue stronger decoders, but the problem is also constrained by a representation-interface mismatch: fMRI responses are delayed, temporally integrated, and spatially distributed, whereas generic EEG encodings often entangle temporal lag, channel identity, and frequency-band structure. We propose Multi-band EEG Latent-state Tokenization (MEL), a coordinate-preserving EEG representation framework that anchors each target fMRI response to its preceding EEG history and organizes it into lag-channel-frequency neural-state tokens. By explicitly capturing hemodynamic latency and spectral-spatial dynamics, MEL aligns fMRI-pertinent EEG representations with capacity-controlled readouts without depending entirely on model scaling. Experiments on VU EEG-fMRI benchmarks and external Oddball data show that MEL improves prediction over strong NeuroBOLT baselines. Ablations and controls further indicate that the gains come from structured EEG representation rather than leakage, shortcut statistics, or decoder capacity.