English

WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities

Signal Processing 2025-10-02 v1 Artificial Intelligence Computation and Language Machine Learning Neurons and Cognition

Abstract

Electroencephalography (EEG) interpretation using multimodal large language models (MLLMs) offers a novel approach for analyzing brain signals. However, the complex nature of brain activity introduces critical challenges: EEG signals simultaneously encode both cognitive processes and intrinsic neural states, creating a mismatch in EEG paired-data modality that hinders effective cross-modal representation learning. Through a pivot investigation, we uncover complementary relationships between these modalities. Leveraging this insight, we propose mapping EEG signals and their corresponding modalities into a unified semantic space to achieve generalized interpretation. To fully enable conversational capabilities, we further introduce WaveMind-Instruct-338k, the first cross-task EEG dataset for instruction tuning. The resulting model demonstrates robust classification accuracy while supporting flexible, open-ended conversations across four downstream tasks, thereby offering valuable insights for both neuroscience research and the development of general-purpose EEG models.

Keywords

Cite

@article{arxiv.2510.00032,
  title  = {WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities},
  author = {Ziyi Zeng and Zhenyang Cai and Yixi Cai and Xidong Wang and Junying Chen and Rongsheng Wang and Yipeng Liu and Siqi Cai and Benyou Wang and Zhiguo Zhang and Haizhou Li},
  journal= {arXiv preprint arXiv:2510.00032},
  year   = {2025}
}
R2 v1 2026-07-01T06:08:33.689Z