English

MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification

Sound 2026-03-04 v1 Artificial Intelligence Audio and Speech Processing

Abstract

We propose MEBM-Phoneme, a multi-scale enhanced neural decoder for phoneme classification from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Phoneme integrates a short-term multi-scale convolutional module to augment the native mid-term encoder, with fused representations via depthwise separable convolution for efficient cross-scale integration. A convolutional attention layer dynamically weights temporal dependencies to refine feature aggregation. To address class imbalance and session-specific distributional shifts, we introduce a stacking-based local validation set alongside weighted cross-entropy loss and random temporal augmentation. Comprehensive evaluations on LibriBrain Competition 2025 Track2 demonstrate robust generalization, achieving competitive phoneme decoding accuracy on the validation and official test leaderboard. These results underscore the value of hierarchical temporal modeling and training stabilization for advancing MEG-based speech perception analysis.

Keywords

Cite

@article{arxiv.2603.02254,
  title  = {MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification},
  author = {Liang Jinghua and Zhang Zifeng and Li Songyi and Zheng Linze},
  journal= {arXiv preprint arXiv:2603.02254},
  year   = {2026}
}

Comments

5 pages, 1 figure. To appear in the PNPL Competition Workshop at NeurIPS 2025

R2 v1 2026-07-01T10:59:50.047Z