English

sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal Biosignals

Machine Learning 2026-02-17 v1 Signal Processing

Abstract

Tasks ranging from sleep staging to clinical diagnosis traditionally rely on standard polysomnography (PSG) devices, bedside monitors and wearable devices, which capture diverse nocturnal biosignals (e.g., EEG, EOG, ECG, SpO2_2). However, heterogeneity across devices and frequent sensor dropout pose significant challenges for unified modelling of these multimodal signals. We present \texttt{sleep2vec}, a foundation model for diverse and incomplete nocturnal biosignals that learns a shared representation via cross-modal alignment. \texttt{sleep2vec} is contrastively pre-trained on 42,249 overnight recordings spanning nine modalities using a \textit{Demography, Age, Site \& History-aware InfoNCE} objective that incorporates physiological and acquisition metadata (\textit{e.g.}, age, gender, recording site) to dynamically weight negatives and mitigate cohort-specific shortcuts. On downstream sleep staging and clinical outcome assessment, \texttt{sleep2vec} consistently outperforms strong baselines and remains robust to any subset of available modalities and sensor dropout. We further characterize, to our knowledge for the first time, scaling laws for nocturnal biosignals with respect to modality diversity and model capacity. Together, these results show that unified cross-modal alignment, coupled with principled scaling, enables label-efficient, general-purpose modelling of real-world nocturnal biosignals.

Keywords

Cite

@article{arxiv.2602.13857,
  title  = {sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal Biosignals},
  author = {Weixuan Yuan and Zengrui Jin and Yichen Wang and Donglin Xie and Ziyi Ye and Chao Zhang and Xuesong Chen},
  journal= {arXiv preprint arXiv:2602.13857},
  year   = {2026}
}
R2 v1 2026-07-01T10:37:02.362Z