Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms
Abstract
Contrastive learning is a widely adopted self-supervised pretraining strategy, yet its dependence on cohort composition remains underexplored. We present Contrasting by Patient Augmented Electrocardiograms (CAPE) foundation model and pretrain on four cohorts (n = 5,203,352), from diverse populations across three continents (North America, South America, Asia). We systematically assess how cohort demographics, health status, and population diversity influence the downstream performance for prediction tasks also including two additional cohorts from another continent (Europe). We find that downstream performance depends on the distributional properties of the pretraining cohort, including demographics and health status. Moreover, while pretraining with a multi-centre, demographically diverse cohort improves in-distribution accuracy, it reduces out-of-distribution (OOD) generalisation of our contrastive approach by encoding cohort-specific artifacts. To address this, we propose the In-Distribution Batch (IDB) strategy, which preserves intra-cohort consistency during pretraining and enhances OOD robustness. This work provides important insights for developing clinically fair and generalisable foundation models.
Keywords
Cite
@article{arxiv.2509.10369,
title = {Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms},
author = {Gul Rukh Khattak and Konstantinos Patlatzoglou and Joseph Barker and Libor Pastika and Boroumand Zeidaabadi and Ahmed El-Medany and Hesham Aggour and Yixiu Liang and Antonio H. Ribeiro and Jeffrey Annis and Antonio Luiz Pinho Ribeiro and Junbo Ge and Daniel B. Kramer and Jonathan W. Waks and Evan Brittain and Nicholas Peters and Fu Siong Ng and Arunashis Sau},
journal= {arXiv preprint arXiv:2509.10369},
year = {2025}
}
Comments
Currently under review at npj Digital Medicine