English

HeAR -- Health Acoustic Representations

Machine Learning 2024-03-06 v1 Artificial Intelligence

Abstract

Health acoustic sounds such as coughs and breaths are known to contain useful health signals with significant potential for monitoring health and disease, yet are underexplored in the medical machine learning community. The existing deep learning systems for health acoustics are often narrowly trained and evaluated on a single task, which is limited by data and may hinder generalization to other tasks. To mitigate these gaps, we develop HeAR, a scalable self-supervised learning-based deep learning system using masked autoencoders trained on a large dataset of 313 million two-second long audio clips. Through linear probes, we establish HeAR as a state-of-the-art health audio embedding model on a benchmark of 33 health acoustic tasks across 6 datasets. By introducing this work, we hope to enable and accelerate further health acoustics research.

Keywords

Cite

@article{arxiv.2403.02522,
  title  = {HeAR -- Health Acoustic Representations},
  author = {Sebastien Baur and Zaid Nabulsi and Wei-Hung Weng and Jake Garrison and Louis Blankemeier and Sam Fishman and Christina Chen and Sujay Kakarmath and Minyoi Maimbolwa and Nsala Sanjase and Brian Shuma and Yossi Matias and Greg S. Corrado and Shwetak Patel and Shravya Shetty and Shruthi Prabhakara and Monde Muyoyeta and Diego Ardila},
  journal= {arXiv preprint arXiv:2403.02522},
  year   = {2024}
}

Comments

4 tables, 4 figures, 6 supplementary tables, 3 supplementary figures

R2 v1 2026-06-28T15:09:08.007Z