English

CryCeleb: A Speaker Verification Dataset Based on Infant Cry Sounds

Sound 2024-03-22 v7 Artificial Intelligence Computation and Language Audio and Speech Processing

Abstract

This paper describes the Ubenwa CryCeleb dataset - a labeled collection of infant cries - and the accompanying CryCeleb 2023 task, which is a public speaker verification challenge based on cry sounds. We released more than 6 hours of manually segmented cry sounds from 786 newborns for academic use, aiming to encourage research in infant cry analysis. The inaugural public competition attracted 59 participants, 11 of whom improved the baseline performance. The top-performing system achieved a significant improvement scoring 25.8% equal error rate, which is still far from the performance of state-of-the-art adult speaker verification systems. Therefore, we believe there is room for further research on this dataset, potentially extending beyond the verification task.

Cite

@article{arxiv.2305.00969,
  title  = {CryCeleb: A Speaker Verification Dataset Based on Infant Cry Sounds},
  author = {David Budaghyan and Charles C. Onu and Arsenii Gorin and Cem Subakan and Doina Precup},
  journal= {arXiv preprint arXiv:2305.00969},
  year   = {2024}
}

Comments

ICASSP 2024

R2 v1 2026-06-28T10:22:42.311Z