English
Related papers

Related papers: ICSD: An Open-source Dataset for Infant Cry and Sn…

200 papers

Neonatal pain assessment in clinical environments is challenging as it is discontinuous and biased. Facial/body occlusion can occur in such settings due to clinical condition, developmental delays, prone position, or other external factors.…

Computer Vision and Pattern Recognition · Computer Science 2020-03-25 Md Sirajus Salekin , Ghada Zamzmi , Rahul Paul , Dmitry Goldgof , Rangachar Kasturi , Thao Ho , Yu Sun

This paper addresses the problem of infants' cry fundamental frequency estimation. The fundamental frequency is estimated using a modified simple inverse filtering tracking (SIFT) algorithm. The performance of the modified SIFT is studied…

Sound · Computer Science 2010-09-16 Dror Lederman

Despite continuing medical advances, the rate of newborn morbidity and mortality globally remains high, with over 6 million casualties every year. The prediction of pathologies affecting newborns based on their cry is thus of significant…

Machine Learning · Computer Science 2020-03-20 Charles C. Onu , Jonathan Lebensold , William L. Hamilton , Doina Precup

Sound event detection (SED) is typically posed as a supervised learning problem requiring training data with strong temporal labels of sound events. However, the production of datasets with strong labels normally requires unaffordable labor…

Sound · Computer Science 2018-11-02 Dezhi Wang , Lilun Zhang , Changchun Bao , Kele Xu , Boqing Zhu , Qiuqiang Kong

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

As the burden of respiratory diseases continues to fall on society worldwide, this paper proposes a high-quality and reliable dataset of human sounds for studying respiratory illnesses, including pneumonia and COVID-19. It consists of…

Sound · Computer Science 2023-08-07 Truong V. Hoang , Quang H. Nguyen , Cuong Q. Nguyen , Phong X. Nguyen , Hoang D. Nguyen

Noisy labels are inevitable, even in well-annotated datasets. The detection of noisy labels is of significant importance to enhance the robustness of speaker recognition models. In this paper, we propose a novel noisy label detection…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-21 Yao Shen , Yingying Gao , Yaqian Hao , Chenguang Hu , Fulin Zhang , Junlan Feng , Shilei Zhang

We present the Noisy Ostracods, a noisy dataset for genus and species classification of crustacean ostracods with specialists' annotations. Over the 71466 specimens collected, 5.58% of them are estimated to be noisy (possibly problematic)…

Machine Learning · Computer Science 2024-12-04 Jiamian Hu , Yuanyuan Hong , Yihua Chen , He Wang , Moriaki Yasuhara

Most of the existing isolated sound event datasets comprise a small number of sound event classes, usually 10 to 15, restricted to a small domain, such as domestic and urban sound events. In this work, we introduce GISE-51, a dataset…

Sound · Computer Science 2021-10-08 Sarthak Yadav , Mary Ellen Foster

ICD coding is the international standard for capturing and reporting health conditions and diagnosis for revenue cycle management in healthcare. Manually assigning ICD codes is prone to human error due to the large code vocabulary and the…

Machine Learning · Computer Science 2021-03-16 Youngwoo Kim , Cheng Li , Bingyang Ye , Amir Tahmasebi , Javed Aslam

The issue of domain shift remains a problematic phenomenon in most real-world datasets and clinical audio is no exception. In this work, we study the nature of domain shift in a clinical database of infant cry sounds acquired across…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-04 Charles C. Onu , Hemanth K. Sheetha , Arsenii Gorin , Doina Precup

Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED,…

Sound · Computer Science 2021-02-04 Weiquan Fan , Xiangmin Xu , Xiaofen Xing , Weidong Chen , Dongyan Huang

Early detection of asthma in children is crucial to prevent long-term respiratory complications and reduce emergency interventions. This work presents an AI-powered diagnostic pipeline that leverages Googles Health Acoustic Representations…

Sound · Computer Science 2025-04-30 Abul Ehtesham , Saket Kumar , Aditi Singh , Tala Talaei Khoei

Accurate and interpretable classification of infant cry paralinguistics is essential for early detection of neonatal distress and clinical decision support. However, many existing deep learning methods rely on correlation-driven acoustic…

Sound · Computer Science 2025-12-19 Geofrey Owino , Bernard Shibwabo Kasamani , Ahmed M. Abdelmoniem , Edem Wornyo

In this paper, we introduce ASDKit, a toolkit for anomalous sound detection (ASD) task. Our aim is to facilitate ASD research by providing an open-source framework that collects and carefully evaluates various ASD methods. First, ASDKit…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-15 Takuya Fujimura , Kevin Wilkinghoff , Keisuke Imoto , Tomoki Toda

Infrared small target detection (IRSTD) is critical for applications like remote sensing and surveillance, which aims to identify small, low-contrast targets against complex backgrounds. However, existing methods often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Shuying Li , Qiang Ma , San Zhang , Chuang Yang

Automatic recognition of insect sound could help us understand changing biodiversity trends around the world -- but insect sounds are challenging to recognize even for deep learning. We present a new dataset comprised of 26399 audio files,…

Sound · Computer Science 2025-03-20 Marius Faiß , Burooj Ghani , Dan Stowell

Most sound event detection (SED) systems perform well on clean datasets but degrade significantly in noisy environments. Language-queried audio source separation (LASS) models show promise for robust SED by separating target events;…

Sound · Computer Science 2025-08-12 Yuanjian Chen , Yang Xiao , Han Yin , Yadong Guan , Xubo Liu

Accurately detecting voiced intervals in speech signals is a critical step in pitch tracking and has numerous applications. While conventional signal processing methods and deep learning algorithms have been proposed for this task, their…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-07 Yixuan Zhang , Heming Wang , DeLiang Wang