English
Related papers

Related papers: CryCeleb: A Speaker Verification Dataset Based on …

200 papers

Facial recognition system is one of the major successes of Artificial intelligence and has been used a lot over the last years. But, images are not the only biometric present: audio is another possible biometric that can be used as an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-26 Prateek Mishra

This paper presents the "Ethiopian" system for the SLT 2021 Children Speech Recognition Challenge. Various data processing and augmentation techniques are proposed to tackle children's speech recognition problem, especially the lack of the…

Sound · Computer Science 2020-11-10 Guoguo Chen , Xingyu Na , Yongqing Wang , Zhiyong Yan , Junbo Zhang , Sifan Ma , Yujun Wang

In this paper, we focus on Whisper, a recent automatic speech recognition model trained with a massive 680k hour labeled speech corpus recorded in diverse conditions. We first show an interesting finding that while Whisper is very robust…

Sound · Computer Science 2023-10-10 Yuan Gong , Sameer Khurana , Leonid Karlinsky , James Glass

Automatic Speech Recognition (ASR) systems struggle with child speech due to its distinct acoustic and linguistic variability and limited availability of child speech datasets, leading to high transcription error rates. While ASR error…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Natarajan Balaji Shankar , Zilai Wang , Kaiyuan Zhang , Mohan Shi , Abeer Alwan

The development of Automatic Speech Recognition (ASR) systems for low-resource African languages remains challenging due to limited transcribed speech data. While recent advances in large multilingual models like OpenAI's Whisper offer…

Computation and Language · Computer Science 2025-10-09 Benjamin Akera , Evelyn Nafula , Patrick Walukagga , Gilbert Yiga , John Quinn , Ernest Mwebaze

Creating Speaker Verification (SV) systems for classroom settings that are robust to classroom noises such as babble noise is crucial for the development of AI tools that assist educational environments. In this work, we study the efficacy…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Saba Tabatabaee , Jing Liu , Carol Espy-Wilson

We held the second installment of the VoxCeleb Speaker Recognition Challenge in conjunction with Interspeech 2020. The goal of this challenge was to assess how well current speaker recognition technology is able to diarise and recognize…

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range…

Computation and Language · Computer Science 2025-02-04 Turi Abu , Ying Shi , Thomas Fang Zheng , Dong Wang

In the U.S., approximately 15-17% of children 2-8 years of age are estimated to have at least one diagnosed mental, behavioral or developmental disorder. However, such disorders often go undiagnosed, and the ability to evaluate and treat…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-30 Jialu Li , Mark Hasegawa-Johnson , Nancy L. McElwain

Even with recent advances in speech synthesis models, the evaluation of such models is based purely on human judgement as a single naturalness score, such as the Mean Opinion Score (MOS). The score-based metric does not give any further…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-27 Kyumin Park , Keon Lee , Daeyoung Kim , Dongyeop Kang

This paper reports on the design and outcomes of the ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids. The scenario was a listener attending to a target speaker in a noisy, domestic environment. There were multiple…

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Yanir Marmor , Arad Zulti , David Krongauz , Adam Gabet , Yoad Snapir , Yair Lifshitz , Eran Segal

The 2020 Personalized Voice Trigger Challenge (PVTC2020) addresses two different research problems a unified setup: joint wake-up word detection with speaker verification on close-talking single microphone data and far-field multi-channel…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-07 Yan Jia , Xingming Wang , Xiaoyi Qin , Yinping Zhang , Xuyang Wang , Junjie Wang , Ming Li

We present an overview of the CLEF-2018 CheckThat! Lab on Automatic Identification and Verification of Political Claims, with focus on Task 1: Check-Worthiness. The task asks to predict which claims in a political debate should be…

Neonatal respiratory distress is a common condition that if left untreated, can lead to short- and long-term complications. This paper investigates the usage of digital stethoscope recorded chest sounds taken within 1min post-delivery, to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-13 Ethan Grooby , Chiranjibi Sitaula , Kenneth Tan , Lindsay Zhou , Arrabella King , Ashwin Ramanathan , Atul Malhotra , Guy A. Dumont , Faezeh Marzbanrad

The development of high-performing, robust, and reliable speech technologies depends on large, high-quality datasets. However, African languages -- including our focus, Igbo, Hausa, and Yoruba -- remain under-represented due to insufficient…

This paper addresses the problem of infants' cry fundamental frequency estimation. The fundamental frequency is estimated using a modified simple inverse filtering tracking (SIFT) algorithm. The performance of the modified SIFT is studied…

Sound · Computer Science 2010-09-16 Dror Lederman

We present our shared task on text-based emotion detection, covering more than 30 languages from seven distinct language families. These languages are predominantly low-resource and are spoken across various continents. The data instances…

This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles,…