English
Related papers

Related papers: Data Efficient Child-Adult Speaker Diarization wit…

200 papers

Joint automatic speech recognition (ASR) and speaker diarization aim to answer the question "who spoke what" in multi-speaker scenarios. In this paper, we present an end-to-end speech large language model (Speech-LLM) for Joint strEamable…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-21 Mohan Shi , Xiong Xiao , Ruchao Fan , Shaoshi Ling , Jinyu Li

Speech technology systems struggle with many downstream tasks for child speech due to small training corpora and the difficulties that child speech pose. We apply a novel dataset, SpeechMaturity, to state-of-the-art transformer models to…

Computation and Language · Computer Science 2025-06-11 Theo Zhang , Madurya Suresh , Anne S. Warlaumont , Kasia Hitczenko , Alejandrina Cristia , Margaret Cychosz

This paper presents a computationally efficient and distributed speaker diarization framework for networked IoT-style audio devices. The work proposes a Federated Learning model which can identify the participants in a conversation without…

Sound · Computer Science 2024-12-02 Amit Kumar Bhuyan , Hrishikesh Dutta , Subir Biswas

Automatic Speech Recognition (ASR) systems often struggle with transcribing child speech due to the lack of large child speech datasets required to accurately train child-friendly ASR models. However, there are huge amounts of annotated…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-26 Rishabh Jain , Andrei Barcovschi , Mariam Yiwere , Peter Corcoran , Horia Cucu

Speaker diarization (SD) struggles in real-world scenarios due to dynamic environments and unknown speaker counts. SD is rarely used alone and is often paired with automatic speech recognition (ASR), but non-modular methods that jointly…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Yu-Wen Chen , William Ho , Maxim Topaz , Julia Hirschberg , Zoran Kostic

Children's speech recognition is considered a low-resource task mainly due to the lack of publicly available data. There are several reasons for such data scarcity, including expensive data collection and annotation processes, and data…

Computation and Language · Computer Science 2024-06-25 Vrunda N. Sukhadia , Shammur Absar Chowdhury

Speech applications dealing with conversations require not only recognizing the spoken words, but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of two separate…

Computation and Language · Computer Science 2019-07-12 Laurent El Shafey , Hagen Soltau , Izhak Shafran

Longform audio recordings obtained with microphones worn by children-also known as child-centered daylong recordings-have become a standard method for studying children's language experiences and their impact on subsequent language…

Sound · Computer Science 2025-06-16 Daniil Kocharov , Okko Räsänen

The performance of Artificial Intelligence (AI) systems fundamentally depends on high-quality training data. However, low-resource languages like Arabic suffer from severe data scarcity. Moreover, the absence of child-specific speech…

Computation and Language · Computer Science 2025-10-28 Mouhand Alkadri , Dania Desouki , Khloud Al Jallad

Speech recognition (ASR) and speaker diarization (SD) models have traditionally been trained separately to produce rich conversation transcripts with speaker labels. Recent advances have shown that joint ASR and SD models can learn to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-06 Huanru Henry Mao , Shuyang Li , Julian McAuley , Garrison Cottrell

Speaker diarization is the task of partitioning audio into segments according to speaker identity, answering the question of "who spoke when" in multi-speaker conversation recordings. While diarization is an essential task for many…

Sound · Computer Science 2025-10-01 Luca A. Lanzendörfer , Florian Grötschla , Cesare Blaser , Roger Wattenhofer

The goal of this paper is speaker diarisation of videos collected 'in the wild'. We make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker…

Sound · Computer Science 2021-08-17 Joon Son Chung , Jaesung Huh , Arsha Nagrani , Triantafyllos Afouras , Andrew Zisserman

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech…

In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning models. To address these two problems, firstly, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-01 Shilong Wu

Whispering is a distinct form of speech known for its soft, breathy, and hushed characteristics, often used for private communication. The acoustic characteristics of whispered speech differ substantially from normally phonated speech and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-08 Zhaofeng Lin , Tanvina Patel , Odette Scharenborg

We proposed a novel machine learning framework to conduct real-time multi-speaker diarization and recognition without prior registration and pretraining in a fully online learning setting. Our contributions are two-fold. First, we proposed…

Machine Learning · Computer Science 2021-12-28 Baihan Lin , Xinxin Zhang

Automatic text-based diacritic restoration models generally have high diacritic error rates when applied to speech transcripts as a result of domain and style shifts in spoken language. In this work, we explore the possibility of improving…

Computation and Language · Computer Science 2024-04-09 Sara Shatnawi , Sawsan Alqahtani , Hanan Aldarmaki

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-09 Yuke Lin , Ming Cheng , Ze Li , Beilong Tang , Ming Li

This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-16 Tae Jin Park , Kyu J. Han , Jing Huang , Xiaodong He , Bowen Zhou , Panayiotis Georgiou , Shrikanth Narayanan

Existing conversational datasets consist either of written proxies for dialog or small-scale transcriptions of natural speech. We introduce 'Interview': a large-scale (105K conversations) media dialog dataset collected from news interview…

Computation and Language · Computer Science 2020-04-08 Bodhisattwa Prasad Majumder , Shuyang Li , Jianmo Ni , Julian McAuley