English
Related papers

Related papers: LibriheavyMix: A 20,000-Hour Dataset for Single-Ch…

200 papers

Speaker extraction and diarization are two enabling techniques for real-world speech applications. Speaker extraction aims to extract a target speaker's voice from a speech mixture, while speaker diarization demarcates speech segments by…

Sound · Computer Science 2025-01-17 Junyi Ao , Mehmet Sinan Yıldırım , Ruijie Tao , Meng Ge , Shuai Wang , Yanmin Qian , Haizhou Li

Self-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in low-resource settings, over the past few years. Despite this, evaluations on…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-18 Séverin Baroudi , Hervé Bredin , Joseph Razik , Ricard Marxer

This paper introduces the second DIHARD challenge, the second in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variation in recording equipment, noise conditions, and conversational…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-20 Neville Ryant , Kenneth Church , Christopher Cieri , Alejandrina Cristia , Jun Du , Sriram Ganapathy , Mark Liberman

ASR systems often struggle with maintaining syntactic and semantic accuracy in long audio transcripts, impacting tasks like Named Entity Recognition (NER), capitalization, and punctuation. We propose a novel approach that enhances ASR by…

Computation and Language · Computer Science 2025-08-20 Duygu Altinok

Recent progress in separating the speech signals from multiple overlapping speakers using a single audio channel has brought us closer to solving the cocktail party problem. However, most studies in this area use a constrained problem…

Speaker diarization (SD) struggles in real-world scenarios due to dynamic environments and unknown speaker counts. SD is rarely used alone and is often paired with automatic speech recognition (ASR), but non-modular methods that jointly…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Yu-Wen Chen , William Ho , Maxim Topaz , Julia Hirschberg , Zoran Kostic

In recent years, wsj0-2mix has become the reference dataset for single-channel speech separation. Most deep learning-based speech separation models today are benchmarked on it. However, recent studies have shown important performance drops…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-25 Joris Cosentino , Manuel Pariente , Samuele Cornell , Antoine Deleforge , Emmanuel Vincent

The CHiME-7 and 8 distant speech recognition (DASR) challenges focus on multi-channel, generalizable, joint automatic speech recognition (ASR) and diarization of conversational speech. With participation from 9 teams submitting 32 diverse…

Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper, to better address these tasks, we first introduce speaker…

Sound · Computer Science 2023-12-19 Peng Shen , Xugang Lu , Hisashi Kawai

Speaker diarization is the task of answering Who spoke and when? in an audio stream. Pipeline systems rely on speech segmentation to extract speakers' segments and achieve robust speaker diarization. This paper proposes a common framework…

Sound · Computer Science 2023-06-08 Théo Mariotte , Anthony Larcher , Silvio Montrésor , Jean-Hugh Thomas

Multi-speaker automatic speech recognition (MASR) aims to predict ''who spoke when and what'' from multi-speaker speech, a key technology for multi-party dialogue understanding. However, most existing approaches decouple temporal modeling…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Yifan Hu , Peiji Yang , Zhisheng Wang , Yicheng Zhong , Rui Liu

We present a unified network for voice separation of an unknown number of speakers. The proposed approach is composed of several separation heads optimized together with a speaker classification branch. The separation is carried out in the…

Sound · Computer Science 2020-11-05 Shlomo E. Chazan , Lior Wolf , Eliya Nachmani , Yossi Adi

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both single-channel and multichannel conditions. However, severe performance degradation is still observed in the…

Speech applications dealing with conversations require not only recognizing the spoken words, but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of two separate…

Computation and Language · Computer Science 2019-07-12 Laurent El Shafey , Hagen Soltau , Izhak Shafran

Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, publicly available…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-13 Anfeng Xu , Tiantian Feng , Helen Tager-Flusberg , Catherine Lord , Shrikanth Narayanan

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-07 Naijun Zheng , Na Li , Xixin Wu , Lingwei Meng , Jiawen Kang , Haibin Wu , Chao Weng , Dan Su , Helen Meng

Self-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diarisation and speech recognition using a single model, a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-11 Xianrui Zheng , Chao Zhang , Philip C. Woodland

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

Sound · Computer Science 2020-01-03 Rongzhi Gu , Yuexian Zou

We propose a diarization system, that estimates "who spoke when" based on spatial information, to be used as a front-end of a meeting transcription system running on the signals gathered from an acoustic sensor network (ASN). Although the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-28 Tobias Gburrek , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-17 Alexander Polok , Dominik Klement , Matthew Wiesner , Sanjeev Khudanpur , Jan Černocký , Lukáš Burget