English
Related papers

Related papers: Meeting Transcription Using Virtual Microphone Arr…

200 papers

The performances of the automatic speaker verification (ASV) systems degrade due to the reduction in the amount of speech used for enrollment and verification. Combining multiple systems based on different features and classifiers…

Computer Vision and Pattern Recognition · Computer Science 2019-02-01 Arnab Poddar , Md Sahidullah , Goutam Saha

High quality transcription data is crucial for training automatic speech recognition (ASR) systems. However, the existing industry-level data collection pipelines are expensive to researchers, while the quality of crowdsourced transcription…

Computation and Language · Computer Science 2023-09-27 Jian Gao , Hanbo Sun , Cheng Cao , Zheng Du

Creating abstractive summaries from meeting transcripts has proven to be challenging due to the limited amount of labeled data available for training neural network models. Moreover, Transformer-based architectures have proven to beat…

Computation and Language · Computer Science 2021-08-16 Nima Sadri , Bohan Zhang , Bihan Liu

This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings. By…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-16 Weiqing Wang , Ming Li

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal,…

We present the recent advances along with an error analysis of the IBM speaker recognition system for conversational speech. Some of the key advancements that contribute to our system include: a nearest-neighbor discriminant analysis (NDA)…

Computation and Language · Computer Science 2016-05-06 Seyed Omid Sadjadi , Jason Pelecanos , Sriram Ganapathy

Automatic speech recognition systems have accomplished remarkable improvements in transcription accuracy in recent years. On some domains, models now achieve near-human performance. However, transcription performance on oral history has not…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-19 Michael Gref , Nike Matthiesen , Christoph Schmidt , Sven Behnke , Joachim Köhler

Automatic speech recognition (ASR) has been an essential component of computer assisted language learning (CALL) and computer assisted language testing (CALT) for many years. As this technology continues to develop rapidly, it is important…

Computation and Language · Computer Science 2025-04-01 Michael McGuire

Automatic speech recognition (ASR) models are typically designed to operate on a single input data type, e.g. a single or multi-channel audio streamed from a device. This design decision assumes the primary input data source does not change…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-30 Gokce Keskin , Minhua Wu , Brian King , Harish Mallidi , Yang Gao , Jasha Droppo , Ariya Rastrow , Roland Maas

In this paper, we propose a novel end-to-end neural-network-based speaker diarization method. Unlike most existing methods, our proposed method does not have separate modules for extraction and clustering of speaker representations.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-16 Yusuke Fujita , Naoyuki Kanda , Shota Horiguchi , Kenji Nagamatsu , Shinji Watanabe

Speaker diarization is an essential step for processing multi-speaker audio. Although an end-to-end neural diarization (EEND) method achieved state-of-the-art performance, it is limited to a fixed number of speakers. In this paper, we solve…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-03 Yusuke Fujita , Shinji Watanabe , Shota Horiguchi , Yawen Xue , Jing Shi , Kenji Nagamatsu

Conventionally, Automatic Speech Recognition (ASR) systems are evaluated on their ability to correctly recognize each word contained in a speech signal. In this context, the word error rate (WER) metric is the reference for evaluating…

Computation and Language · Computer Science 2026-05-06 Thibault Bañeras Roux , Jane Wottawa , Mickael Rouvier , Teva Merlin , Richard Dufour

Sequence to Sequence models, in particular the Transformer, achieve state of the art results in Automatic Speech Recognition. Practical usage is however limited to cases where full utterance latency is acceptable. In this work we introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-25 George Sterpu , Christian Saam , Naomi Harte

This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to empower the clustering-based speaker diarization system to handle overlapped…

Sound · Computer Science 2022-02-11 Chen Shen , Yi Liu , Wenzhi Fan , Bin Wang , Shixue Wen , Yao Tian , Jun Zhang , Jingsheng Yang , Zejun Ma

Speaker diarization is the task of partitioning audio into segments according to speaker identity, answering the question of "who spoke when" in multi-speaker conversation recordings. While diarization is an essential task for many…

Sound · Computer Science 2025-10-01 Luca A. Lanzendörfer , Florian Grötschla , Cesare Blaser , Roger Wattenhofer

This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)-based dereverberation as a front end, then applies…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-25 Naohiro Tawara , Marc Delcroix , Atsushi Ando , Atsunori Ogawa

Whisper is a multitask and multilingual speech model covering 99 languages. It yields commendable automatic speech recognition (ASR) results in a subset of its covered languages, but the model still underperforms on a non-negligible number…

Computation and Language · Computer Science 2025-12-02 Thomas Palmeira Ferraz , Marcely Zanon Boito , Caroline Brun , Vassilina Nikoulina

We propose DarkStream, a streaming speech synthesis model for real-time speaker anonymization. To improve content encoding under strict latency constraints, DarkStream combines a causal waveform encoder, a short lookahead buffer, and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-08 Waris Quamer , Ricardo Gutierrez-Osuna

Speaker diarization consists of assigning speech signals to people engaged in a dialogue. An audio-visual spatiotemporal diarization model is proposed. The model is well suited for challenging scenarios that consist of several participants…

Computer Vision and Pattern Recognition · Computer Science 2018-10-15 Israel D. Gebru , Silèye Ba , Xiaofei Li , Radu Horaud

Speaker diarization(SD) is a classic task in speech processing and is crucial in multi-party scenarios such as meetings and conversations. Current mainstream speaker diarization approaches consider acoustic information only, which result in…

Computation and Language · Computer Science 2023-05-23 Luyao Cheng , Siqi Zheng , Zhang Qinglin , Hui Wang , Yafeng Chen , Qian Chen