English
Related papers

Related papers: DiaCorrect: End-to-end error correction for speake…

200 papers

A method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulating it as a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-29 Shota Horiguchi , Shinji Watanabe , Paola Garcia , Yuki Takashima , Yohei Kawaguchi

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved…

Speaker diarization relies on the assumption that speech segments corresponding to a particular speaker are concentrated in a specific region of the speaker space; a region which represents that speaker's identity. These identities are not…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Nikolaos Flemotomos , Panayiotis Georgiou , Shrikanth Narayanan

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Ming Cheng , Weiqing Wang , Yucong Zhang , Xiaoyi Qin , Ming Li

Automatic speech recognition (ASR) is a relevant area in multiple settings because it provides a natural communication mechanism between applications and users. ASRs often fail in environments that use language specific to particular…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-16 Rafael Viana-Cámara , Mario Campos-Soberanis , Diego Campos-Sobrino

Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric speech dataset from…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Seohyun Park , Chitralekha Gupta , Michelle Kah Yian Kwan , Xinhui Fung , Alexander Wenjun Yip , Suranga Nanayakkara

In this paper, we conduct a comparative study on speaker-attributed automatic speech recognition (SA-ASR) in the multi-party meeting scenario, a topic with increasing attention in meeting rich transcription. Specifically, three approaches…

Sound · Computer Science 2022-07-04 Fan Yu , Zhihao Du , Shiliang Zhang , Yuxiao Lin , Lei Xie

We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we develop a novel centroid refinement decoder whose usefulness…

Machine Learning · Computer Science 2025-12-18 Elio Gruttadauria , Mathieu Fontaine , Jonathan Le Roux , Slim Essid

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-17 Alexander Polok , Dominik Klement , Matthew Wiesner , Sanjeev Khudanpur , Jan Černocký , Lukáš Burget

Since the first speech recognition systems were built more than 30 years ago, improvement in voice technology has enabled applications such as smart assistants and automated customer support. However, conversation intelligence of the future…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-15 Desh Raj

Although Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and…

Sound · Computer Science 2026-02-27 Sanjid Hasan , Risalat Labib , A H M Fuad , Bayazid Hasan

Dysarthric speech reconstruction (DSR) typically employs a cascaded system that combines automatic speech recognition (ASR) and sentence-level text-to-speech (TTS) to convert dysarthric speech into normally-prosodied speech. However,…

Sound · Computer Science 2026-03-03 Minghui Wu , Haitao Tang , Jiahuan Fan , Ruizhi Liao , Yanyong Zhang

Dysarthric speech reconstruction (DSR), which aims to improve the quality of dysarthric speech, remains a challenge, not only because we need to restore the speech to be normal, but also must preserve the speaker's identity. The speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-21 Disong Wang , Songxiang Liu , Xixin Wu , Hui Lu , Lifa Sun , Xunying Liu , Helen Meng

Even state-of-the-art speaker diarization systems exhibit high variance in error rates across different datasets, representing numerous use cases and domains. Furthermore, comparing across systems requires careful application of best…

Sound · Computer Science 2025-08-07 Eduardo Pacheco , Atila Orhon , Berkin Durmus , Blaise Munyampirwa , Andrey Leonov

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-02 Tobias Cord-Landwehr , Christoph Boeddeker , Cătălin Zorilă , Rama Doddipatla , Reinhold Haeb-Umbach

Deep clustering is a recently introduced deep learning architecture that uses discriminatively trained embeddings as the basis for clustering. It was recently applied to spectrogram segmentation, resulting in impressive results on…

Machine Learning · Computer Science 2016-07-11 Yusuf Isik , Jonathan Le Roux , Zhuo Chen , Shinji Watanabe , John R. Hershey

We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Mingyue Huo , Yiwen Shao , Yuheng Zhang

Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to…

Sound · Computer Science 2020-07-30 Viet-Trung Dang , Tianyu Zhao , Sei Ueno , Hirofumi Inaguma , Tatsuya Kawahara

End-to-end transformer-based automatic speech recognition (ASR) systems often capture multiple speech traits in their learned representations that are highly entangled, leading to a lack of interpretability. In this study, we propose the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Pu Wang , Hugo Van hamme

Automatic speaker diarization techniques typically involve a two-stage processing approach where audio segments of fixed duration are converted to vector representations in the first stage. This is followed by an unsupervised clustering of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-15 Prachi Singh , Sriram Ganapathy