English
Related papers

Related papers: Joint speaker diarisation and tracking in switchin…

200 papers

Over the last few years, deep learning has grown in popularity for speaker verification, identification, and diarization. Inarguably, a significant part of this success is due to the demonstrated effectiveness of their speaker…

Sound · Computer Science 2022-10-07 Yehoshua Dissen , Felix Kreuk , Joseph Keshet

In this paper, we present a novel training method for speaker change detection models. Speaker change detection is often viewed as a binary sequence labelling problem. The main challenges with this approach are the vagueness of annotated…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-17 Joonas Kalda , Tanel Alumäe

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-02 Tobias Cord-Landwehr , Christoph Boeddeker , Cătălin Zorilă , Rama Doddipatla , Reinhold Haeb-Umbach

We investigate the effect of speaker localization on the performance of speech recognition systems in a multispeaker, multichannel environment. Given the speaker location information, speech separation is performed in three stages. In the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-25 Sunit Sivasankaran , Emmaneul Vincent , Dominique Fohr

Speaker diarization, which is to find the speech segments of specific speakers, has been widely used in human-centered applications such as video conferences or human-computer interaction systems. In this paper, we propose a self-supervised…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-14 Yifan Ding , Yong Xu , Shi-Xiong Zhang , Yahuan Cong , Liqiang Wang

Multi-talker overlapped speech poses a significant challenge for speech recognition and diarization. Recent research indicated that these two tasks are inter-dependent and complementary, motivating us to explore a unified modeling method to…

Sound · Computer Science 2023-05-26 Lingwei Meng , Jiawen Kang , Mingyu Cui , Haibin Wu , Xixin Wu , Helen Meng

Existing speaker diarization systems typically rely on large amounts of manually annotated data, which is labor-intensive and difficult to obtain, especially in real-world scenarios. Additionally, language-specific constraints in these…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-13 Phat Lam , Lam Pham , Truong Nguyen , Dat Ngo , Thinh Pham , Tin Nguyen , Loi Khanh Nguyen , Alexander Schindler

Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Lianwu Chen , Yuexian Zou , Dong Yu

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-09 Yuke Lin , Ming Cheng , Ze Li , Beilong Tang , Ming Li

Due to the high performance of multi-channel speech processing, we can use the outputs from a multi-channel model as teacher labels when training a single-channel model with knowledge distillation. To the contrary, it is also known that…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-10 Shota Horiguchi , Yuki Takashima , Shinji Watanabe , Paola Garcia

Speaker diarization remains challenging due to the need for structured speaker representations, efficient modeling, and robustness to varying conditions. We propose a performant, compact diarization framework that integrates conformer…

Sound · Computer Science 2025-06-16 David Palzer , Matthew Maciejewski , Eric Fosler-Lussier

Speech applications dealing with conversations require not only recognizing the spoken words, but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of two separate…

Computation and Language · Computer Science 2019-07-12 Laurent El Shafey , Hagen Soltau , Izhak Shafran

Current methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker. Although this strategy can be enough for addressing single-speaker scenarios, it prevents accurate detection when the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-21 Juan Leon Alcazar , Fabian Caba Heilbron , Long Mai , Federico Perazzi , Joon-Young Lee , Pablo Arbelaez , Bernard Ghanem

Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras and microphones are…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Davide Berghi , Philip J. B. Jackson

Speaker diarization systems segment a conversation recording based on the speakers' identity. Such systems can misclassify the speaker of a portion of audio due to a variety of factors, such as speech pattern variation, background noise,…

Sound · Computer Science 2024-06-26 Anurag Chowdhury , Abhinav Misra , Mark C. Fuhs , Monika Woszczyna

Recent speaker diarisation systems often convert variable length speech segments into fixed-length vector representations for speaker clustering, which are known as speaker embeddings. In this paper, the content-aware speaker embeddings…

Sound · Computer Science 2021-02-15 G. Sun , D. Liu , C. Zhang , P. C. Woodland

We propose a streaming diarization method based on an end-to-end neural diarization (EEND) model, which handles flexible numbers of speakers and overlapping speech. In our previous study, the speaker-tracing buffer (STB) mechanism was…

Recognizing who is speaking in a crowded scene is a key challenge towards the understanding of the social interactions going on within. Detecting speaking status from body movement alone opens the door for the analysis of social scenes in…

Computer Vision and Pattern Recognition · Computer Science 2022-11-02 Jose Vargas-Quiros , Laura Cabrera-Quiros , Hayley Hung

Dialogue State Tracking (DST) is core research in dialogue systems and has received much attention. In addition, it is necessary to define a new problem that can deal with dialogue between users as a step toward the conversational AI that…

Computation and Language · Computer Science 2023-01-19 Hyungtak Choi , Hyeonmok Ko , Gurpreet Kaur , Lohith Ravuru , Kiranmayi Gandikota , Manisha Jhawar , Simma Dharani , Pranamya Patil

We present a speaker-aware approach for simulating multi-speaker conversations that captures temporal consistency and realistic turn-taking dynamics. Prior work typically models aggregate conversational statistics under an independence…

Sound · Computer Science 2026-05-25 Máté Gedeon , Péter Mihajlik