English
Related papers

Related papers: Toeplitz Inverse Covariance based Robust Speaker C…

200 papers

Speaker diarization of audio streams turns out to be particularly challenging when applied to fictional films, where many characters talk in various acoustic conditions (background music, sound effects, variations in intonation...). Despite…

Multimedia · Computer Science 2018-12-19 Xavier Bost , Georges Linares

Speaker diarization (SD) is the task of answering "who spoke when" in a multi-speaker audio stream. Classically, an SD system clusters segments of speech belonging to an individual speaker's identity. Recent years have seen substantial…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-24 Nikhil Raghav

With the rise in multimedia content over the years, more variety is observed in the recording environments of audio. An audio processing system might benefit when it has a module to identify the acoustic domain at its front-end. In this…

Sound · Computer Science 2022-08-09 A Kishore Kumar , Shefali Waldekar , Md Sahidullah , Goutam Saha

We consider the problem of speaker diarization, the problem of segmenting an audio recording of a meeting into temporal segments corresponding to individual speakers. The problem is rendered particularly difficult by the fact that we are…

Methodology · Statistics 2015-03-13 Emily B. Fox , Erik B. Sudderth , Michael I. Jordan , Alan S. Willsky

This paper describes the submissions of team TalTech-IRIT-LIS to the DISPLACE 2024 challenge. Our team participated in the speaker diarization and language diarization tracks of the challenge. In the speaker diarization track, our best…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-18 Joonas Kalda , Tanel Alumäe , Martin Lebourdais , Hervé Bredin , Séverin Baroudi , Ricard Marxer

In this paper, we propose TitaNet, a novel neural network architecture for extracting speaker representations. We employ 1D depth-wise separable convolutions with Squeeze-and-Excitation (SE) layers with global context followed by channel…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-12 Nithin Rao Koluguri , Taejin Park , Boris Ginsburg

Speaker diarization is typically considered a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore the use of neural network-based generative methods for speaker diarization…

Sound · Computer Science 2024-09-20 Zhengyang Chen , Bing Han , Shuai Wang , Yidi Jiang , Yanmin Qian

The performance of speaker diarization is strongly affected by its clustering algorithm at the test stage. However, it is known that clustering algorithms are sensitive to random noises and small variations, particularly when the clustering…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-25 Meng-Zhen Li , Xiao-Lei Zhang

We present a speaker-aware approach for simulating multi-speaker conversations that captures temporal consistency and realistic turn-taking dynamics. Prior work typically models aggregate conversational statistics under an independence…

Sound · Computer Science 2026-05-25 Máté Gedeon , Péter Mihajlik

End-to-End Neural Diarization (EEND) systems produce frame-level probabilistic speaker activity estimates, yet since evaluation focuses primarily on Diarization Error Rate (DER), the reliability and calibration of these confidence scores…

Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques (e.g. ultrasound tongue imaging, lip video). Real-time MRI (rtMRI) of the vocal tract has not been used before for…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Tamás Gábor Csapó

This paper introduces an efficient and accurate pipeline for text-dependent speaker verification (TDSV), designed to address the need for high-performance biometric systems. The proposed system incorporates a Fast-Conformer-based ASR module…

Sound · Computer Science 2024-11-26 Mohammadreza Molavi , Reza Khodadadi

We propose a new speaker diarization system based on a recently introduced unsupervised clustering technique namely, generative adversarial network mixture model (GANMM). The proposed system uses x-vectors as front-end representation.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-28 Monisankha Pal , Manoj Kumar , Raghuveer Peri , Shrikanth Narayanan

The recently proposed VBx diarization method uses a Bayesian hidden Markov model to find speaker clusters in a sequence of x-vectors. In this work we perform an extensive comparison of performance of the VBx diarization with other…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-01 Federico Landini , Ján Profant , Mireia Diez , Lukáš Burget

Computational modeling of naturalistic conversations in clinical applications has seen growing interest in the past decade. An important use-case involves child-adult interactions within the autism diagnosis and intervention domain. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-30 Nithin Rao Koluguri , Manoj Kumar , So Hyun Kim , Catherine Lord , Shrikanth Narayanan

In this paper, we make the explicit connection between image segmentation methods and end-to-end diarization methods. From these insights, we propose a novel, fully end-to-end diarization model, EEND-M2F, based on the Mask2Former…

Sound · Computer Science 2024-01-24 Marc Härkönen , Samuel J. Broughton , Lahiru Samarakoon

Speaker clustering is the task of differentiating speakers in a recording. In a way, the aim is to answer "who spoke when" in audio recordings. A common method used in industry is feature extraction directly from the recording thanks to…

Sound · Computer Science 2018-03-23 Maxime Jumelle , Taqiyeddine Sakmeche

This paper aims to improve the widely used deep speaker embedding x-vector model. We propose the following improvements: (1) a hybrid neural network structure using both time delay neural network (TDNN) and long short-term memory neural…

Computation and Language · Computer Science 2019-02-22 Yun Tang , Guohong Ding , Jing Huang , Xiaodong He , Bowen Zhou

Speech recognition and other natural language tasks have long benefited from voting-based algorithms as a method to aggregate outputs from several systems to achieve a higher accuracy than any of the individual systems. Diarization, the…

Computation and Language · Computer Science 2020-02-06 Andreas Stolcke , Takuya Yoshioka

In speaker diarisation, speaker embedding extraction models often suffer from the mismatch between their training loss functions and the speaker clustering method. In this paper, we propose the method of spectral clustering-aware learning…

Sound · Computer Science 2023-03-16 Evonne P. C. Lee , Guangzhi Sun , Chao Zhang , Philip C. Woodland