English
Related papers

Related papers: Speaker Diarization With Lexical Information

200 papers

Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Alexandre Ducorroy , Rachid Riad

In speaker diarization, traditional clustering-based methods remain widely used in real-world applications. However, these methods struggle with the complex distribution of speaker embeddings and overlapping speech segments. To address…

Sound · Computer Science 2025-06-04 Zhaoyang Li , Jie Wang , XiaoXiao Li , Wangjie Li , Longjie Luo , Lin Li , Qingyang Hong

We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Model (VMFMM) for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-25 Tobias Cord-Landwehr , Christoph Boeddeker , Reinhold Haeb-Umbach

We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-channel training data…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Tobias Cord-Landwehr , Tobias Gburrek , Marc Deegen , Reinhold Haeb-Umbach

Recently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis. However, current models always treat overlapped speaker diarization as a multi-label classification…

Sound · Computer Science 2022-11-21 Zhihao Du , Shiliang Zhang , Siqi Zheng , Zhijie Yan

We propose a diarization system, that estimates "who spoke when" based on spatial information, to be used as a front-end of a meeting transcription system running on the signals gathered from an acoustic sensor network (ASN). Although the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-28 Tobias Gburrek , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

In traditional speaker diarization systems, a well-trained speaker model is a key component to extract representations from consecutive and partially overlapping segments in a long speech session. To be more consistent with the back-end…

Sound · Computer Science 2022-04-01 Yu-Huai Peng , Hung-Shin Lee , Pin-Tuan Huang , Hsin-Min Wang

The goal of this paper is text-independent speaker verification where utterances come from 'in the wild' videos and may contain irrelevant signal. While speaker verification is naturally a pair-wise problem, existing methods to produce the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-04 Seong Min Kye , Yoohwan Kwon , Joon Son Chung

State-of-the-art speaker verification systems are inherently dependent on some kind of human supervision as they are trained on massive amounts of labeled data. However, manually annotating utterances is slow, expensive and not scalable to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-25 Théo Lepage , Réda Dehak

Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic information in the…

Sound · Computer Science 2024-08-27 Zhaoxi Mu , Xinyu Yang , Sining Sun , Qing Yang

Identifying the identity of the speaker of short segments in human dialogue has been considered one of the most challenging problems in speech signal processing. Speaker representations of short speech segments tend to be unreliable,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-23 Tae Jin Park , Manoj Kumar , Shrikanth Narayanan

Informed speaker extraction aims to extract a target speech signal from a mixture of sources given prior knowledge about the desired speaker. Recent deep learning-based methods leverage a speaker discriminative model that maps a reference…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Mohamed Elminshawi , Wolfgang Mack , Emanuël A. P. Habets

In this paper, we present state-of-the-art diarization error rates (DERs) on multiple publicly available datasets, including AliMeeting-far, AliMeeting-near, AMI-Mix, AMI-SDM, DIHARD III, and MagicData RAMC. Leveraging EEND-TA, a single…

Sound · Computer Science 2025-09-19 Samuel J. Broughton , Lahiru Samarakoon

Information on speaker characteristics can be useful as side information in improving speaker recognition accuracy. However, such information is often private. This paper investigates how privacy-preserving learning can improve a speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Filip Granqvist , Matt Seigel , Rogier van Dalen , Áine Cahill , Stephen Shum , Matthias Paulik

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

Speaker diarization is usually referred to as the task that determines ``who spoke when'' in a recording. Until a few years ago, all competitive approaches were modular. Systems based on this framework reached state-of-the-art performance…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-15 Federico Landini

Speaker diarization, the task of segmenting an audio recording based on speaker identity, constitutes an important speech pre-processing step for several downstream applications.The conventional approach to diarization involves multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Prachi Singh , Sriram Ganapathy

Speaker diarization, usually denoted as the ''who spoke when'' task, turns out to be particularly challenging when applied to fictional films, where many characters talk in various acoustic conditions (background music, sound effects...).…

Multimedia · Computer Science 2019-04-22 Xavier Bost , Georges Linares

Smart devices serviced by large-scale AI models necessitates user data transfer to the cloud for inference. For speech applications, this means transferring private user information, e.g., speaker identity. Our paper proposes a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-26 Md Asif Jalal , Pablo Peso Parada , Jisi Zhang , Karthikeyan Saravanan , Mete Ozay , Myoungji Han , Jung In Lee , Seokyeong Jung

Despite the rapid progress of automatic speech recognition (ASR) technologies targeting normal speech in recent decades, accurate recognition of dysarthric and elderly speech remains highly challenging tasks to date. Sources of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-18 Mengzhe Geng , Xurong Xie , Zi Ye , Tianzi Wang , Guinan Li , Shujie Hu , Xunying Liu , Helen Meng
‹ Prev 1 8 9 10 Next ›