English
Related papers

Related papers: Collar-aware Training for Streaming Speaker Change…

200 papers

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-17 Alexander Polok , Dominik Klement , Matthew Wiesner , Sanjeev Khudanpur , Jan Černocký , Lukáš Burget

This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-16 Tae Jin Park , Kyu J. Han , Jing Huang , Xiaodong He , Bowen Zhou , Panayiotis Georgiou , Shrikanth Narayanan

Automatic Speech Scoring (ASS) is the computer-assisted evaluation of a candidate's speaking proficiency in a language. ASS systems face many challenges like open grammar, variable pronunciations, and unstructured or semi-structured…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-07 Yaman Kumar Singla , Avykat Gupta , Shaurya Bagga , Changyou Chen , Balaji Krishnamurthy , Rajiv Ratn Shah

Spoken language recognition (SLR) is the task of automatically identifying the language present in a speech signal. Existing SLR models are either too computationally expensive or too large to run effectively on devices with limited…

Computation and Language · Computer Science 2023-06-06 Oriol Nieto , Zeyu Jin , Franck Dernoncourt , Justin Salamon

This paper presents a framework for the analysis of changes in visual streams: ordered sequences of images, possibly separated by significant time gaps. We propose a new approach to incorporating unlabeled data into training to generate…

Computer Vision and Pattern Recognition · Computer Science 2020-04-13 Davis Gilton , Ruotian Luo , Rebecca Willett , Greg Shakhnarovich

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

Speaker change detection (SCD) is an important task in dialog modeling. Our paper addresses the problem of text-based SCD, which differs from existing audio-based studies and is useful in various scenarios, for example, processing dialog…

Computation and Language · Computer Science 2018-10-01 Zhao Meng , Lili Mou , Zhi Jin

Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the multiple acoustic…

Sound · Computer Science 2021-04-26 Chau Luu , Peter Bell , Steve Renals

The goal of this work is to train effective representations for keyword spotting via metric learning. Most existing works address keyword spotting as a closed-set classification problem, where both target and non-target keywords are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Jaesung Huh , Minjae Lee , Heesoo Heo , Seongkyu Mun , Joon Son Chung

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki

Recent studies on learning with noisy labels have shown remarkable performance by exploiting a small clean dataset. In particular, model agnostic meta-learning-based label correction methods further improve performance by correcting noisy…

Machine Learning · Computer Science 2022-07-13 Seong Min Kye , Kwanghee Choi , Joonyoung Yi , Buru Chang

Speech tokenization serves as the foundation of speech language model (LM), enabling them to perform various tasks such as spoken language modeling, text-to-speech, speech-to-text, etc. Most speech tokenizers are trained independently of…

Computation and Language · Computer Science 2024-09-11 Arnon Turetzky , Yossi Adi

We propose a Perceiver-based sequence classifier to detect abnormalities in speech reflective of several neurological disorders. We combine this classifier with a Universal Speech Model (USM) that is trained (unsupervised) on 12 million…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-23 Hagen Soltau , Izhak Shafran , Alex Ottenwess , Joseph R. JR Duffy , Rene L. Utianski , Leland R. Barnard , John L. Stricker , Daniela Wiepert , David T. Jones , Hugo Botha

Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-03 Ismail Rasim Ulgen , Zongyang Du , Carlos Busso , Berrak Sisman

The problem of identifying voice commands has always been a challenge due to the presence of noise and variability in speed, pitch, etc. We will compare the efficacies of several neural network architectures for the speech recognition…

Machine Learning · Statistics 2020-11-25 Sanjay Krishna Gouda , Salil Kanetkar , David Harrison , Manfred K Warmuth

Although few-shot learning has attracted much attention from the fields of image and audio classification, few efforts have been made on few-shot speaker identification. In the task of few-shot learning, overfitting is a tough problem…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-26 Yanxiong Li , Wucheng Wang , Hao Chen , Wenchang Cao , Wei Li , Qianhua He

We introduce machine unlearning for speech tasks, a novel and underexplored research problem that aims to efficiently and effectively remove the influence of specific data from trained speech models without full retraining. This has…

Machine Learning · Computer Science 2025-06-03 Jiali Cheng , Hadi Amiri

Sortformer is an encoder-based speaker diarization model designed for supervising speaker tagging in speech-to-text models. Instead of relying solely on permutation invariant loss (PIL), Sortformer introduces Sort Loss to resolve the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-22 Taejin Park , Ivan Medennikov , Kunal Dhawan , Weiqing Wang , He Huang , Nithin Rao Koluguri , Krishna C. Puvvada , Jagadeesh Balam , Boris Ginsburg

The challenges in applying contrastive learning to speaker verification (SV) are that the softmax-based contrastive loss lacks discriminative power and that the hard negative pairs can easily influence learning. To overcome the first…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-14 Zhe Li , Man-Wai Mak , Helen Mei-Ling Meng

Speech separation has been extensively explored to tackle the cocktail party problem. However, these studies are still far from having enough generalization capabilities for real scenarios. In this work, we raise a common strategy named…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-26 Jing Shi , Jiaming Xu , Yusuke Fujita , Shinji Watanabe , Bo Xu