English
Related papers

Related papers: Training Dynamics-Aware Multi-Factor Curriculum Le…

200 papers

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

Multimedia · Computer Science 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

Speech separation has been extensively explored to tackle the cocktail party problem. However, these studies are still far from having enough generalization capabilities for real scenarios. In this work, we raise a common strategy named…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-26 Jing Shi , Jiaming Xu , Yusuke Fujita , Shinji Watanabe , Bo Xu

The common target speech separation directly estimate the target source, ignoring the interrelationship between different speakers at each frame. We propose a multiple-target speech separation model (MTSS) to simultaneously extract each…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Bang Zeng , Hongbing Suo , Yulong Wan , Ming Li

Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain underexplored, as…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 The Hieu Pham , Phuong Thanh Tran Nguyen , Xuan Tho Nguyen , Tan Dat Nguyen , Duc Dung Nguyen

Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Taewoo Kim , Guisik Kim , Choongsang Cho , Young Han Lee

Personalized speech enhancement (PSE) is a real-time SE approach utilizing a speaker embedding of a target person to remove background noise, reverberation, and interfering voices. To deploy a PSE model for full duplex communications, the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-29 Sefik Emre Eskimez , Takuya Yoshioka , Alex Ju , Min Tang , Tanel Parnamaa , Huaming Wang

The goal of this paper is to train effective self-supervised speaker representations without identity labels. We propose two curriculum learning strategies within a self-supervised learning framework. The first strategy aims to gradually…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-15 Hee-Soo Heo , Jee-weon Jung , Jingu Kang , Youngki Kwon , You Jin Kim , Bong-Jin Lee , Joon Son Chung

Deep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms when training data is insufficient, data augmentation is a…

Sound · Computer Science 2024-09-17 Junjie Li , Ke Zhang , Shuai Wang , Haizhou Li , Man-Wai Mak , Kong Aik Lee

Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE methods encode visual and audio features simultaneously,…

Sound · Computer Science 2025-11-13 Zixuan Li , Xueliang Zhang , Lei Miao , Zhipeng Yan , Ying Sun , Chong Zhu

Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is largely caused by the lack of generalizability and…

Sound · Computer Science 2025-05-29 Zhenghai You , Zhenyu Zhou , Lantian Li , Dong Wang

Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system performance. Although this problem can be solved by carefully…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-07 Yexin Yang , Shuai Wang , Xun Gong , Yanmin Qian , Kai Yu

Leveraging additional speaker information to facilitate speech separation has received increasing attention in recent years. Recent research includes extracting target speech by using the target speaker's voice snippet and jointly…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-21 Cong Han , Yi Luo , Chenda Li , Tianyan Zhou , Keisuke Kinoshita , Shinji Watanabe , Marc Delcroix , Hakan Erdogan , John R. Hershey , Nima Mesgarani , Zhuo Chen

Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic information in the…

Sound · Computer Science 2024-08-27 Zhaoxi Mu , Xinyu Yang , Sining Sun , Qing Yang

The diversity of speaker profiles in multi-speaker TTS systems is a crucial aspect of its performance, as it measures how many different speaker profiles TTS systems could possibly synthesize. However, this important aspect is often…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-08 Jie Pu , Yixiong Meng , Oguz Elibol

Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchronization between…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-11 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

The past decade has witnessed substantial growth of data-driven speech enhancement (SE) techniques thanks to deep learning. While existing approaches have shown impressive performance in some common datasets, most of them are designed only…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-19 Wangyou Zhang , Kohei Saijo , Zhong-Qiu Wang , Shinji Watanabe , Yanmin Qian

Conventional speech enhancement (SE) aims to improve speech perception and intelligibility by suppressing noise without requiring enrollment speech as reference, whereas personalized SE (PSE) addresses the cocktail party problem by…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-20 Ziling Huang , Haixin Guan , Yanhua Long

Speaker recognition, recognizing speaker identities based on voice alone, enables important downstream applications, such as personalization and authentication. Learning speaker representations, in the context of supervised learning,…

Machine Learning · Computer Science 2022-07-13 Metehan Cekic , Ruirui Li , Zeya Chen , Yuguang Yang , Andreas Stolcke , Upamanyu Madhow

State-of-the-art speaker diarization systems utilize knowledge from external data, in the form of a pre-trained distance metric, to effectively determine relative speaker identities to unseen data. However, much of recent focus has been on…

Machine Learning · Statistics 2018-11-02 Vivek Sivaraman Narayanaswamy , Jayaraman J. Thiagarajan , Huan Song , Andreas Spanias

Audio-Visual Target Speaker Extraction (AVTSE) aims to separate a target speaker's voice from a mixed audio signal using the corresponding visual cues. While most existing AVTSE methods rely exclusively on frontal-view videos, this…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-12 Peijun Yang , Zhan Jin , Juan Liu , Ming Li