English
Related papers

Related papers: Training Strategies for Modality Dropout Resilient…

200 papers

Target speaker extraction (TSE) aims to recover the speech of a desired speaker from a mixture given a short enrollment utterance, while speech enhancement (SE) focuses on improving speech quality under noisy conditions. Most existing TSE…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Bang Zeng , Beilong Tang , Wang Xiang , Ming Li

Traditional speech separation and speaker diarization approaches rely on prior knowledge of target speakers or a predetermined number of participants in audio signals. To address these limitations, recent advances focus on developing…

Deep-learning based speech separation models confront poor generalization problem that even the state-of-the-art models could abruptly fail when evaluating them in mismatch conditions. To address this problem, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-04 Max W. Y. Lam , Jun Wang , Dan Su , Dong Yu

We propose a multi-task universal speech enhancement (MUSE) model that can perform five speech enhancement (SE) tasks: dereverberation, denoising, speech separation (SS), target speaker extraction (TSE), and speaker counting. This is…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-13 Kohei Saijo , Wangyou Zhang , Zhong-Qiu Wang , Shinji Watanabe , Tetsunori Kobayashi , Tetsuji Ogawa

Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real-time TSE is challenging as the computational complexity…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-03 Hiroshi Sato , Takafumi Moriya , Masato Mimura , Shota Horiguchi , Tsubasa Ochiai , Takanori Ashihara , Atsushi Ando , Kentaro Shinayama , Marc Delcroix

Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an offline fashion and are not suited to the low-latency causal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-22 Shrishail Baligar , Mikolaj Kegler , Bryce Irvin , Marko Stamenovic , Shawn Newsam

Expressive text-to-speech (TTS) aims to synthesize different speaking style speech according to human's demands. Nowadays, there are two common ways to control speaking styles: (1) Pre-defining a group of speaking style and using…

Sound · Computer Science 2023-06-27 Dongchao Yang , Songxiang Liu , Rongjie Huang , Chao Weng , Helen Meng

In the development of neural text-to-speech systems, model pre-training with a large amount of non-target speakers' data is a common approach. However, in terms of ultimately achieved system performance for target speaker(s), the actual…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-11 Guangyan Zhang , Yichong Leng , Daxin Tan , Ying Qin , Kaitao Song , Xu Tan , Sheng Zhao , Tan Lee

Despite recent advancements in offline multi-task reinforcement learning (MTRL) have harnessed the powerful capabilities of the Transformer architecture, most approaches focus on a limited number of tasks, with scaling to extremely massive…

Machine Learning · Computer Science 2025-06-02 Yilun Kong , Guozheng Ma , Qi Zhao , Haoyu Wang , Li Shen , Xueqian Wang , Dacheng Tao

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-25 Hongyu Wang , Chenda Li , Xin Zhou , Shuai Wang , Yanmin Qian

Social media enables data-driven analysis of public opinion on contested issues. Target-Stance Extraction (TSE) is the task of identifying the target discussed in a document and the document's stance towards that target. Many works classify…

Computation and Language · Computer Science 2025-10-28 Ethan Mines , Bonnie Dorr

Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We present GenTSE, a two-stage decoder-only generative LM approach for TSE:…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-25 Haoyang Li , Xuyi Zhuang , Azmat Adnan , Ye Ni , Wei Rao , Shreyas Gopal , Eng Siong Chng

We propose a multichannel-to-multichannel target sound extraction (M2M-TSE) framework for separating multichannel target signals from a multichannel mixture of sound sources. Target sound extraction (TSE) isolates a specific target signal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Dayun Choi , Jung-Woo Choi

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always available due to various…

Sound · Computer Science 2025-04-01 Junjie Li , Ke Zhang , Shuai Wang , Kong Aik Lee , Man-Wai Mak , Haizhou Li

This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Processing (MISP) 2023 Challenge. Specifically, our approach adopts…

Sound · Computer Science 2024-03-08 Runduo Han , Xiaopeng Yan , Weiming Xu , Pengcheng Guo , Jiayao Sun , He Wang , Quan Lu , Ning Jiang , Lei Xie

Single channel target speaker separation (TSS) aims at extracting a speaker's voice from a mixture of multiple talkers given an enrollment utterance of that speaker. A typical deep learning TSS framework consists of an upstream model that…

Sound · Computer Science 2022-10-27 Xiaoyu Liu , Xu Li , Joan Serrà

Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occurs'' according to the semantic characteristics of speech. We…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Yidi Jiang , Ruijie Tao , Zhengyang Chen , Yanmin Qian , Haizhou Li

With the acceleration of globalization, more and more people are willing or required to learn second languages (L2). One of the major remaining challenges facing current mispronunciation and diagnosis (MDD) models for use in…

Multimedia · Computer Science 2021-10-05 Shao-Wei Fan Jiang , Bi-Cheng Yan , Tien-Hong Lo , Fu-An Chao , Berlin Chen

Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the problems of weakly-labelled data, jointly learning and new…

Sound · Computer Science 2022-04-05 Helin Wang , Dongchao Yang , Chao Weng , Jianwei Yu , Yuexian Zou

Audio-visual target speech extraction (AV-TSE) is one of the enabling technologies in robotics and many audio-visual applications. One of the challenges of AV-TSE is how to effectively utilize audio-visual synchronization information in the…

Sound · Computer Science 2024-03-26 Wenxuan Wu , Xueyuan Chen , Xixin Wu , Haizhou Li , Helen Meng