English
Related papers

Related papers: pTSE-T: Presentation Target Speaker Extraction usi…

200 papers

To extract the voice of a target speaker when mixed with a variety of other sounds, such as white and ambient noises or the voices of interfering speakers, we extend the Transformer network to attend the most relevant information with…

Social media enables data-driven analysis of public opinion on contested issues. Target-Stance Extraction (TSE) is the task of identifying the target discussed in a document and the document's stance towards that target. Many works classify…

Computation and Language · Computer Science 2025-10-28 Ethan Mines , Bonnie Dorr

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the…

Sound · Computer Science 2023-09-18 Junjie Li , Ruijie Tao , Zexu Pan , Meng Ge , Shuai Wang , Haizhou Li

Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the sliding window approach,…

Sound · Computer Science 2022-11-09 Hee-Soo Heo , Youngki Kwon , Bong-Jin Lee , You Jin Kim , Jee-weon Jung

The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of…

Sound · Computer Science 2023-10-18 Yu Chen , Xinyuan Qian , Zexu Pan , Kainan Chen , Haizhou Li

The reconstruction of clipped speech signals is an important task in audio signal processing to achieve an enhanced audio quality for further processing. In this paper, Frequency Selective Extrapolation (FSE), which is commonly used for…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-11 Markus Jonscher , Jürgen Seiler , André Kaup

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Helin Wang , Jiarui Hai , Yen-Ju Lu , Karan Thakkar , Mounya Elhilali , Najim Dehak

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-05 Meng Ge , Chenglin Xu , Longbiao Wang , Eng Siong Chng , Jianwu Dang , Haizhou Li

Developing a single-microphone speech denoising or dereverberation front-end for robust automatic speaker verification (ASV) in noisy far-field speaking scenarios is challenging. To address this problem, we present a novel front-end design…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-28 Joon-Young Yang , Joon-Hyuk Chang

We propose a knowledge-driven approach to speech target extraction in the presence of background sound effects already recorded in cinematic audio. The specific knowledge sources studied are manners of articulation that are detected in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-01 Chun-wei Ho , Sabato Marco Siniscalchi , Kai Li , Chin-Hui Lee

Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes remains challenging, as the output often inherits both the accent and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-09 Mu Yang , John H. L. Hansen

Target speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the mainstream audio-visual approaches which usually require…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Leyuan Qu , Cornelius Weber , Stefan Wermter

Speaker-conditioned target speaker extraction systems rely on auxiliary information about the target speaker to extract the target speaker signal from a mixture of multiple speakers. Typically, a deep neural network is applied to isolate…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-12 Ragini Sinha , Marvin Tammen , Christian Rollwage , Simon Doclo

The complete decomposition performed by blind source separation is computationally demanding and superfluous when only the speech of one specific target speaker is desired. In this paper, we propose a computationally efficient blind speech…

Sound · Computer Science 2020-08-04 Lele Liao , Zhaoyi Gu , Jing Lu

The common target speech separation directly estimate the target source, ignoring the interrelationship between different speakers at each frame. We propose a multiple-target speech separation model (MTSS) to simultaneously extract each…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Bang Zeng , Hongbing Suo , Yulong Wan , Ming Li

Discrete audio tokens derived from self-supervised learning models have gained widespread usage in speech generation. However, current practice of directly utilizing audio tokens poses challenges for sequence modeling due to the length of…

Sound · Computer Science 2024-01-17 Feiyu Shen , Yiwei Guo , Chenpeng Du , Xie Chen , Kai Yu

Recently, end-to-end speaker extraction has attracted increasing attention and shown promising results. However, its performance is often inferior to that of a blind source separation (BSS) counterpart with a similar network architecture,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-05 Zifeng Zhao , Dongchao Yang , Rongzhi Gu , Haoran Zhang , Yuexian Zou

Multi-modal cues, including spatial information, facial expression and voiceprint, are introduced to the speech separation and speaker extraction tasks to serve as complementary information to achieve better performance. However, the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-12 Qinghua Liu , Yating Huang , Yunzhe Hao , Jiaming Xu , Bo Xu

Text-based speech editing (TSE) allows users to edit speech by modifying the corresponding text directly without altering the original recording. Current TSE techniques often focus on minimizing discrepancies between generated speech and…

Computation and Language · Computer Science 2024-12-10 Rui Liu , Jiatian Xi , Ziyue Jiang , Haizhou Li

In this paper, we present SANE-TTS, a stable and natural end-to-end multilingual TTS model. By the difficulty of obtaining multilingual corpus for given speaker, training multilingual TTS model with monolingual corpora is unavoidable. We…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Hyunjae Cho , Wonbin Jung , Junhyeok Lee , Sang Hoon Woo