English
Related papers

Related papers: CLAPSep: Leveraging Contrastive Pre-trained Model …

200 papers

Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled speech is crucial to performance. However, these embeddings may…

Sound · Computer Science 2025-08-12 Shu Wu , Anbin Qi , Yanzhang Xie , Xiang Xie

Recent years have seen progress beyond domain-specific sound separation for speech or music towards universal sound separation for arbitrary sounds. Prior work on universal sound separation has investigated separating a target sound out of…

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target…

Speech Emotion Recognition (SER) is fundamental to affective computing and human-computer interaction, yet existing models struggle to generalize across diverse acoustic conditions. While Contrastive Language-Audio Pretraining (CLAP)…

Sound · Computer Science 2025-07-08 Jiacheng Shi , Yanfu Zhang , Ye Gao

Universal sound separation faces a fundamental misalignment: models optimized for low-level signal metrics often produce semantically contaminated outputs, failing to suppress perceptually salient interference from acoustically similar…

Sound · Computer Science 2026-02-18 Zihan Zhang , Xize Cheng , Zhennan Jiang , Dongjie Fu , Jingyuan Chen , Zhou Zhao , Tao Jin

Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an offline fashion and are not suited to the low-latency causal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-22 Shrishail Baligar , Mikolaj Kegler , Bryce Irvin , Marko Stamenovic , Shawn Newsam

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters…

Sound · Computer Science 2024-01-08 Shulin He , Jinjiang liu , Hao Li , Yang Yang , Fei Chen , Xueliang Zhang

Improving text representation has attracted much attention to achieve expressive text-to-speech (TTS). However, existing works only implicitly learn the prosody with masked token reconstruction tasks, which leads to low training efficiency…

Sound · Computer Science 2023-05-19 Zhenhui Ye , Rongjie Huang , Yi Ren , Ziyue Jiang , Jinglin Liu , Jinzheng He , Xiang Yin , Zhou Zhao

Contrastive language-audio pretraining (CLAP) has recently emerged as a method for making audio analysis more generalisable. Specifically, CLAP-style models are able to `answer' a diverse set of language queries, extending the capabilities…

Sound · Computer Science 2024-06-12 Xin Jing , Andreas Triantafyllopoulos , Björn Schuller

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds…

Sound · Computer Science 2025-01-10 Yi Yuan , Xubo Liu , Haohe Liu , Mark D. Plumbley , Wenwu Wang

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Target speaker extraction (TSE) is a technique for isolating a target speaker's voice from mixed speech using auxiliary features associated with the target speaker. It is another attempt at addressing the cocktail party problem and is…

Sound · Computer Science 2024-11-26 Chang Sun , Bo Qin

Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spoke when,' while Target Speaker Extraction (TSE) or Target…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-20 Bang Zeng , Ming Li

Diffusion model-based speech enhancement has received increased attention since it can generate very natural enhanced signals and generalizes well to unseen conditions. Diffusion models have been explored for several sub-tasks of speech…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-21 Naoyuki Kamo , Marc Delcroix , Tomohiro Nakatani

Previously, Target Speaker Extraction (TSE) has yielded outstanding performance in certain application scenarios for speech enhancement and source separation. However, obtaining auxiliary speaker-related information is still challenging in…

Universal sound separation (USS) is a task to separate arbitrary sounds from an audio mixture. Existing USS systems are capable of separating arbitrary sources, given a few examples of the target sources as queries. However, separating…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-01 Yuzhuo Liu , Xubo Liu , Yan Zhao , Yuanyuan Wang , Rui Xia , Pingchuan Tain , Yuxuan Wang

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Shuai Wang , Ke Zhang , Shaoxiong Lin , Junjie Li , Xuefei Wang , Meng Ge , Jianwei Yu , Yanmin Qian , Haizhou Li

Personalized or target speech extraction (TSE) typically needs a clean enrollment -- hard to obtain in real-world crowded environments. We remove the essential need for enrollment by predicting, from the mixture itself, a small set of…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-06 FNU Sidharth , Meysam Asgari , Hao-Wen Dong , Dhruv Jain

The query-based audio separation usually employs specific queries to extract target sources from a mixture of audio signals. Currently, most query-based separation models need additional networks to obtain query embedding. In this way,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-27 Yuanyuan Wang , Hangting Chen , Dongchao Yang , Jianwei Yu , Chao Weng , Zhiyong Wu , Helen Meng

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Mohamed Elminshawi , Wolfgang Mack , Srikanth Raj Chetupalli , Soumitro Chakrabarty , Emanuël A. P. Habets