English
Related papers

Related papers: Plug-and-Steer: Decoupling Separation and Selectio…

200 papers

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-23 Kohei Saijo , Janek Ebbers , François G. Germain , Sameer Khurana , Gordon Wichern , Jonathan Le Roux

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provides a natural and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Xubo Liu , Qiuqiang Kong , Yan Zhao , Haohe Liu , Yi Yuan , Yuzhuo Liu , Rui Xia , Yuxuan Wang , Mark D. Plumbley , Wenwu Wang

In this paper, we propose long short term memory speech enhancement network (LSTMSE-Net), an audio-visual speech enhancement (AVSE) method. This innovative method leverages the complementary nature of visual and audio information to boost…

Sparse autoencoders (SAEs) have recently emerged as a powerful tool for language model steering. Prior work has explored top-k SAE latents for steering, but we observe that many dimensions among the top-k latents capture non-semantic…

Computation and Language · Computer Science 2025-10-03 Jiaqing Xie

Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks. However, there is less exploration concerning how SAM works on audio-visual tasks, such as visual sound localization and segmentation.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-04 Shentong Mo , Yapeng Tian

Humans can easily isolate a single speaker from a complex acoustic environment, a capability referred to as the "Cocktail Party Effect." However, replicating this ability has been a significant challenge in the field of target speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Xiang Hao , Jibin Wu , Jianwei Yu , Chenglin Xu , Kay Chen Tan

Traditional audio-visual methods rely on independent audio and visual backbones, which is costly and not scalable. In this work, we investigate using an audio-visual siamese network (AVSiam) for efficient and scalable audio-visual…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Yan-Bo Lin , Gedas Bertasius

Streaming recognition and segmentation of multi-party conversations with overlapping speech is crucial for the next generation of voice assistant applications. In this work we address its challenges discovered in the previous work on…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Ilya Sklyar , Anna Piunova , Christian Osendorfer

Audio-visual semantic segmentation (AVSS) represents an extension of the audio-visual segmentation (AVS) task, necessitating a semantic understanding of audio-visual scenes beyond merely identifying sound-emitting objects at the visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yujian Lee , Peng Gao , Yongqi Xu , Wentao Fan

This paper aims to improve the widely used deep speaker embedding x-vector model. We propose the following improvements: (1) a hybrid neural network structure using both time delay neural network (TDNN) and long short-term memory neural…

Computation and Language · Computer Science 2019-02-22 Yun Tang , Guohong Ding , Jing Huang , Xiaodong He , Bowen Zhou

Personalized or target speech extraction (TSE) typically needs a clean enrollment -- hard to obtain in real-world crowded environments. We remove the essential need for enrollment by predicting, from the mixture itself, a small set of…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-06 FNU Sidharth , Meysam Asgari , Hao-Wen Dong , Dhruv Jain

While existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios.…

Sound · Computer Science 2024-07-31 Tianrui Pan , Jie Liu , Bohan Wang , Jie Tang , Gangshan Wu

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Cunhang Fan , Ying Chen , Jian Zhou , Zexu Pan , Jingjing Zhang , Youdian Gao , Xiaoke Yang , Zhengqi Wen , Zhao Lv

We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mechanisms to integrate…

Sound · Computer Science 2024-09-18 Beilong Tang , Bang Zeng , Ming Li

Target speaker extraction (TSE) aims to isolate individual speaker voices from complex speech environments. The effectiveness of TSE systems is often compromised when the speaker characteristics are similar to each other. Recent research…

Sound · Computer Science 2024-10-08 Yun Liu , Xuechen Liu , Junichi Yamagishi

The goal of this paper is to introduce SPADE, a framework for Structured Pruning and Adaptive Distillation for Efficient Large Language Model-based text-to-speech (LLM-TTS). Recent LLM-TTS systems achieve strong controllability and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Tan Dat Nguyen , Jaehun Kim , Ji-Hoon Kim , Shukjae Choi , Youshin Lim , Joon Son Chung

We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speech recognition system. Delivering such a model presents…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Quan Wang , Ignacio Lopez Moreno , Mert Saglam , Kevin Wilson , Alan Chiao , Renjie Liu , Yanzhang He , Wei Li , Jason Pelecanos , Marily Nika , Alexander Gruenstein

Steering has emerged as a promising approach in controlling large language models (LLMs) without modifying model parameters. However, most existing steering methods rely on large-scale datasets to learn clear behavioral information, which…

Machine Learning · Computer Science 2025-10-06 Anyi Wang , Xuansheng Wu , Dong Shu , Yunpu Ma , Ninghao Liu

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological information for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-31 Runwu Shi , Benjamin Yen , Kazuhiro Nakadai

Audio-Visual Segmentation (AVS) aims to produce pixel-level masks of sound producing objects in videos, by jointly learning from audio and visual signals. However, real-world environments are inherently dynamic, causing audio and visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Siddeshwar Raghavan , Gautham Vinod , Bruce Coburn , Fengqing Zhu
‹ Prev 1 4 5 6 7 8 10 Next ›