English
Related papers

Related papers: Plug-and-Steer: Decoupling Separation and Selectio…

200 papers

Variational auto-encoder (VAE) is an effective neural network architecture to disentangle a speech utterance into speaker identity and linguistic content latent embeddings, then generate an utterance for a target speaker from that of a…

Sound · Computer Science 2022-08-23 Ziang Long , Yunling Zheng , Meng Yu , Jack Xin

Deep learning has shown a great potential for speech separation, especially for speech and non-speech separation. However, it encounters permutation problem for multi-speaker separation where both target and interference are speech.…

Sound · Computer Science 2021-03-29 Hao Li , Xueliang Zhang , Guanglai Gao

Audio-visual speech enhancement (AV-SE) aims to enhance degraded speech along with extra visual information such as lip videos, and has been shown to be more effective than audio-only speech enhancement. This paper proposes the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Rui-Chen Zheng , Yang Ai , Zhen-Hua Ling

We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify…

Sound · Computer Science 2022-07-22 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds…

Sound · Computer Science 2025-01-10 Yi Yuan , Xubo Liu , Haohe Liu , Mark D. Plumbley , Wenwu Wang

Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due to different interacting factors. Previous curriculum…

Sound · Computer Science 2026-03-06 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

Audio-Visual Speech Recognition (AVSR) enhances robustness in noisy environments by integrating visual cues. While recent advances integrate Large Language Models (LLMs) into AVSR, their high computational cost hinders deployment in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-22 Umberto Cappellazzo , Minsu Kim , Stavros Petridis , Daniele Falavigna , Alessio Brutti

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior…

Sound · Computer Science 2023-08-02 Chen Liu , Peike Li , Xingqun Qi , Hu Zhang , Lincheng Li , Dadong Wang , Xin Yu

Image retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-12 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Xinyue Song , Xianghu Yue

Target speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the mainstream audio-visual approaches which usually require…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Leyuan Qu , Cornelius Weber , Stefan Wermter

Integrating diverse visual capabilities into a unified model is a significant trend in Multimodal Large Language Models (MLLMs). Among these, the inclusion of segmentation poses a distinct set of challenges. To equip MLLMs with pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Jiazhen Liu , Long Chen

The objective of this paper is to perform visual sound separation: i) we study visual sound separation on spectrograms of different temporal resolutions; ii) we propose a new light yet efficient three-stream framework V-SlowFast that…

Computer Vision and Pattern Recognition · Computer Science 2021-09-22 Lingyu Zhu , Esa Rahtu

End-to-end transformer-based automatic speech recognition (ASR) systems often capture multiple speech traits in their learned representations that are highly entangled, leading to a lack of interpretability. In this study, we propose the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Pu Wang , Hugo Van hamme

While self-supervised learning (SSL) has revolutionized audio representation, the excessive parameterization and quadratic computational cost of standard Transformers limit their deployment on resource-constrained devices. To address this…

Sound · Computer Science 2026-03-30 Harunori Kawano , Takeshi Sasaki

Vision-Language Models (VLMs) excel at zero-shot inference but often degrade under test-time domain shifts. For this reason, episodic test-time adaptation strategies have recently emerged as powerful techniques for adapting VLMs to a single…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Konstantinos M. Dafnis , Dimitris N. Metaxas

We propose a novel Neural Steering technique that adapts the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-23 Martin Strauss , Wolfgang Mack , María Luis Valero , Okan Köpüklü

Recent speaker verification studies have achieved notable success by leveraging layer-wise output from pre-trained Transformer models. However, few have explored the advancements in aggregating these multi-level features beyond the static…

Sound · Computer Science 2025-12-30 Jin Sob Kim , Hyun Joon Park , Wooseok Shin , Sung Won Han

Audio-Visual Segmentation (AVS) aims to precisely outline audible objects in a visual scene at the pixel level. Existing AVS methods require fine-grained annotations of audio-mask pairs in supervised learning fashion. This limits their…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Xiatian Zhu

Recently, end-to-end speaker extraction has attracted increasing attention and shown promising results. However, its performance is often inferior to that of a blind source separation (BSS) counterpart with a similar network architecture,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-05 Zifeng Zhao , Dongchao Yang , Rongzhi Gu , Haoran Zhang , Yuexian Zou