中文
相关论文

相关论文: Muse: Multi-modal target speaker extraction with v…

200 篇论文

In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the…

声音 · 计算机科学 2024-09-17 Dashanka De Silva , Siqi Cai , Saurav Pahuja , Tanja Schultz , Haizhou Li

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

多媒体 · 计算机科学 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Lianwu Chen , Yuexian Zou , Dong Yu

Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on speaker cues, such as pre-enrolled speech, to identify and…

音频与语音处理 · 电气工程与系统科学 2026-02-10 Haoyu Li , Yu Xi , Yidi Jiang , Shuai Wang , Kate Knill , Mark Gales , Haizhou Li , Kai Yu

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full…

音频与语音处理 · 电气工程与系统科学 2021-04-05 Meng Ge , Chenglin Xu , Longbiao Wang , Eng Siong Chng , Jianwu Dang , Haizhou Li

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker…

音频与语音处理 · 电气工程与系统科学 2019-09-24 Jian Wu , Yong Xu , Shi-Xiong Zhang , Lian-Wu Chen , Meng Yu , Lei Xie , Dong Yu

Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled speech is crucial to performance. However, these embeddings may…

声音 · 计算机科学 2025-08-12 Shu Wu , Anbin Qi , Yanzhang Xie , Xiang Xie

This paper describes an audio-visual speech enhancement (AV-SE) method that estimates from noisy input audio a mixture of the speech of the speaker appearing in an input video (on-screen target speech) and of a selected speaker not…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Tomoya Yoshinaga , Keitaro Tanaka , Shigeo Morishima

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semantics, to support…

声音 · 计算机科学 2025-06-17 Wenxuan Wu , Shuai Wang , Xixin Wu , Helen Meng , Haizhou Li

In machine lip-reading, which is identification of speech from visual-only information, there is evidence to show that visual speech is highly dependent upon the speaker [1]. Here, we use a phoneme-clustering method to form new…

计算机视觉与模式识别 · 计算机科学 2018-04-26 Helen L. Bear , Stephen J. Cox , Richard W. Harvey

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of…

音频与语音处理 · 电气工程与系统科学 2024-12-12 Ke Zhang , Junjie Li , Shuai Wang , Yangjie Wei , Yi Wang , Yannan Wang , Haizhou Li

The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a…

音频与语音处理 · 电气工程与系统科学 2023-03-10 Zexu Pan , Wupeng Wang , Marvin Borsdorf , Haizhou Li

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Mohamed Elminshawi , Wolfgang Mack , Srikanth Raj Chetupalli , Soumitro Chakrabarty , Emanuël A. P. Habets

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order, age, pitch level) are grouped by relative differences, while…

音频与语音处理 · 电气工程与系统科学 2025-06-10 Wang Dai , Archontis Politis , Tuomas Virtanen

Our objective is an audio-visual model for separating a single speaker from a mixture of sounds such as other speakers and background noise. Moreover, we wish to hear the speaker even when the visual cues are temporarily absent due to…

计算机视觉与模式识别 · 计算机科学 2019-07-12 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Multi-channel target speaker extraction (MC-TSE) aims to extract a target speaker's voice from multi-speaker signals captured by multiple microphones. Existing methods often rely on auxiliary clues such as direction-of-arrival (DOA) or…

音频与语音处理 · 电气工程与系统科学 2025-10-20 Tongtao Ling , Shulin He , Pengjie Shen , Zhong-Qiu Wang

We propose a multi-task universal speech enhancement (MUSE) model that can perform five speech enhancement (SE) tasks: dereverberation, denoising, speech separation (SS), target speaker extraction (TSE), and speaker counting. This is…

音频与语音处理 · 电气工程与系统科学 2023-10-13 Kohei Saijo , Wangyou Zhang , Zhong-Qiu Wang , Shinji Watanabe , Tetsunori Kobayashi , Tetsuji Ogawa

The remarkable ability of humans to selectively focus on a target speaker in cocktail party scenarios is facilitated by binaural audio processing. In this paper, we present a binaural time-domain Target Speaker Extraction model based on the…

音频与语音处理 · 电气工程与系统科学 2024-06-19 Hanyu Meng , Qiquan Zhang , Xiangyu Zhang , Vidhyasaharan Sethu , Eliathamby Ambikairajah