English
Related papers

Related papers: Muse: Multi-modal target speaker extraction with v…

200 papers

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The…

Sound · Computer Science 2023-02-23 Meng Liu , Kong Aik Lee , Longbiao Wang , Hanyi Zhang , Chang Zeng , Jianwu Dang

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance…

This manuscript proposes a novel robust procedure for the extraction of a speaker of interest (SOI) from a mixture of audio sources. The estimation of the SOI is performed via independent vector extraction (IVE). Since the blind IVE cannot…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-29 Jiri Malek , Jakub Jansky , Zbynek Koldovsky , Tomas Kounovsky , Jaroslav Cmejla , Jindrich Zdansky

Large language models reveal deep comprehension and fluent generation in the field of multi-modality. Although significant advancements have been achieved in audio multi-modality, existing methods are rarely leverage language model for…

Sound · Computer Science 2024-08-06 Hualei Wang , Jianguo Mao , Zhifang Guo , Jiarui Wan , Hong Liu , Xiangdong Wang

Extracting the desired speech from a mixture is a meaningful and challenging task. The end-to-end DNN-based methods, though attractive, face the problem of generalization. In this paper, we explore a sequential approach for target speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-02 Zhaoyi Gu , Lele Liao , Kai Chen , Jing Lu

Recent speaker extraction methods using deep non-linear spatial filtering perform exceptionally well when the target direction is known and stationary. However, spatially dynamic scenarios are considerably more challenging due to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Jakob Kienegger , Timo Gerkmann

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Riki Shimizu , Xilin Jiang , Nima Mesgarani

Humans can listen to a target speaker even in challenging acoustic conditions that have noise, reverberation, and interfering speakers. This phenomenon is known as the cocktail-party effect. For decades, researchers have focused on…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-17 Katerina Zmolikova , Marc Delcroix , Tsubasa Ochiai , Keisuke Kinoshita , Jan Černocký , Dong Yu

We propose speaker separation using speaker inventories and estimated speech (SSUSIES), a framework leveraging speaker profiles and estimated speech for speaker separation. SSUSIES contains two methods, speaker separation using speaker…

Sound · Computer Science 2020-10-22 Peidong Wang , Zhuo Chen , DeLiang Wang , Jinyu Li , Yifan Gong

Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Hao Zhu , Huaibo Huang , Yi Li , Aihua Zheng , Ran He

Time-domain single-channel speech enhancement (SE) still remains challenging to extract the target speaker without any prior information on multi-talker conditions. It has been shown via auditory attention decoding that the brain activity…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-18 Jie Zhang , Qing-Tian Xu , Qiu-Shi Zhu , Zhen-Hua Ling

Speaker extraction and diarization are two enabling techniques for real-world speech applications. Speaker extraction aims to extract a target speaker's voice from a speech mixture, while speaker diarization demarcates speech segments by…

Sound · Computer Science 2025-01-17 Junyi Ao , Mehmet Sinan Yıldırım , Ruijie Tao , Meng Ge , Shuai Wang , Yanmin Qian , Haizhou Li

Speech Event Extraction (SpeechEE) is a challenging task that lies at the intersection of Automatic Speech Recognition (ASR) and Natural Language Processing (NLP), requiring the identification of structured event information from spoken…

Computation and Language · Computer Science 2025-09-30 Máté Gedeon

Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users…

Computation and Language · Computer Science 2024-03-27 Dominik Wagner , Alexander Churchill , Siddharth Sigtia , Panayiotis Georgiou , Matt Mirsamadi , Aarshee Mishra , Erik Marchi

In this paper, we present our solutions for the Multimodal Sentiment Analysis Challenge (MuSe) 2022, which includes MuSe-Humor, MuSe-Reaction and MuSe-Stress Sub-challenges. The MuSe 2022 focuses on humor detection, emotional reactions and…

Computer Vision and Pattern Recognition · Computer Science 2022-08-15 Jia Li , Ziyang Zhang , Junjie Lang , Yueqi Jiang , Liuwei An , Peng Zou , Yangyang Xu , Sheng Gao , Jie Lin , Chunxiao Fan , Xiao Sun , Meng Wang

Personalized speech enhancement (PSE) models achieve promising results compared with unconditional speech enhancement models due to their ability to remove interfering speech in addition to background noise. Unlike unconditional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-08 Hassan Taherian , Sefik Emre Eskimez , Takuya Yoshioka

Argument structure extraction (ASE) aims to identify the discourse structure of arguments within documents. Previous research has demonstrated that contextual information is crucial for developing an effective ASE model. However, we observe…

Computation and Language · Computer Science 2023-10-10 Yun Luo , Zhen Yang , Fandong Meng , Yingjie Li , Jie Zhou , Yue Zhang

This paper presents a novel approach to target speaker extraction (TSE) using Curriculum Learning (CL) techniques, addressing the challenge of distinguishing a target speaker's voice from a mixture containing interfering speakers. For…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

Extracting the speech of participants in a conversation amidst interfering speakers and noise presents a challenging problem. In this paper, we introduce the novel task of target conversation extraction, where the goal is to extract the…

Computation and Language · Computer Science 2024-09-26 Tuochao Chen , Qirui Wang , Bohan Wu , Malek Itani , Sefik Emre Eskimez , Takuya Yoshioka , Shyamnath Gollakota

When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be…

Multimedia · Computer Science 2020-08-19 Ying Cheng , Ruize Wang , Zhihao Pan , Rui Feng , Yuejie Zhang
‹ Prev 1 8 9 10 Next ›