中文
相关论文

相关论文: Two-stage Audio-Visual Target Speaker Extraction S…

200 篇论文

This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological information for…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Runwu Shi , Benjamin Yen , Kazuhiro Nakadai

Humans are adept at leveraging visual cues from lip movements for recognizing speech in adverse listening conditions. Audio-Visual Speech Recognition (AVSR) models follow similar approach to achieve robust speech recognition in noisy…

音频与语音处理 · 电气工程与系统科学 2024-05-24 Maxime Burchi , Krishna C. Puvvada , Jagadeesh Balam , Boris Ginsburg , Radu Timofte

Diffusion model-based speech enhancement has received increased attention since it can generate very natural enhanced signals and generalizes well to unseen conditions. Diffusion models have been explored for several sub-tasks of speech…

音频与语音处理 · 电气工程与系统科学 2023-08-21 Naoyuki Kamo , Marc Delcroix , Tomohiro Nakatani

Despite recent advances in voice separation methods, many challenges remain in realistic scenarios such as noisy recording and the limits of available data. In this work, we propose to explicitly incorporate the phonetic and linguistic…

In this research, we present an innovative, parameter-efficient model that utilizes the attention U-Net architecture for the automatic detection and eradication of non-speech vocal sounds, specifically breath sounds, in vocal recordings.…

声音 · 计算机科学 2024-09-10 Nidula Elgiriyewithana , N. D. Kodikara

Voice activity detection (VAD) is an important pre-processing step for speech technology applications. The task consists of deriving segment boundaries of audio signals which contain voicing information. In recent years, it has been shown…

音频与语音处理 · 电气工程与系统科学 2023-03-28 Eklavya Sarkar , RaviShankar Prasad , Mathew Magimai. -Doss

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle overlapping speech.…

Target speech extraction (TSE) has achieved strong performance in relatively simple conditions such as one-speaker-plus-noise and two-speaker mixtures, but its performance remains unsatisfactory in noisy multi-speaker scenarios. To address…

音频与语音处理 · 电气工程与系统科学 2026-03-16 Ziling Huang , Junnan Wu , Lichun Fan , Zhenbo Luo , Jian Luan , Haixin Guan , Yanhua Long

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain underexplored, as…

音频与语音处理 · 电气工程与系统科学 2025-06-10 The Hieu Pham , Phuong Thanh Tran Nguyen , Xuan Tho Nguyen , Tan Dat Nguyen , Duc Dung Nguyen

The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining…

音频与语音处理 · 电气工程与系统科学 2023-10-31 Suyeon Lee , Chaeyoung Jung , Youngjoon Jang , Jaehun Kim , Joon Son Chung

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

声音 · 计算机科学 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

We propose listen to extract (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting the target speaker from the…

音频与语音处理 · 电气工程与系统科学 2025-11-06 Pengjie Shen , Kangrui Chen , Shulin He , Pengru Chen , Shuqi Yuan , He Kong , Xueliang Zhang , Zhong-Qiu Wang

The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment and visual feeds corresponding to the target speaker. MTSE…

音频与语音处理 · 电气工程与系统科学 2025-07-10 Srikanth Korse , Mohamed Elminshawi , Emanuel A. P. Habets , Srikanth Raj Chetupalli

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a…

声音 · 计算机科学 2022-07-14 Joanna Hong , Minsu Kim , Daehun Yoo , Yong Man Ro

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been…

音频与语音处理 · 电气工程与系统科学 2021-03-16 Daniel Michelsanti , Zheng-Hua Tan , Shi-Xiong Zhang , Yong Xu , Meng Yu , Dong Yu , Jesper Jensen

Personalization of on-device speech recognition (ASR) has seen explosive growth in recent years, largely due to the increasing popularity of personal assistant features on mobile devices and smart home speakers. In this work, we present…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Shaojin Ding , Rajeev Rikhye , Qiao Liang , Yanzhang He , Quan Wang , Arun Narayanan , Tom O'Malley , Ian McGraw

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text…

音频与语音处理 · 电气工程与系统科学 2024-09-23 Kohei Saijo , Janek Ebbers , François G. Germain , Sameer Khurana , Gordon Wichern , Jonathan Le Roux

The common target speech separation directly estimate the target source, ignoring the interrelationship between different speakers at each frame. We propose a multiple-target speech separation model (MTSS) to simultaneously extract each…

音频与语音处理 · 电气工程与系统科学 2023-11-21 Bang Zeng , Hongbing Suo , Yulong Wan , Ming Li
‹ 上一页 1 8 9 10 下一页 ›