中文
相关论文

相关论文: ReVISE: Self-Supervised Speech Resynthesis with Vi…

200 篇论文

Speech enhancement for voice pickup in hearables aims to improve the user's voice by suppressing noise and interfering talkers, while maintaining own-voice quality. For single-channel methods, it is particularly challenging to distinguish…

音频与语音处理 · 电气工程与系统科学 2026-02-05 Mattes Ohlenbusch , Mikolaj Kegler , Marko Stamenovic

Humans are adept at leveraging visual cues from lip movements for recognizing speech in adverse listening conditions. Audio-Visual Speech Recognition (AVSR) models follow similar approach to achieve robust speech recognition in noisy…

音频与语音处理 · 电气工程与系统科学 2024-05-24 Maxime Burchi , Krishna C. Puvvada , Jagadeesh Balam , Boris Ginsburg , Radu Timofte

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

声音 · 计算机科学 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions where traditional…

声音 · 计算机科学 2025-05-16 Detao Bai , Zhiheng Ma , Xihan Wei , Liefeng Bo

Automatic speech recognition (ASR) systems degrade significantly under noisy conditions. Recently, speech enhancement (SE) is introduced as front-end to reduce noise for ASR, but it also suppresses some important speech information, i.e.,…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Yuchen Hu , Nana Hou , Chen Chen , Eng Siong Chng

In this study, we try to address the problem of leveraging visual signals to improve Automatic Speech Recognition (ASR), also known as visual context-aware ASR (VC-ASR). We explore novel VC-ASR approaches to leverage video and text…

音频与语音处理 · 电气工程与系统科学 2020-11-10 Shahram Ghorbani , Yashesh Gaur , Yu Shi , Jinyu Li

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

声音 · 计算机科学 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

Unsupervised Zero-Shot Voice Conversion (VC) aims to modify the speaker characteristic of an utterance to match an unseen target speaker without relying on parallel training data. Recently, self-supervised learning of speech representation…

声音 · 计算机科学 2022-02-14 Trung Dang , Dung Tran , Peter Chin , Kazuhito Koishida

Speech is a means of communication which relies on both audio and visual information. The absence of one modality can often lead to confusion or misinterpretation of information. In this paper we present an end-to-end temporal model capable…

音频与语音处理 · 电气工程与系统科学 2019-06-17 Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Audio-visual automatic speech recognition (AV-ASR) models are very effective at reducing word error rates on noisy speech, but require large amounts of transcribed AV training data. Recently, audio-visual self-supervised learning (SSL)…

声音 · 计算机科学 2023-12-18 Avner May , Dmitriy Serdyuk , Ankit Parag Shah , Otavio Braga , Olivier Siohan

Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input speech to sound as though it was recorded…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Tingle Li , Renhao Wang , Po-Yao Huang , Andrew Owens , Gopala Anumanchipalli

Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to use deep neural…

音频与语音处理 · 电气工程与系统科学 2019-05-17 Ahmed Hussen Abdelaziz , Barry-John Theobald , Justin Binder , Gabriele Fanelli , Paul Dixon , Nicholas Apostoloff , Thibaut Weise , Sachin Kajareker

Enhancing speech signal quality in adverse acoustic environments is a persistent challenge in speech processing. Existing deep learning based enhancement methods often struggle to effectively remove background noise and reverberation in…

音频与语音处理 · 电气工程与系统科学 2023-09-19 Heming Wang , Meng Yu , Hao Zhang , Chunlei Zhang , Zhongweiyang Xu , Muqiao Yang , Yixuan Zhang , Dong Yu

Many expressive visualizations are shared online only as bitmap images, making them difficult to redesign or adapt to new data. Reusing such image-based visualizations requires substantial expertise and is often time-consuming, even for…

人机交互 · 计算机科学 2026-04-20 Xiaolin Wen , Changlin Li , Manusha Karunathilaka , Can Liu , Fangzhuo Jin , Yong Wang

This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form audio (e.g.,…

This paper discusses the task of face-based speech synthesis, a kind of personalized speech synthesis where the synthesized voices are constrained to perceptually match with a reference face image. Due to the lack of TTS-quality…

音频与语音处理 · 电气工程与系统科学 2026-01-07 Yao Shi , Yunfei Xu , Hongbin Suo , Yulong Wan , Haifeng Liu

Audio-Visual Speech Recognition (AVSR) combines auditory and visual speech cues to enhance the accuracy and robustness of speech recognition systems. Recent advancements in AVSR have improved performance in noisy environments compared to…

音频与语音处理 · 电气工程与系统科学 2025-04-29 Zhaofeng Lin , Naomi Harte

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework, where the input and output of the system are multimodal (i.e., audio and visual speech). With the proposed AV2AV, two key…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Jeongsoo Choi , Se Jin Park , Minsu Kim , Yong Man Ro

Speech synthesis (text to speech, TTS) and recognition (automatic speech recognition, ASR) are important speech tasks, and require a large amount of text and speech pairs for model training. However, there are more than 6,000 languages in…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Jin Xu , Xu Tan , Yi Ren , Tao Qin , Jian Li , Sheng Zhao , Tie-Yan Liu