中文
相关论文

相关论文: VocaLiST: An Audio-Visual Synchronisation Model fo…

200 篇论文

In this paper, we present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extremely simple approach to generating (weak) speech…

多媒体 · 计算机科学 2017-06-02 Ken Hoover , Sourish Chaudhuri , Caroline Pantofaru , Malcolm Slaney , Ian Sturdy

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

计算与语言 · 计算机科学 2025-04-11 Lakshmipathi Balaji , Karan Singla

Speech is the most used communication method between humans and it involves the perception of auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, although the video can provide information…

计算机视觉与模式识别 · 计算机科学 2017-04-27 Adriana Fernandez-Lopez , Oriol Martinez , Federico M. Sukno

In this work, we re-think the task of speech enhancement in unconstrained real-world environments. Current state-of-the-art methods use only the audio stream and are limited in their performance in a wide range of real-world noises. Recent…

计算机视觉与模式识别 · 计算机科学 2020-12-22 Sindhu B Hegde , K R Prajwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C. V. Jawahar

We present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disentangles this…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Kun Cheng , Xiaodong Cun , Yong Zhang , Menghan Xia , Fei Yin , Mingrui Zhu , Xuan Wang , Jue Wang , Nannan Wang

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated…

Audio-visual synchronization aims to determine whether the mouth movements and speech in the video are synchronized. VocaLiST reaches state-of-the-art performance by incorporating multimodal Transformers to model audio-visual interact…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Xuanjun Chen , Haibin Wu , Chung-Che Wang , Hung-yi Lee , Jyh-Shing Roger Jang

Audiovisual synchronisation is the task of determining the time offset between speech audio and a video recording of the articulators. In child speech therapy, audio and ultrasound videos of the tongue are captured using instruments which…

计算与语言 · 计算机科学 2019-11-28 Aciel Eshky , Manuel Sam Ribeiro , Korin Richmond , Steve Renals

Lip reading, also known as visual speech recognition, aims to recognize the speech content from videos by analyzing the lip dynamics. There have been several appealing progress in recent years, benefiting much from the rapidly developed…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Dalu Feng , Shuang Yang , Shiguang Shan , Xilin Chen

This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is based on a two-stage network. Motion cues are obtained with…

声音 · 计算机科学 2022-07-20 Juan F. Montesinos , Venkatesh S. Kadandale , Gloria Haro

Our objective is an audio-visual model for separating a single speaker from a mixture of sounds such as other speakers and background noise. Moreover, we wish to hear the speaker even when the visual cues are temporarily absent due to…

计算机视觉与模式识别 · 计算机科学 2019-07-12 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Tao Feng , Yifan Xie , Xun Guan , Jiyuan Song , Zhou Liu , Fei Ma , Fei Yu

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Chunyu Li , Chao Zhang , Weikai Xu , Jingyu Lin , Jinghui Xie , Weiguo Feng , Bingyue Peng , Cunjian Chen , Weiwei Xing

In this paper, we present StyleLipSync, a style-based personalized lip-sync video generative model that can generate identity-agnostic lip-synchronizing video from arbitrary audio. To generate a video of arbitrary identities, we leverage…

计算机视觉与模式识别 · 计算机科学 2024-02-13 Taekyung Ki , Dongchan Min

Cross-modality generation is an emerging topic that aims to synthesize data in one modality based on information in a different modality. In this paper, we consider a task of such: given an arbitrary audio speech and one lip image of…

计算机视觉与模式识别 · 计算机科学 2018-05-23 Lele Chen , Zhiheng Li , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Silent speech interfaces have been recently proposed as a way to enable communication when the acoustic signal is not available. This introduces the need to build visual speech recognition systems for silent and whispered speech. However,…

计算机视觉与模式识别 · 计算机科学 2018-02-20 Stavros Petridis , Jie Shen , Doruk Cetin , Maja Pantic

Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Ke Gu , Zhicong Wu , Peng Bai , Sitong Qiao , Zhiqi Jiang , Junchen Lu , Xiaodong Shi , Xinyuan Qian

Lip-reading aims to recognize speech content from videos via visual analysis of speakers' lip movements. This is a challenging task due to the existence of homophemes-words which involve identical or highly similar lip movements, as well as…

计算机视觉与模式识别 · 计算机科学 2019-09-04 Chenhao Wang

Previous studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facial area, leading to non-realistic results. In this work, we…

计算机视觉与模式识别 · 计算机科学 2022-12-12 Yasheng Sun , Hang Zhou , Kaisiyuan Wang , Qianyi Wu , Zhibin Hong , Jingtuo Liu , Errui Ding , Jingdong Wang , Ziwei Liu , Hideki Koike

Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences in videos to words poses significant challenges due to…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Fei Yu , Jun Wang