中文
相关论文

相关论文: End to End Lip Synchronization with a Temporal Aut…

200 篇论文

The goal of automatic dubbing is to perform speech-to-speech translation while achieving audiovisual coherence. This entails isochrony, i.e., translating the original speech by also matching its prosodic structure into phrases and pauses,…

计算与语言 · 计算机科学 2022-04-07 Yogesh Virkar , Marcello Federico , Robert Enyedi , Roberto Barra-Chicote

Recent advances in diffusion-based lip-syncing generative models have demonstrated their ability to produce highly synchronized talking face videos for visual dubbing. Although these models excel at lip synchronization, they often struggle…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yanyu Zhu , Lichen Bai , Jintao Xu , Hai-tao Zheng

Previous studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facial area, leading to non-realistic results. In this work, we…

计算机视觉与模式识别 · 计算机科学 2022-12-12 Yasheng Sun , Hang Zhou , Kaisiyuan Wang , Qianyi Wu , Zhibin Hong , Jingtuo Liu , Errui Ding , Jingdong Wang , Ziwei Liu , Hideki Koike

The goal of this work is to temporally align asynchronous subtitles in sign language videos. In particular, we focus on sign-language interpreted TV broadcast data comprising (i) a video of continuous signing, and (ii) subtitles…

计算机视觉与模式识别 · 计算机科学 2021-05-07 Hannah Bull , Triantafyllos Afouras , Gül Varol , Samuel Albanie , Liliane Momeni , Andrew Zisserman

Recent talking head synthesis works typically adopt speech features extracted from large-scale pre-trained acoustic models. However, the intrinsic many-to-many relationship between speech and lip motion causes phoneme-viseme alignment…

图形学 · 计算机科学 2025-10-16 Yihuan Huang , Jiajun Liu , Yanzhen Ren , Jun Xue , Wuyang Liu , Zongkun Sun

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and…

声音 · 计算机科学 2024-12-12 Yifan Xie , Tao Feng , Xin Zhang , Xiangyang Luo , Zixuan Guo , Weijiang Yu , Heng Chang , Fei Ma , Fei Richard Yu

Recent advances in diffusion models have showcased promising results in the text-to-video (T2V) synthesis task. However, as these T2V models solely employ text as the guidance, they tend to struggle in modeling detailed temporal dynamics.…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Seungwoo Lee , Chaerin Kong , Donghyeon Jeon , Nojun Kwak

In the domain of photorealistic avatar generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Deng Junli , Luo Yihao , Yang Xueting , Li Siyou , Wang Wei , Guo Jinyang , Shi Ping

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of a single speaker, but a method that accounts for multiple…

音频与语音处理 · 电气工程与系统科学 2021-05-21 Dan Oneata , Adriana Stan , Horia Cucu

A fitting soundtrack can help a video better convey its content and provide a better immersive experience. This paper introduces a novel approach utilizing self-supervised learning and contrastive learning to automatically recommend audio…

多媒体 · 计算机科学 2025-03-10 Shimiao Liu , Alexander Lerch

Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on…

计算机视觉与模式识别 · 计算机科学 2020-05-14 Hao Zhu , Huaibo Huang , Yi Li , Aihua Zheng , Ran He

Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip movements, VSR tasks require effectively utilizing any…

声音 · 计算机科学 2024-10-23 Zehua Liu , Xiaolou Li , Chen Chen , Li Guo , Lantian Li , Dong Wang

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

计算机视觉与模式识别 · 计算机科学 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Audio-driven talking face generation aims to synthesize video with lip movements synchronized to input audio. However, current generative techniques face challenges in preserving intricate regional textures (skin, teeth). To address the…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Lingyu Xiong , Xize Cheng , Jintao Tan , Xianjia Wu , Xiandong Li , Lei Zhu , Fei Ma , Minglei Li , Huang Xu , Zhihu Hu

Lip reading, also known as visual speech recognition, aims to recognize the speech content from videos by analyzing the lip dynamics. There have been several appealing progress in recent years, benefiting much from the rapidly developed…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Dalu Feng , Shuang Yang , Shiguang Shan , Xilin Chen

The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of homophenes and (2)…

音频与语音处理 · 电气工程与系统科学 2024-01-05 Ji-Hoon Kim , Jaehun Kim , Joon Son Chung

The goal of Automatic Voice Over (AVO) is to generate speech in sync with a silent video given its text script. Recent AVO frameworks built upon text-to-speech synthesis (TTS) have shown impressive results. However, the current AVO learning…

音频与语音处理 · 电气工程与系统科学 2023-06-30 Junchen Lu , Berrak Sisman , Mingyang Zhang , Haizhou Li

The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder…

声音 · 计算机科学 2017-03-16 Tsubasa Ochiai , Shinji Watanabe , Takaaki Hori , John R. Hershey

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standing challenge to be further investigated. Motivated by xxx, in…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Ganglai Wang , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

We investigate unsupervised learning of correspondences between sound events and textual phrases through aligning audio clips with textual captions describing the content of a whole audio clip. We align originally unaligned and unannotated…

音频与语音处理 · 电气工程与系统科学 2022-02-22 Huang Xie , Okko Räsänen , Konstantinos Drossos , Tuomas Virtanen