English
Related papers

Related papers: A Neural Lip-Sync Framework for Synthesizing Photo…

200 papers

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first…

Co-speech gesture generation is to synthesize a gesture sequence that not only looks real but also matches with the input speech audio. Our method generates the movements of a complete upper body, including arms, hands, and the head.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Shenhan Qian , Zhi Tu , Yihao Zhi , Wen Liu , Shenghua Gao

Understanding the relationship between vocal tract motion during speech and the resulting acoustic signal is crucial for aided clinical assessment and developing personalized treatment and rehabilitation strategies. Toward this goal, we…

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and…

Sound · Computer Science 2024-12-12 Yifan Xie , Tao Feng , Xin Zhang , Xiangyang Luo , Zixuan Guo , Weijiang Yu , Heng Chang , Fei Ma , Fei Richard Yu

Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Ke Gu , Zhicong Wu , Peng Bai , Sitong Qiao , Zhiqi Jiang , Junchen Lu , Xiaodong Shi , Xinyuan Qian

Visual speech recognition (VSR), also known as lip reading, is the task of recognizing speech from silent video. Despite significant advancements in VSR over recent decades, most existing methods pay limited attention to real-world visual…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Tianyue Wang , Shuang Yang , Shiguang Shan , Xilin Chen

We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We address these limitations with a flow matching based framework.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Haiyang Liu , Xiaolin Hong , Xuancheng Yang , Yudi Ruan , Xiang Lian , Michael Lingelbach , Hongwei Yi , Wei Li

Lip-reading models have been significantly improved recently thanks to powerful deep learning architectures. However, most works focused on frontal or near frontal views of the mouth. As a consequence, lip-reading performance seriously…

Computer Vision and Pattern Recognition · Computer Science 2019-11-15 Shiyang Cheng , Pingchuan Ma , Georgios Tzimiropoulos , Stavros Petridis , Adrian Bulat , Jie Shen , Maja Pantic

Speech-driven 3D face animation technique, extending its applications to various multimedia fields. Previous research has generated promising realistic lip movements and facial expressions from audio signals. However, traditional regression…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Ziqiao Peng , Yihao Luo , Yue Shi , Hao Xu , Xiangyu Zhu , Jun He , Hongyan Liu , Zhaoxin Fan

The creation of high-quality human-labeled image-caption datasets presents a significant bottleneck in the development of Visual-Language Models (VLMs). In this work, we investigate an approach that leverages the strengths of Large Language…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Sahand Sharifzadeh , Christos Kaplanis , Shreya Pathak , Dharshan Kumaran , Anastasija Ilic , Jovana Mitrovic , Charles Blundell , Andrea Banino

Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent…

Computer Vision and Pattern Recognition · Computer Science 2020-10-14 Jianrong Wang , Tong Wu , Shanyu Wang , Mei Yu , Qiang Fang , Ju Zhang , Li Liu

Compositing is one of the most important editing operations for images and videos. The process of improving the realism of composite results is often called harmonization. Previous approaches for harmonization mainly focus on images. In…

Computer Vision and Pattern Recognition · Computer Science 2018-09-06 Haozhi Huang , Senzhe Xu , Junxiong Cai , Wei Liu , Shimin Hu

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Although current deep learning-based face forgery detectors achieve impressive performance in constrained scenarios, they are vulnerable to samples created by unseen manipulation methods. Some recent works show improvements in…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Alexandros Haliassos , Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Although image captioning models have made significant advancements in recent years, the majority of them heavily depend on high-quality datasets containing paired images and texts which are costly to acquire. Previous works leverage the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Zhiyue Liu , Jinyuan Liu , Fanrong Ma

Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos into classes…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-03 Yaman Kumar , Rohit Jain , Khwaja Mohd. Salik , Rajiv Ratn Shah , Yifang yin , Roger Zimmermann

When we speak, the prosody and content of the speech can be inferred from the movement of our lips. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate speech given only the lip movements of a speaker…

Computer Vision and Pattern Recognition · Computer Science 2022-06-29 Christen Millerdurai , Lotfy Abdel Khaliq , Timon Ulrich

In this paper, we propose a novel Lip-to-Speech synthesis (L2S) framework, for synthesizing intelligible speech from a silent lip movement video. Specifically, to complement the insufficient supervisory signal of the previous L2S model, we…

Sound · Computer Science 2023-06-01 Jeongsoo Choi , Minsu Kim , Yong Man Ro

Data-driven approaches hold promise for audio captioning. However, the development of audio captioning methods can be biased due to the limited availability and quality of text-audio data. This paper proposes a SynthAC framework, which…

Sound · Computer Science 2023-09-19 Feiyang Xiao , Qiaoxi Zhu , Jian Guan , Xubo Liu , Haohe Liu , Kejia Zhang , Wenwu Wang

In this paper, we formulate a novel task to synthesize speech in sync with a silent pre-recorded video, denoted as automatic voice over (AVO). Unlike traditional speech synthesis, AVO seeks to generate not only human-sounding speech, but…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-03 Junchen Lu , Berrak Sisman , Rui Liu , Mingyang Zhang , Haizhou Li