English
Related papers

Related papers: A Comprehensive Multi-scale Approach for Speech an…

200 papers

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Speech-driven 3D facial animation has recently garnered attention due to its cost-effective usability in multimedia production. However, most current advances overlook the intelligibility of lip movements, limiting the realism of facial…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Han EunGi , Oh Hyun-Bin , Kim Sung-Bin , Corentin Nivelet Etcheberry , Suekyeong Nam , Janghoon Joo , Tae-Hyun Oh

We aim to edit the lip movements in talking video according to the given speech while preserving the personal identity and visual details. The task can be decomposed into two sub-problems: (1) speech-driven lip motion generation and (2)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Runyi Yu , Tianyu He , Ailing Zhang , Yuchi Wang , Junliang Guo , Xu Tan , Chang Liu , Jie Chen , Jiang Bian

In many applications, synchronizing audio with visuals is crucial, such as in creating graphic animations for films or games, translating movie audio into different languages, and developing metaverse applications. This review explores…

Talking face generation is the challenging task of synthesizing a natural and realistic face that requires accurate synchronization with a given audio. Due to co-articulation, where an isolated phone is influenced by the preceding or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Se Jin Park , Minsu Kim , Jeongsoo Choi , Yong Man Ro

Multi-view capture systems have been an important tool in research for recording human motion under controlling conditions. Most existing systems are specified around video streams and provide little or no support for audio acquisition and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Xiangwei Shi , Gara Dorta , Ruud de Jong , Ojas Shirekar , Chirag Raman

Speech-driven 3D facial animation with accurate lip synchronization has been widely studied. However, synthesizing realistic motions for the entire face during speech has rarely been explored. In this work, we present a joint audio-text…

Computer Vision and Pattern Recognition · Computer Science 2021-12-08 Yingruo Fan , Zhaojiang Lin , Jun Saito , Wenping Wang , Taku Komura

Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial movement over the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Yang Song , Jingwen Zhu , Dawei Li , Xiaolong Wang , Hairong Qi

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(speaking style) and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

Talking face generation has historically struggled to produce head movements and natural facial expressions without guidance from additional reference videos. Recent developments in diffusion-based generative models allow for more realistic…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Michał Stypułkowski , Konstantinos Vougioukas , Sen He , Maciej Zięba , Stavros Petridis , Maja Pantic

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Shuchen Weng , Haojie Zheng , Zheng Chang , Si Li , Boxin Shi , Xinlong Wang

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video motion, making it a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Jibin Song , Mingi Kwon , Jaeseok Jeong , Youngjung Uh

Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in response generation latency: (1) generating a spoken response…

Computation and Language · Computer Science 2024-10-04 Kentaro Mitsui , Koh Mitsuda , Toshiaki Wakatsuki , Yukiya Hono , Kei Sawada

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Tao Feng , Yifan Xie , Xun Guan , Jiyuan Song , Zhou Liu , Fei Ma , Fei Yu

In this work, we investigate the problem of lip-syncing a talking face video of an arbitrary identity to match a target speech segment. Current works excel at producing accurate lip movements on a static image or videos of specific people…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 K R Prajwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the unique speaking styles of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Weizhi Zhong , Jichang Li , Yinqi Cai , Ming Li , Feng Gao , Liang Lin , Guanbin Li

Recently, interactive digital human video generation has attracted widespread attention and achieved remarkable progress. However, building such a practical system that can interact with diverse input signals in real time remains…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Ming Chen , Liyuan Cui , Wenyuan Zhang , Haoxian Zhang , Yan Zhou , Xiaohan Li , Songlin Tang , Jiwen Liu , Borui Liao , Hejia Chen , Xiaoqiang Liu , Pengfei Wan

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual datasets, such as GRID,…

Sound · Computer Science 2023-03-02 Zhe Niu , Brian Mak

Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Phyo Thet Yee , Dimitrios Kollias , Sudeepta Mishra , Abhinav Dhall