English
Related papers

Related papers: DREAM-Talk: Diffusion-based Realistic Emotional Au…

200 papers

We introduce a novel approach for high-resolution talking head generation from a single image and audio input. Prior methods using explicit face models, like 3D morphable models (3DMM) and facial landmarks, often fall short in generating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sejong Yang , Seoung Wug Oh , Yang Zhou , Seon Joo Kim

Generating synchronized and natural lip movement with speech is one of the most important tasks in creating realistic virtual characters. In this paper, we present a combined deep neural network of one-dimensional convolutions and LSTM to…

Sound · Computer Science 2022-05-03 Xiaohong Li , Xiang Wang , Kai Wang , Shiguo Lian

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zhen Ye , Xu Tan , Aoxiong Yin , Hongzhan Lin , Guangyan Zhang , Peiwen Sun , Yiming Li , Chi-Min Chan , Wei Ye , Shikun Zhang , Wei Xue

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to create realistic-looking…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Sahil Goyal , Shagun Uppal , Sarthak Bhagat , Yi Yu , Yifang Yin , Rajiv Ratn Shah

We propose an end to end deep learning approach for generating real-time facial animation from just audio. Specifically, our deep architecture employs deep bidirectional long short-term memory network and attention mechanism to discover the…

Machine Learning · Computer Science 2019-05-28 Guanzhong Tian , Yi Yuan , Yong liu

Producing expressive facial animations from static images is a challenging task. Prior methods relying on explicit geometric priors (e.g., facial landmarks or 3DMM) often suffer from artifacts in cross reenactment and struggle to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Qiang Wang , Mengchao Wang , Fan Jiang , Yaqi Fan , Yonggang Qi , Mu Xu

Generating consecutive images of lip movements that align with a given speech in audio-driven lip synthesis is a challenging task. While previous studies have made strides in synchronization and visual quality, lip intelligibility and video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Shiyan Liu , Rui Qu , Yan Jin

There has been significant progress in emotional Text-To-Speech (TTS) synthesis technology in recent years. However, existing methods primarily focus on the synthesis of a limited number of emotion types and have achieved unsatisfactory…

Sound · Computer Science 2023-06-02 Haobin Tang , Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the unique speaking styles of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Weizhi Zhong , Jichang Li , Yinqi Cai , Ming Li , Feng Gao , Liang Lin , Guanbin Li

Generating realistic listener facial motions in dyadic conversations remains challenging due to the high-dimensional action space and temporal dependency requirements. Existing approaches usually consider extracting 3D Morphable Model…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Zesheng Wang , Alexandre Bruckert , Patrick Le Callet , Guangtao Zhai

We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disentangled latent…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Duomin Wang , Yu Deng , Zixin Yin , Heung-Yeung Shum , Baoyuan Wang

This paper introduces DreamDiffusion, a novel method for generating high-quality images directly from brain electroencephalogram (EEG) signals, without the need to translate thoughts into text. DreamDiffusion leverages pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2023-07-03 Yunpeng Bai , Xintao Wang , Yan-pei Cao , Yixiao Ge , Chun Yuan , Ying Shan

Although previous co-speech gesture generation methods are able to synthesize motions in line with speech content, it is still not enough to handle diverse and complicated motion distribution. The key challenges are: 1) the one-to-many…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Lianying Yin , Yijun Wang , Tianyu He , Jinming Liu , Wei Zhao , Bohan Li , Xin Jin , Jianxin Lin

Generating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods that introduce 3D global representation into diffusion models have shown the potential to generate…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Zehuan Huang , Hao Wen , Junting Dong , Yaohui Wang , Yangguang Li , Xinyuan Chen , Yan-Pei Cao , Ding Liang , Yu Qiao , Bo Dai , Lu Sheng

Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging because of \textbf{1)} generalizable talking-face…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Bo Han , Heqing Zou , Haoyang Li , Guangcong Wang , Chng Eng Siong

Existing audio-driven video digital human generation models rely on multi-step denoising, resulting in substantial computational overhead that severely limits their deployment in real-world settings. While one-step distillation approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Xiangyu Liu , Feng Gao , Xiaomei Zhang , Yong Zhang , Xiaoming Wei , Zhen Lei , Xiangyu Zhu

Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchronization. This paper…

Sound · Computer Science 2026-02-03 Zhipeng Chen , Xinheng Wang , Lun Xie , Haijie Yuan , Hang Pan

Denoising Diffusion Probabilistic Models have shown extraordinary ability on various generative tasks. However, their slow inference speed renders them impractical in speech synthesis. This paper proposes a linear diffusion model (LinDiff)…

Sound · Computer Science 2023-06-13 Haogeng Liu , Tao Wang , Jie Cao , Ran He , Jianhua Tao

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that emphasize lip…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Chenxu Zhang , Zenan Li , Hongyi Xu , You Xie , Xiaochen Zhao , Tianpei Gu , Guoxian Song , Xin Chen , Chao Liang , Jianwen Jiang , Linjie Luo

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

Sound · Computer Science 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee