中文
相关论文

相关论文: Hierarchical Cross-Modal Talking Face Generationwi…

200 篇论文

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. While audio-driven…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Changpeng Cai , Guinan Guo , Jiao Li , Junhao Su , Fei Shen , Chenghao He , Jing Xiao , Yuanxu Chen , Lei Dai , Feiyu Zhu

We present a novel approach for generating realistic speaking and talking faces by synthesizing a person's voice and facial movements from a static image, a voice profile, and a target text. The model encodes the prompt/driving text, the…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Aashish Chandra , Aashutosh A , Abhijit Das

We propose a multi-scale GAN model to hallucinate realistic context (forehead, hair, neck, clothes) and background pixels automatically from a single input face mask. Instead of swapping a face on to an existing picture, our model directly…

计算机视觉与模式识别 · 计算机科学 2020-01-14 Sandipan Banerjee , Walter J. Scheirer , Kevin W. Bowyer , Patrick J. Flynn

We present variational generative adversarial networks, a general learning framework that combines a variational auto-encoder with a generative adversarial network, for synthesizing images in fine-grained categories, such as faces of a…

计算机视觉与模式识别 · 计算机科学 2018-02-06 Jianmin Bao , Dong Chen , Fang Wen , Houqiang Li , Gang Hua

Face parsing is a fundamental task in computer vision, enabling applications such as identity verification, facial editing, and controllable image synthesis. However, existing face parsing models often lack fairness and robustness, leading…

计算机视觉与模式识别 · 计算机科学 2025-02-10 Sophia J. Abraham , Jonathan D. Hauenstein , Walter J. Scheirer

Taking a photo outside, can we predict the immediate future, e.g., how would the cloud move in the sky? We address this problem by presenting a generative adversarial network (GAN) based two-stage approach to generating realistic time-lapse…

计算机视觉与模式识别 · 计算机科学 2018-04-02 Wei Xiong , Wenhan Luo , Lin Ma , Wei Liu , Jiebo Luo

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photo-realistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Shivangi Aneja , Artem Sevastopolsky , Tobias Kirschstein , Justus Thies , Angela Dai , Matthias Nießner

Talking head generation intends to produce vivid and realistic talking head videos from a single portrait and speech audio clip. Although significant progress has been made in diffusion-based talking head generation, almost all methods rely…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Hanbo Cheng , Limin Lin , Chenyu Liu , Pengcheng Xia , Pengfei Hu , Jiefeng Ma , Jun Du , Jia Pan

Generative models have advanced rapidly, enabling impressive talking head generation that brings AI to life. However, most existing methods focus solely on one-way portrait animation. Even the few that support bidirectional conversational…

音频与语音处理 · 电气工程与系统科学 2025-11-25 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

Articulation, emotion, and personality play strong roles in the orofacial movements. To improve the naturalness and expressiveness of virtual agents (VAs), it is important that we carefully model the complex interplay between these factors.…

人机交互 · 计算机科学 2023-05-15 Najmeh Sadoughi , Carlos Busso

We present a new multi-modal face image generation method that converts a text prompt and a visual input, such as a semantic mask or scribble map, into a photo-realistic face image. To do this, we combine the strengths of Generative…

计算机视觉与模式识别 · 计算机科学 2024-05-08 Jihyun Kim , Changjae Oh , Hoseok Do , Soohyun Kim , Kwanghoon Sohn

The task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data…

计算机视觉与模式识别 · 计算机科学 2022-01-19 Tianyi Xie , Liucheng Liao , Cheng Bi , Benlai Tang , Xiang Yin , Jianfei Yang , Mingjie Wang , Jiali Yao , Yang Zhang , Zejun Ma

In this paper, we propose an effective face completion algorithm using a deep generative model. Different from well-studied background completion, the face completion task is more challenging as it often requires to generate semantically…

计算机视觉与模式识别 · 计算机科学 2017-04-20 Yijun Li , Sifei Liu , Jimei Yang , Ming-Hsuan Yang

In this paper, we propose a novel algorithm for matching faces with temporal variations caused due to age progression. The proposed generative adversarial network algorithm is a unified framework that combines facial age estimation and…

计算机视觉与模式识别 · 计算机科学 2020-11-12 Daksha Yadav , Naman Kohli , Mayank Vatsa , Richa Singh , Afzel Noore

Generative adversarial networks have been widely used in image synthesis in recent years and the quality of the generated image has been greatly improved. However, the flexibility to control and decouple facial attributes (e.g., eyes, nose,…

计算机视觉与模式识别 · 计算机科学 2021-08-26 Xiao Cui , Wengang Zhou , Yang Hu , Weilun Wang , Houqiang Li

Recent research has demonstrated the ability to estimate gaze on mobile devices by performing inference on the image from the phone's front-facing camera, and without requiring specialized hardware. While this offers wide potential…

计算机视觉与模式识别 · 计算机科学 2017-11-28 Matan Sela , Pingmei Xu , Junfeng He , Vidhya Navalpakkam , Dmitry Lagun

High-resolution video generation has emerged as a crucial task in computer vision, with wide-ranging applications in entertainment, simulation, and data augmentation. However, generating temporally coherent and visually realistic videos…

图像与视频处理 · 电气工程与系统科学 2025-07-08 Abhinav Sagar

Facial video re-targeting is a challenging problem aiming to modify the facial attributes of a target subject in a seamless manner by a driving monocular sequence. We leverage the 3D geometry of faces and Generative Adversarial Networks…

计算机视觉与模式识别 · 计算机科学 2021-09-29 Michail Christos Doukas , Mohammad Rami Koujan , Viktoriia Sharmanska , Anastasios Roussos

Audio-driven talking head generation aims to create vivid and realistic videos from a static portrait and speech. Existing AR-based methods rely on intermediate facial representations, which limit their expressiveness and realism.…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Yuzhe Weng , Haotian Wang , Yuanhong Yu , Jun Du , Shan He , Xiaoyan Wu , Haoran Xu