English
Related papers

Related papers: Superior and Pragmatic Talking Face Generation wit…

200 papers

One-shot talking head video generation uses a source image and driving video to create a synthetic video where the source person's facial movements imitate those of the driving video. However, differences in scale between the source and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Fa-Ting Hong , Dan Xu

Generating realistic talking faces is an interesting and long-standing topic in the field of computer vision. Although significant progress has been made, it is still challenging to generate high-quality dynamic faces with personalized…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Bo Ding , Zhenfeng Fan , Shuang Yang , Shihong Xia

Talking face generation has been extensively investigated owing to its wide applicability. The two primary frameworks used for talking face generation comprise a text-driven framework, which generates synchronized speech and talking faces…

Computer Vision and Pattern Recognition · Computer Science 2023-05-22 Kentaro Mitsui , Yukiya Hono , Kei Sawada

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few…

Computer Vision and Pattern Recognition · Computer Science 2023-04-21 Shuai Shen , Wenliang Zhao , Zibin Meng , Wanhua Li , Zheng Zhu , Jie Zhou , Jiwen Lu

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu

Text-conditioned image editing has greatly benefitted from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Huanghao Yin , Shenkun Xu , Kanle Shi , Junhai Yong , Bin Wang

People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Haozhe Wu , Jia Jia , Haoyu Wang , Yishun Dou , Chao Duan , Qingshan Deng

To the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural…

Graphics · Computer Science 2021-09-27 Yuanxun Lu , Jinxiang Chai , Xun Cao

In this work, we introduce Speech-Copilot, a modular framework for instruction-oriented speech-processing tasks that minimizes human effort in toolset construction. Unlike end-to-end methods using large audio-language models, Speech-Copilot…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Chun-Yi Kuan , Chih-Kai Yang , Wei-Ping Huang , Ke-Han Lu , Hung-yi Lee

Talking face generation has gained immense popularity in the computer vision community, with various applications including AR, VR, teleconferencing, digital assistants, and avatars. Traditional methods are mainly audio-driven, which have…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Xingjian Diao , Ming Cheng , Wayner Barrios , SouYoung Jin

Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie, and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape generation.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Cuican Yu , Guansong Lu , Yihan Zeng , Jian Sun , Xiaodan Liang , Huibin Li , Zongben Xu , Songcen Xu , Wei Zhang , Hang Xu

Recent neural talking radiance field methods have shown great success in photorealistic audio-driven talking face synthesis. In this paper, we propose a novel interactive framework that utilizes human instructions to edit such implicit…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Yuqi Sun , Ruian He , Weimin Tan , Bo Yan

We present a method for generating a video of a talking face. The method takes as inputs: (i) still images of the target face, and (ii) an audio speech segment; and outputs a video of the target face lip synched with the audio. The method…

Computer Vision and Pattern Recognition · Computer Science 2017-07-19 Joon Son Chung , Amir Jamaludin , Andrew Zisserman

Talking face generation has a wide range of potential applications in the field of virtual digital humans. However, rendering high-fidelity facial video while ensuring lip synchronization is still a challenge for existing audio-driven…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Yaosen Chen , Yu Yao , Zhiqiang Li , Wei Wang , Yanru Zhang , Han Yang , Xuming Wen

Speech-driven animation has gained significant traction in recent years, with current methods achieving near-photorealistic results. However, the field remains underexplored regarding non-verbal communication despite evidence demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Antoni Bigata Casademunt , Rodrigo Mira , Nikita Drobyshev , Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Deepfake is a technology dedicated to creating highly realistic facial images and videos under specific conditions, which has significant application potential in fields such as entertainment, movie production, digital human creation, to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Gan Pei , Jiangning Zhang , Menghan Hu , Zhenyu Zhang , Chengjie Wang , Yunsheng Wu , Guangtao Zhai , Jian Yang , Dacheng Tao

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

Face recognition datasets are often collected by crawling Internet and without individuals' consents, raising ethical and privacy concerns. Generating synthetic datasets for training face recognition models has emerged as a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Hatef Otroshi Shahreza , Sébastien Marcel

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Zhenhui Ye , Tianyun Zhong , Yi Ren , Jiaqi Yang , Weichuang Li , Jiawei Huang , Ziyue Jiang , Jinzheng He , Rongjie Huang , Jinglin Liu , Chen Zhang , Xiang Yin , Zejun Ma , Zhou Zhao

Cross-lingual synthesis can be defined as the task of letting a speaker generate fluent synthetic speech in another language. This is a challenging task, and resulting speech can suffer from reduced naturalness, accented speech, and/or loss…

Sound · Computer Science 2022-04-04 Marcel de Korte , Jaebok Kim , Aki Kunikoshi , Adaeze Adigwe , Esther Klabbers
‹ Prev 1 4 5 6 7 8 10 Next ›