English
Related papers

Related papers: DH-FaceVid-1K: A Large-Scale High-Quality Dataset …

200 papers

The proliferation of generative video technologies has intensified the need for reliable methods to detect and characterize synthetic media. To address this challenge, we organized the \href{https://safe-video-2025.dsri.org}{SAFE: Synthetic…

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Yuhan Wang , Fangzhou Hong , Shuai Yang , Liming Jiang , Wayne Wu , Chen Change Loy

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Weizhi Zhong , Chaowei Fang , Yinqi Cai , Pengxu Wei , Gangming Zhao , Liang Lin , Guanbin Li

Face recognition datasets are often collected by crawling Internet and without individuals' consents, raising ethical and privacy concerns. Generating synthetic datasets for training face recognition models has emerged as a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Hatef Otroshi Shahreza , Sébastien Marcel

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ziqi Zhang , Cheng Deng

In this paper, we introduce a preview of the Deepfakes Detection Challenge (DFDC) dataset consisting of 5K videos featuring two facial modification algorithms. A data collection campaign has been carried out where participating actors have…

Computer Vision and Pattern Recognition · Computer Science 2019-10-25 Brian Dolhansky , Russ Howes , Ben Pflaum , Nicole Baram , Cristian Canton Ferrer

Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Liyuan Cui , Xiaogang Xu , Wenqi Dong , Zesong Yang , Hujun Bao , Zhaopeng Cui

One-shot talking face generation aims at synthesizing a high-quality talking face video from an arbitrary portrait image, driven by a video or an audio segment. One challenging quality factor is the resolution of the output video: higher…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 Fei Yin , Yong Zhang , Xiaodong Cun , Mingdeng Cao , Yanbo Fan , Xuan Wang , Qingyan Bai , Baoyuan Wu , Jue Wang , Yujiu Yang

As AI-generated video becomes increasingly pervasive across media platforms, the ability to reliably distinguish synthetic content from authentic footage has become both urgent and essential. Existing approaches have primarily treated this…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Yifeng Gao , Yifan Ding , Hongyu Su , Juncheng Li , Yunhan Zhao , Lin Luo , Zixing Chen , Li Wang , Xin Wang , Yixu Wang , Xingjun Ma , Yu-Gang Jiang

The rapid advancement in generative artificial intelligence have enabled the creation of 3D human faces (HFs) for applications including media production, virtual reality, security, healthcare, and game development, etc. However, assessing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Woo Yi Yang , Jiarui Wang , Sijing Wu , Huiyu Duan , Yuxin Zhu , Liu Yang , Kang Fu , Guangtao Zhai , Xiongkuo Min

While recent research has made significant progress in speech-driven talking face generation, the quality of the generated video still lags behind that of real recordings. One reason for this is the use of handcrafted intermediate…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Chenpeng Du , Qi Chen , Tianyu He , Xu Tan , Xie Chen , Kai Yu , Sheng Zhao , Jiang Bian

Recent Multi-modal Large Language Models (MLLMs) have made great progress in video understanding. However, their performance on videos involving human actions is still limited by the lack of high-quality data. To address this, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Xiao Wang , Jingyun Hua , Weihong Lin , Yuanxing Zhang , Fuzheng Zhang , Jianlong Wu , Di Zhang , Liqiang Nie

This paper presents a generative adversarial learning-based human upper body video synthesis approach to generate an upper body video of target person that is consistent with the body motion, face expression, and pose of the person in…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Zhaoxiang Liu , Huan Hu , Zipeng Wang , Kai Wang , Jinqiang Bai , Shiguo Lian

Scalable learning of humanoid robots is crucial for their deployment in real-world applications. While traditional approaches primarily rely on reinforcement learning or teleoperation to achieve whole-body control, they are often limited by…

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Zhengwentai Sun , Keru Zheng , Chenghong Li , Hongjie Liao , Xihe Yang , Heyuan Li , Yihao Zhi , Shuliang Ning , Shuguang Cui , Xiaoguang Han

Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial movement over the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Yang Song , Jingwen Zhu , Dawei Li , Xiaolong Wang , Hairong Qi

Numerous synthesized videos from generative models, especially human-centric ones that simulate realistic human actions, pose significant threats to human information security and authenticity. While progress has been made in binary forgery…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Chang Liu , Yunfan Ye , Fan Zhang , Qingyang Zhou , Yuchuan Luo , Zhiping Cai

This research explores the positive application of deepfake technology for upper body generation, specifically sign language for the Deaf and Hard of Hearing (DHoH) community. Given the complexity of sign language and the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Shahzeb Naeem , Muhammad Riyyan Khan , Usman Tariq , Abhinav Dhall , Carlos Ivan Colon , Hasan Al-Nashash