English
Related papers

Related papers: SyncBreaker:Stage-Aware Multimodal Adversarial Att…

200 papers

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Hui Fu , Zeqing Wang , Ke Gong , Keze Wang , Tianshui Chen , Haojie Li , Haifeng Zeng , Wenxiong Kang

Along with the widespread use of face recognition systems, their vulnerability has become highlighted. While existing face anti-spoofing methods can be generalized between attack types, generic solutions are still challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Kaicheng Li , Hongyu Yang , Binghui Chen , Pengyu Li , Biao Wang , Di Huang

Speech-driven facial animation aims to synthesize lip-synchronized 3D talking faces following the given speech signal. Prior methods to this task mostly focus on pursuing realism with deterministic systems, yet characterizing the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Chunzhi Gu , Shigeru Kuriyama , Katsuya Hotta

Audio watermarking has been widely applied in copyright protection and source tracing. However, due to the inherent characteristics of audio signals, watermark localization and resistance to desynchronization attacks remain significant…

Cryptography and Security · Computer Science 2025-09-03 Zhenliang Gan , Xiaoxiao Hu , Sheng Li , Zhenxing Qian , Xinpeng Zhang

Recent advances have demonstrated compelling capabilities in synthesizing real individuals into generated videos, reflecting the growing demand for identity-aware content creation. Nevertheless, an openly accessible framework enabling…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Yingjie Chen , Shilun Lin , Cai Xing , Binxin Yang , Long Zhou , Qixin Yan , Wenjing Wang , Dingming Liu , Hao Liu , Chen Li , Jing Lyu

Speech-driven 3D facial animation has been an attractive task in both academia and industry. Traditional methods mostly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Peng Chen , Xiaobao Wei , Ming Lu , Yitong Zhu , Naiming Yao , Xingyu Xiao , Hui Chen

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

Generating realistic talking faces is a complex and widely discussed task with numerous applications. In this paper, we present DiffTalker, a novel model designed to generate lifelike talking faces through audio and landmark co-driving.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Zipeng Qi , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait. Unlike existing models that primarily focus on verbal cues such as lip synchronization and fail to capture the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Tao Liu , Feilong Chen , Shuai Fan , Chenpeng Du , Qi Chen , Xie Chen , Kai Yu

Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrupt the forgery…

Sound · Computer Science 2025-12-10 Qianyue Hu , Junyan Wu , Wei Lu , Xiangyang Luo

The field of portrait image animation, driven by speech audio input, has experienced significant advancements in the generation of realistic and dynamic portraits. This research delves into the complexities of synchronizing facial movements…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Mingwang Xu , Hui Li , Qingkun Su , Hanlin Shang , Liwei Zhang , Ce Liu , Jingdong Wang , Yao Yao , Siyu Zhu

Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video and audio while balancing appropriateness, realism, and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Jiaming Li , Sheng Wang , Xin Wang , Yitao Zhu , Honglin Xiong , Zixu Zhuang , Qian Wang

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for applications that demand…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Yujian Liu , Shidang Xu , Jing Guo , Dingbin Wang , Zairan Wang , Xianfeng Tan , Xiaoli Liu

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Ziqiao Peng , Yi Chen , Yifeng Ma , Guozhen Zhang , Zhiyao Sun , Zixiang Zhou , Youliang Zhang , Zhengguang Zhou , Zhaoxin Fan , Hongyan Liu , Yuan Zhou , Qinglin Lu , Jun He

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

Real-time speech-driven 3D facial animation has been attractive in academia and industry. Traditional methods mainly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Peng Chen , Xiaobao Wei , Ming Lu , Hui Chen , Feng Tian

Talking face generation has historically struggled to produce head movements and natural facial expressions without guidance from additional reference videos. Recent developments in diffusion-based generative models allow for more realistic…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Michał Stypułkowski , Konstantinos Vougioukas , Sen He , Maciej Zięba , Stavros Petridis , Maja Pantic

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in virtual interaction…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Xi Liu , Ying Guo , Cheng Zhen , Tong Li , Yingying Ao , Pengfei Yan

This work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Madhav Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar
‹ Prev 1 4 5 6 7 8 10 Next ›