English
Related papers

Related papers: FTFDNet: Learning to Detect Talking Face Video Man…

200 papers

Face authentication usually utilizes deep learning models to verify users with high recognition accuracy. However, face authentication systems are vulnerable to various attacks that cheat the models by manipulating the digital counterparts…

Cryptography and Security · Computer Science 2021-06-16 Man Zhou , Qian Wang , Qi Li , Peipei Jiang , Jingxiao Yang , Chao Shen , Cong Wang , Shouhong Ding

The recent realistic creation and dissemination of so-called deepfakes poses a serious threat to social life, civil rest, and law. Celebrity defaming, election manipulation, and deepfakes as evidence in court of law are few potential…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Muhammad Umar Farooq , Ali Javed , Khalid Mahmood Malik , Muhammad Anas Raza

The spread of misinformation through synthetically generated yet realistic images and videos has become a significant problem, calling for robust manipulation detection methods. Despite the predominant effort of detecting face manipulation…

Computer Vision and Pattern Recognition · Computer Science 2019-05-17 Ekraam Sabir , Jiaxin Cheng , Ayush Jaiswal , Wael AbdAlmageed , Iacopo Masi , Prem Natarajan

Audio and visual signals typically occur simultaneously, and humans possess an innate ability to correlate and synchronize information from these two modalities. Recently, a challenging problem known as Audio-Visual Segmentation (AVS) has…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Yuxuan Wang , Jinchao Zhu , Feng Dong , Shuyue Zhu

Multimodal generative models are rapidly evolving, leading to a surge in the generation of realistic video and audio that offers exciting possibilities but also serious risks. Deepfake videos, which can convincingly impersonate individuals,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Hannah Lee , Changyeon Lee , Kevin Farhat , Lin Qiu , Steve Geluso , Aerin Kim , Oren Etzioni

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Recent advances in diffusion-based lip-syncing generative models have demonstrated their ability to produce highly synchronized talking face videos for visual dubbing. Although these models excel at lip synchronization, they often struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yanyu Zhu , Lichen Bai , Jintao Xu , Hai-tao Zheng

In recent years, the abuse of a face swap technique called deepfake has raised enormous public concerns. So far, a large number of deepfake videos (known as "deepfakes") have been crafted and uploaded to the internet, calling for effective…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Bojia Zi , Minghao Chang , Jingjing Chen , Xingjun Ma , Yu-Gang Jiang

Facial Expression Recognition (FER) in the wild is extremely challenging due to occlusions, variant head poses, face deformation and motion blur under unconstrained conditions. Although substantial progresses have been made in automatic FER…

Computer Vision and Pattern Recognition · Computer Science 2022-05-12 Fuyan Ma , Bin Sun , Shutao Li

Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to the audio, has…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Rongliang Wu , Yingchen Yu , Fangneng Zhan , Jiahui Zhang , Xiaoqin Zhang , Shijian Lu

Face forgery has attracted increasing attention in recent applications of computer vision. Existing detection techniques using the two-branch framework benefit a lot from a frequency perspective, yet are restricted by their fixed frequency…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Neng Wang , Yang Bai , Kun Yu , Yong Jiang , Shu-tao Xia , Yan Wang

With the rapid development of face recognition (FR) systems, the privacy of face images on social media is facing severe challenges due to the abuse of unauthorized FR systems. Some studies utilize adversarial attack techniques to defend…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Yuhao Sun , Lingyun Yu , Hongtao Xie , Jiaming Li , Yongdong Zhang

News media, particularly video-based platforms, have become deeply embed-ded in daily life, concurrently amplifying the risks of misinformation dissem-ination. Consequently, multimodal fake news detection has garnered signifi-cant research…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Yihao Wang , Zhong Qian , Peifeng Li

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

The widespread emergence of face-swap Deepfake videos poses growing risks to digital security, privacy, and media integrity, necessitating effective forensic tools for identifying the source of such manipulations. Although most prior…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Wasim Ahmad , Yan-Tsung Peng , Yuan-Hao Chang

With the rapid development of deep learning techniques, the generation and counterfeiting of multimedia material are becoming increasingly straightforward to perform. At the same time, sharing fake content on the web has become so simple…

Multimedia · Computer Science 2022-09-19 Davide Salvi , Brian Hosler , Paolo Bestagini , Matthew C. Stamm , Stefano Tubaro

In recent years, neural network-based Wake Word Spotting achieves good performance on clean audio samples but struggles in noisy environments. Audio-Visual Wake Word Spotting (AVWWS) receives lots of attention because visual lip movement…

Sound · Computer Science 2024-03-05 Haoxu Wang , Ming Cheng , Qiang Fu , Ming Li

Advanced manipulation techniques have provided criminals with opportunities to make social panic or gain illicit profits through the generation of deceptive media, such as forged face images. In response, various deepfake detection methods…

Computer Vision and Pattern Recognition · Computer Science 2023-07-07 Ruiyang Xia , Decheng Liu , Jie Li , Lin Yuan , Nannan Wang , Xinbo Gao

Deepfake (DF) attacks pose a growing threat as generative models become increasingly advanced. However, our study reveals that existing DF datasets fail to deceive human perception, unlike real DF attacks that influence public discourse. It…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Oguzhan Baser , Ahmet Ege Tanriverdi , Sriram Vishwanath , Sandeep P. Chinchali

The surge of highly realistic synthetic videos produced by contemporary generative systems has significantly increased the risk of malicious use, challenging both humans and existing detectors. Against this backdrop, we take a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Youngseo Kim , Kwan Yun , Seokhyeon Hong , Sihun Cha , Colette Suhjung Koo , Junyong Noh