English
Related papers

Related papers: An Audio-Visual Attention Based Multimodal Network…

200 papers

Significant advancements made in the generation of deepfakes have caused security and privacy issues. Attackers can easily impersonate a person's identity in an image by replacing his face with the target person's face. Moreover, a new…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Hasam Khalid , Minha Kim , Shahroz Tariq , Simon S. Woo

The widespread application of AIGC contents has brought not only unprecedented opportunities, but also potential security concerns, e.g., audio-visual deepfakes. Therefore, it is of great importance to develop an effective and generalizable…

Multimedia · Computer Science 2025-11-25 Fan Nie , Jiangqun Ni , Jian Zhang , Bin Zhang , Weizhe Zhang , Bin Li

With the continuous improvements of deepfake methods, forgery messages have transitioned from single-modality to multi-modal fusion, posing new challenges for existing forgery detection algorithms. In this paper, we propose AVT2-DWF, the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Rui Wang , Dengpan Ye , Long Tang , Yunming Zhang , Jiacheng Deng

The internet is filled with fake face images and videos synthesized by deep generative models. These realistic DeepFakes pose a challenge to determine the authenticity of multimedia content. As countermeasures, artifact-based detection…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Gaojian Wang , Qian Jiang , Xin Jin , Xiaohui Cui

Deep-learning-based technologies such as deepfakes ones have been attracting widespread attention in both society and academia, particularly ones used to synthesize forged face images. These automatic and professional-skill-free face…

Computer Vision and Pattern Recognition · Computer Science 2022-12-08 YuYang Sun , ZhiYong Zhang , Isao Echizen , Huy H. Nguyen , ChangZhen Qiu , Lu Sun

This work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Madhav Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Face forgery by deepfake is widely spread over the internet and has raised severe societal concerns. Recently, how to detect such forgery contents has become a hot research topic and many deepfake detection methods have been proposed. Most…

Computer Vision and Pattern Recognition · Computer Science 2021-03-09 Hanqing Zhao , Wenbo Zhou , Dongdong Chen , Tianyi Wei , Weiming Zhang , Nenghai Yu

This paper reviews the state-of-the-art in deepfake generation and detection, focusing on modern deep learning technologies and tools based on the latest scientific advancements. The rise of deepfakes, leveraging techniques like Variational…

Cryptography and Security · Computer Science 2025-01-14 Arash Dehghani , Hossein Saberi

Audio and visual signals typically occur simultaneously, and humans possess an innate ability to correlate and synchronize information from these two modalities. Recently, a challenging problem known as Audio-Visual Segmentation (AVS) has…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Yuxuan Wang , Jinchao Zhu , Feng Dong , Shuyue Zhu

Existing methods on audio-visual deepfake detection mainly focus on high-level features for modeling inconsistencies between audio and visual data. As a result, these approaches usually overlook finer audio-visual artifacts, which are…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Marcella Astrid , Enjie Ghorbel , Djamila Aouada

Media forensics has attracted a lot of attention in the last years in part due to the increasing concerns around DeepFakes. Since the initial DeepFake databases from the 1st generation such as UADFV and FaceForensics++ up to the latest…

Computer Vision and Pattern Recognition · Computer Science 2020-11-13 Ruben Tolosana , Sergio Romero-Tapiador , Julian Fierrez , Ruben Vera-Rodriguez

One of the most pressing challenges for the detection of face-manipulated videos is generalising to forgery methods not seen during training while remaining effective under common corruptions such as compression. In this paper, we examine…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Alexandros Haliassos , Rodrigo Mira , Stavros Petridis , Maja Pantic

With the rapid advancement of deep learning in image generation, facial forgery techniques have achieved unprecedented realism, posing serious threats to cybersecurity and information authenticity. Most existing deepfake detection…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Haotian Wu , Yue Cheng , Shan Bian

The threat of Audio-Video (AV) forgery is rapidly evolving beyond human-centric deepfakes to include more diverse manipulations across complex natural scenes. However, existing benchmarks are still confined to DeepFake-based forgeries and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Shuhan Xia , Peipei Li , Xuannan Liu , Dongsen Zhang , Xinyu Guo , Zekun Li

AI-created face-swap videos, commonly known as Deepfakes, have attracted wide attention as powerful impersonation attacks. Existing research on Deepfakes mostly focuses on binary detection to distinguish between real and fake videos.…

Computer Vision and Pattern Recognition · Computer Science 2022-03-01 Shan Jia , Xin Li , Siwei Lyu

A major challenge in DeepFake forgery detection is that state-of-the-art algorithms are mostly trained to detect a specific fake method. As a result, these approaches show poor generalization across different types of facial manipulations,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Davide Cozzolino , Andreas Rössler , Justus Thies , Matthias Nießner , Luisa Verdoliva

3D face alignment of monocular images is a crucial process in the recognition of faces with disguise.3D face reconstruction facilitated by alignment can restore the face structure which is helpful in detcting disguise interference.This…

Computer Vision and Pattern Recognition · Computer Science 2019-09-02 Lei Jiang Xiao-Jun Wu Josef Kittler

Speech enhancement can potentially benefit from the visual information from the target speaker, such as lip movement and facial expressions, because the visual aspect of speech is essentially unaffected by acoustic environment. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-24 Xinmeng Xu , Jianjun Hao

The misuse of advanced generative AI models has resulted in the widespread proliferation of falsified data, particularly forged human-centric audiovisual content, which poses substantial societal risks (e.g., financial fraud and social…

Cryptography and Security · Computer Science 2025-10-28 Kangran Zhao , Yupeng Chen , Xiaoyu Zhang , Yize Chen , Weinan Guan , Baicheng Chen , Chengzhe Sun , Soumyya Kanti Datta , Qingshan Liu , Siwei Lyu , Baoyuan Wu

In this work, we describe a new deep learning based method that can effectively distinguish AI-generated fake videos (referred to as {\em DeepFake} videos hereafter) from real videos. Our method is based on the observations that current…

Computer Vision and Pattern Recognition · Computer Science 2019-05-23 Yuezun Li , Siwei Lyu