English
Related papers

Related papers: AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visu…

200 papers

This paper reviews the state-of-the-art in deepfake generation and detection, focusing on modern deep learning technologies and tools based on the latest scientific advancements. The rise of deepfakes, leveraging techniques like Variational…

Cryptography and Security · Computer Science 2025-01-14 Arash Dehghani , Hossein Saberi

As synthetic media, including video, audio, and text, become increasingly indistinguishable from real content, the risks of misinformation, identity fraud, and social manipulation escalate. This survey traces the evolution of deepfake…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Ping Liu , Qiqi Tao , Joey Tianyi Zhou

With the rapid progress of recent years, techniques that generate and manipulate multimedia content can now guarantee a very advanced level of realism. The boundary between real and synthetic media has become very thin. On the one hand,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-19 Luisa Verdoliva

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Tiantian Geng , Teng Wang , Jinming Duan , Runmin Cong , Feng Zheng

Multimodal generative models are rapidly evolving, leading to a surge in the generation of realistic video and audio that offers exciting possibilities but also serious risks. Deepfake videos, which can convincingly impersonate individuals,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Hannah Lee , Changyeon Lee , Kevin Farhat , Lin Qiu , Steve Geluso , Aerin Kim , Oren Etzioni

Recent advances in Artificial Intelligence Generated Content have led to highly realistic synthetic videos, particularly in human-centric scenarios involving speech, gestures, and full-body motion, posing serious threats to information…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Zhipei Xu , Xuanyu Zhang , Qing Huang , Xing Zhou , Jian Zhang

Current text-to-speech algorithms produce realistic fakes of human voices, making deepfake detection a much-needed area of research. While researchers have presented various techniques for detecting audio spoofs, it is often unclear exactly…

Automatic Deception Detection has been a hot research topic for a long time, using machine learning and deep learning to automatically detect deception, brings new light to this old field. In this paper, we proposed a voting-based method…

Machine Learning · Computer Science 2024-03-18 Lana Touma , Mohammad Al Horani , Manar Tailouni , Anas Dahabiah , Khloud Al Jallad

In recent years, detecting fake multimodal content on social media has drawn increasing attention. Two major forms of deception dominate: human-crafted misinformation (e.g., rumors and misleading posts) and AI-generated content produced by…

Artificial Intelligence · Computer Science 2025-10-17 Haiyang Li , Yaxiong Wang , Shengeng Tang , Lianwei Wu , Lechao Cheng , Zhun Zhong

The tremendous recent advances in generative artificial intelligence techniques have led to significant successes and promise in a wide range of different applications ranging from conversational agents and textual content generation to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Hossein Aboutalebi , Dayou Mao , Rongqi Fan , Carol Xu , Chris He , Alexander Wong

The rise of deepfake technology brings forth new questions about the authenticity of various forms of media found online today. Videos and images generated by artificial intelligence (AI) have become increasingly more difficult to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Benjamin Carter , Nathan Dilla , Micheal Callahan , Atuhaire Ambala

Deeplearning has been used to solve complex problems in various domains. As it advances, it also creates applications which become a major threat to our privacy, security and even to our Democracy. Such an application which is being…

Computer Vision and Pattern Recognition · Computer Science 2020-09-17 Rahul U , Ragul M , Raja Vignesh K , Tejeswinee K

The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Hashim Ali , Surya Subramani , Lekha Bollinani , Nithin Sai Adupa , Sali El-Loh , Hafiz Malik

The recent renaissance in generative models, driven primarily by the advent of diffusion models and iterative improvement in GAN methods, has enabled many creative applications. However, each advancement is also accompanied by a rise in the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-28 Sanjay Saha , Rashindrie Perera , Sachith Seneviratne , Tamasha Malepathirana , Sanka Rasnayaka , Deshani Geethika , Terence Sim , Saman Halgamuge

Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations, or unrealistic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Tianxiao Li , Zhenglin Huang , Haiquan Wen , Yiwei He , Xinze Li , Bingyu Zhu , Wuhui Duan , Congang Chen , Zeyu Fu , Yi Dong , Baoyuan Wu , Jason Li , Guangliang Cheng

With the rapid advancement of generative models, the realism of AI-generated images has significantly improved, posing critical challenges for verifying digital content authenticity. Current deepfake detection methods often depend on…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Jiarui Wang , Huiyu Duan , Juntong Wang , Ziheng Jia , Woo Yi Yang , Xiaorong Zhu , Yu Zhao , Jiaying Qian , Yuke Xing , Guangtao Zhai , Xiongkuo Min

With the prevalence of artificial intelligence (AI)-generated content, such as audio deepfakes, a large body of recent work has focused on developing deepfake detection techniques. However, most models are evaluated on a narrow set of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-29 Yi Zhu , Heitor R. Guimarães , Arthur Pimentel , Tiago Falk

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant progress in this domain. This paper surveys 69 datasets used to…

Sound · Computer Science 2025-02-10 Gijs Wijngaard , Elia Formisano , Michele Esposito , Michel Dumontier

Detecting digital face manipulation in images and video has attracted extensive attention due to the potential risk to public trust. To counteract the malicious usage of such techniques, deep learning-based deepfake detection methods have…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Yuhang Lu , Touradj Ebrahimi