中文
相关论文

相关论文: MAVOS-DD: Multilingual Audio-Video Open-Set Deepfa…

200 篇论文

This paper presents a system for detecting fake audio-visual content (i.e., video deepfake), developed for Track 2 of the DDL Challenge. The proposed system employs a two-stage framework, comprising unimodal detection and multimodal score…

多媒体 · 计算机科学 2026-02-03 Qingcao Li , Miao He , Liang Yi , Qing Wen , Yitao Zhang , Hongshuo Jin , Peng Cheng , Zhongjie Ba , Li Lu , Kui Ren

In recent years, the explosive advancement of deepfake technology has posed a critical and escalating threat to public security: diffusion-based digital human generation. Unlike traditional face manipulation methods, such models can…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Jiaxin Liu , Jia Wang , Saihui Hou , Min Ren , Huijia Wu , Long Ma , Renwang Pei , Zhaofeng He

Deepfakes are increasingly realistic and easy to produce, raising concerns about the reliability of human judgments in misinformation settings. We study audiovisual deepfake detection by measuring how consistently crowd workers distinguish…

信息检索 · 计算机科学 2026-05-07 Michael Soprano , Andrea Cioci , Stefano Mizzaro

Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be…

声音 · 计算机科学 2026-02-06 Xueping Zhang , Han Yin , Yang Xiao , Lin Zhang , Ting Dang , Rohan Kumar Das , Ming Li

Talking face generation (TFG) allows for producing lifelike talking videos of any character using only facial images and accompanying text. Abuse of this technology could pose significant risks to society, creating the urgent need for…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Xiaocan Chen , Qilin Yin , Jiarui Liu , Wei Lu , Xiangyang Luo , Jiantao Zhou

Multimodal generative models are rapidly evolving, leading to a surge in the generation of realistic video and audio that offers exciting possibilities but also serious risks. Deepfake videos, which can convincingly impersonate individuals,…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Hannah Lee , Changyeon Lee , Kevin Farhat , Lin Qiu , Steve Geluso , Aerin Kim , Oren Etzioni

There have been emerging a number of benchmarks and techniques for the detection of deepfakes. However, very few works study the detection of incrementally appearing deepfakes in the real-world scenarios. To simulate the wild scenes, this…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Chuqiao Li , Zhiwu Huang , Danda Pani Paudel , Yabin Wang , Mohamad Shahbazi , Xiaopeng Hong , Luc Van Gool

Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving…

声音 · 计算机科学 2026-02-12 Qizhou Wang , Hanxun Huang , Guansong Pang , Sarah Erfani , Christopher Leckie

It is becoming increasingly easy to automatically replace a face of one person in a video with the face of another person by using a pre-trained generative adversarial network (GAN). Recent public scandals, e.g., the faces of celebrities…

计算机视觉与模式识别 · 计算机科学 2018-12-21 Pavel Korshunov , Sebastien Marcel

The increasing use of synthetic media, particularly deepfakes, is an emerging challenge for digital content verification. Although recent studies use both audio and visual information, most integrate these cues within a single model, which…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Sayeem Been Zaman , Wasimul Karim , Arefin Ittesafun Abian , Reem E. Mohamed , Md Rafiqul Islam , Asif Karim , Sami Azam

Video synthesis methods rapidly improved in recent years, allowing easy creation of synthetic humans. This poses a problem, especially in the era of social media, as synthetic videos of speaking humans can be used to spread misinformation…

计算机视觉与模式识别 · 计算机科学 2024-02-09 Gil Knafo , Ohad Fried

Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further…

声音 · 计算机科学 2026-03-09 Daixian Li , Jun Xue , Yanzhen Ren , Zhuolin Yi , Yihuan Huang , Guanxiang Feng , Yi Chai

Fake audio detection is a growing concern and some relevant datasets have been designed for research. However, there is no standard public Chinese dataset under complex conditions.In this paper, we aim to fill in the gap and design a…

声音 · 计算机科学 2023-07-19 Haoxin Ma , Jiangyan Yi , Chenglong Wang , Xinrui Yan , Jianhua Tao , Tao Wang , Shiming Wang , Ruibo Fu

We propose detection of deepfake videos based on the dissimilarity between the audio and visual modalities, termed as the Modality Dissonance Score (MDS). We hypothesize that manipulation of either modality will lead to dis-harmony between…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Komal Chugh , Parul Gupta , Abhinav Dhall , Ramanathan Subramanian

The availability of smart devices leads to an exponential increase in multimedia content. However, advancements in deep learning have also enabled the creation of highly sophisticated Deepfake content, including speech Deepfakes, which pose…

声音 · 计算机科学 2025-07-16 Menglu Li , Yasaman Ahmadiadli , Xiao-Ping Zhang

The rapid evolution of generative AI has increased the threat of realistic audio-visual deepfakes, demanding robust detection methods. Existing solutions primarily address unimodal (audio or visual) forgeries but struggle with multimodal…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Jian Wang , Baoyuan Wu , Li Liu , Qingshan Liu

Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on curated synthetic forgeries. Such synthetic dependence can…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Sahibzada Adil Shahzad , Ammarah Hashmi , Junichi Yamagishi , Yusuke Yasuda , Yu Tsao , Chia-Wen Lin , Yan-Tsung Peng , Hsin-Min Wang

Advancements in audio deepfake technology offers benefits like AI assistants, better accessibility for speech impairments, and enhanced entertainment. However, it also poses significant risks to security, privacy, and trust in digital…

声音 · 计算机科学 2025-06-27 Abhay Kumar , Kunal Verma , Omkar More

In recent years, the abuse of a face swap technique called deepfake has raised enormous public concerns. So far, a large number of deepfake videos (known as "deepfakes") have been crafted and uploaded to the internet, calling for effective…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Bojia Zi , Minghao Chang , Jingjing Chen , Xingjun Ma , Yu-Gang Jiang

In this paper, we present the Global Multimedia Deepfake Detection held concurrently with the Inclusion 2024. Our Multimedia Deepfake Detection aims to detect automatic image and audio-video manipulations including but not limited to…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Yi Zhang , Weize Gao , Changtao Miao , Man Luo , Jianshu Li , Wenzhong Deng , Zhe Li , Bingyu Hu , Weibin Yao , Yunfeng Diao , Wenbo Zhou , Tao Gong , Qi Chu