从单模态到多模态人脸深度伪造检测:进展与挑战
摘要
随着合成媒体(包括视频、音频和文本)日益难以与真实内容区分, misinformation、身份欺诈和社会操纵的风险不断上升。本综述追溯了深度伪造检测从早期单模态方法向整合音视频和文本-视频线索的高级多模态方法的演进。我们提出了结构化的检测技术分类框架,分析了从基于GAN到基于扩散模型驱动的深度伪造技术的转变,这些技术因其提高的逼真度和对检测的鲁棒性而带来新挑战。与专注于单模态检测或更早期深度伪造技术的先前综述不同,本工作报道了迄今为止最全面的研究,涵盖多模态深度伪造检测的最新进展、泛化挑战、主动防御机制以及专为支持新颖可解释性和推理任务而设计的新型数据集。我们进一步探讨了视觉语言模型(Vision-Language Models, VLMs)和多模态大语言模型(Multimodal Large Language Models, MLLMs)在强化检测对手段日益精妙深度伪造攻击鲁棒性方面的作用。通过系统性地对现有方法进行分类并识别 emerging research directions, 本综述为未来在对抗AI生成人脸伪造的进一步发展奠定了基础。相关论文的精选列表可在\href{https://github.com/qiqitao77/Comprehensive-Advances-in-Deepfake-Detection-Spanning-Diverse-Modalities}{https://github.com/qiqitao77/Awesome-Comprehensive-Deepfake-Detection}中找到。
引用
@article{arxiv.2406.06965,
title = {Evolving from Single-modal to Multi-modal Facial Deepfake Detection: Progress and Challenges},
author = {Ping Liu and Qiqi Tao and Joey Tianyi Zhou},
journal= {arXiv preprint arXiv:2406.06965},
year = {2025}
}
备注
P. Liu is with the Department of Computer Science and Engineering, University of Nevada, Reno, NV, 89512. Q. Tao and J. Zhou are with Centre for Frontier AI Research (CFAR), and Institute of High Performance Computing (IHPC), A*STAR, Singapore. J. Zhou is also with Centre for Advanced Technologies in Online Safety (CATOS), A*STAR, Singapore. J. Zhou is the corresponding author