中文

论深度伪造语音检测——呈现才是关键

音频与语音处理 2026-03-16 v2 人工智能

摘要

尽管近年来生成式 AI 的进步极大地推动了恶意音频深度伪造技术的发展,但全球针对欺骗(深度伪造)反向措施的研究却未能同步进步。本文指出,当前深度伪造数据集和研究方法导致系统难以在实际场景中泛化。其主要原因在于原始深度伪造音频与通过通信信道(如电话)呈现的深度伪造音频存在差异。我们提出一种新的数据创建框架和研究方法,使反欺骗措施在实际场景中更有效。遵循本文提出的指南,我们在更稳健且更贴近现实的实验 setup 中提高了深度伪造检测准确率 39%,在真实世界基准上提高了 57%。我们还展示了数据集的改进对深度伪造检测准确率的提升作用远超采用更大规模 SOTA 模型相对于小型模型的选择——即科学界更应投资于全面的数据收集计划,而非仅训练更大计算需求更高的模型。

引用

@article{arxiv.2509.26471,
  title  = {On Deepfake Voice Detection -- It's All in the Presentation},
  author = {Héctor Delgado and Giorgio Ramondetti and Emanuele Dalmasso and Gennady Karvitsky and Daniele Colibro and Haydar Talib},
  journal= {arXiv preprint arXiv:2509.26471},
  year   = {2026}
}

备注

ICASSP 2026. \c{opyright}IEEE Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works