基于GRPO的强化学习微调是否改善二进制语音深度伪造检测的泛化能力?
音频与语音处理
2026-03-04 v1
摘要
构建能够对未见攻击保持泛化的语音深度伪造检测模型仍是一个具有挑战性的难题。尽管该领域已转向使用语音基础模型进行预训练和微调,但大多数方法仅依赖监督微调(SFT)。鼓舞人心的是,近年来在大型语言模型领域中,强化学习(RL)被用于模型微调。我们因此研究了GRPO(Group Relative Policy Optimization)的影响。使用多个检测器和测试集进行的实验结果表明,纯粹的GRPO-based微调在出域测试集上提高了性能,同时保持了目标域测试数据的性能。该方法优于SFT-only和混合方案。我们的消融研究进一步表明,GRPO中的负奖励可能是这一改进的关键因素。
引用
@article{arxiv.2603.02914,
title = {Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?},
author = {Xin Wang and Ge Wanying and Junichi Yamagishi},
journal= {arXiv preprint arXiv:2603.02914},
year = {2026}
}
备注
Submitted to Interspeech 2026; put on arxiv based on requirement of paper open-access rule; quote from Interspeech: "Interspeech no longer enforces an anonymity period for submissions. While uploading a version online is permitted, your official submission to Interspeech must not contain any author-identifying information"