针对基于大语言模型的分词不一致性问题进行 steganography 和 watermarking
声音
2025-12-22 v2 机器学习
摘要
大语言模型显著提升了文本生成的能力和效率。一方面,它们提高了文本基于 steganography 的质量;另一方面,也凸显了 watermarking 作为防止恶意滥用的重要保障。本研究聚焦于 steganography 和 watermarking 中 Alice 和 Bob 之间的分词不一致性(TI),其中 TI 可能削弱鲁棒性。我们的调查揭示,导致 TI 的 problematic token 具有两个关键特征:稀有性和暂时性。基于这些发现,我们提出了针对 TI 消除的两个定制解决方案:针对 steganography 的逐步验证方法和针对 watermarking 的事后回滚方法。实验表明:(1)与 steganography 中传统歧义消除方法相比,直接解决 TI 可在流畅性、不可见性和抗 steganalysis 能力方面获得改进;(2)对于 watermarking,解决 TI 可提高检测性和抗攻击鲁棒性。
引用
@article{arxiv.2508.20717,
title = {Unified Acoustic Representations for Screening Neurological and Respiratory Pathologies from Voice},
author = {Ran Piao and Yuan Lu and Hareld Kemps and Tong Xia and Aaqib Saeed},
journal= {arXiv preprint arXiv:2508.20717},
year = {2025}
}