J-Guard:新闻学引导的 AI 生成新闻对抗鲁棒检测
计算与语言
2023-09-07 v1 人工智能
摘要
AI 生成文本在网络上的迅速扩散正在深刻重塑信息格局。在各种类型的 AI 生成文本中,AI 生成新闻构成了重大威胁,因为它可能成为网络虚假信息的重要来源。尽管近期已有一些工作致力于检测通用的 AI 生成文本,但考虑到其易受简单对抗攻击的脆弱性,这些方法需要更高的可靠性。此外,由于新闻写作的独特性,将这些检测方法应用于 AI 生成新闻可能会产生假阳性,从而可能损害新闻机构的声誉。为了应对这些挑战,我们利用跨学科团队的专业知识,开发了一个名为 J-Guard 的框架,该框架能够引导现有的有监督 AI 文本检测器来检测 AI 生成新闻,同时增强其对抗鲁棒性。通过融入受独特新闻属性启发的文体线索,J-Guard 有效地将真实世界的新闻报道与 AI 生成的新闻文章区分开来。我们在包括 ChatGPT (GPT3.5) 在内的大量 AI 模型生成的新闻文章上进行了实验,结果表明 J-Guard 在增强检测能力方面十分有效,且在面对对抗攻击时,其平均性能下降幅度仅为 7%。
引用
@article{arxiv.2309.03164,
title = {J-Guard: Journalism Guided Adversarially Robust Detection of AI-generated News},
author = {Tharindu Kumarage and Amrita Bhattacharjee and Djordje Padejski and Kristy Roschke and Dan Gillmor and Scott Ruston and Huan Liu and Joshua Garland},
journal= {arXiv preprint arXiv:2309.03164},
year = {2023}
}
备注
This Paper is Accepted to The 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL 2023)