揭示宣传:人工标注与机器分类的风格特征分析
计算与语言
2024-02-27 v3 人工智能
机器学习
摘要
本文探讨了宣传语言及其风格特征。本文提出了PPN数据集,即Propagandist Pseudo-News (虚假新闻),该数据集由来自专家机构识别为宣传来源网站提取的新闻文章组成,涵盖多源、多语言和多模态。该数据集的有限样本随机混合来自常规法国新闻,URL被遮盖,以进行人工标注实验,使用11个 distinct标签。结果表明,人工标注者能够可靠地在每个标签上区分两种新闻类型。我们提出了不同的NLP技术来识别注释者使用的线索,并与机器分类进行比较。包括VAGO分析器用于衡量话语模糊性和主观性、TF-IDF作为基线,以及四个分类器:两个基于RoBERTa的模型、使用语法的CATS,以及一个结合语法和语义特征的XGBoost。
引用
@article{arxiv.2402.03780,
title = {Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification},
author = {Géraud Faye and Benjamin Icard and Morgane Casanova and Julien Chanson and François Maine and François Bancilhon and Guillaume Gadek and Guillaume Gravier and Paul Égré},
journal= {arXiv preprint arXiv:2402.03780},
year = {2024}
}
备注
Paper to appear in the EACL 2024 Proceedings of the Third Workshop on Understanding Implicit and Underspecified Language (UnImplicit 2024)