面向 Roman Urdu 口语化推文的希望言检测:积极情绪在自然语言处理中的转机
摘要
希望是一种涉及对有利未来结果期待的积极情绪,而希望言则指促进乐观、韧性和支持的交流,尤其在逆境情境中。尽管希望言检测在自然语言处理(NLP)中受到关注,现有研究主要聚焦于高资源语言和标准化脚本,常忽视非正式且支持不足的形式,如 Roman Urdu。本研究鉴于目前没有针对 Roman Urdu 口语化 hope speech 进行检测的工作,首次引入精心注释的数据集,填补了面向低资源、非正式语言变体的包容性NLP研究的关键空白。本研究作出四项关键贡献:(1) 引入首个针对 Roman Urdu hope speech 的多类别注释数据集,包含 Generalized Hope、Realistic Hope、Unrealistic Hope 和 Not Hope 四类;(2) 探讨希望的心理学基础并分析其在 code-mixed Roman Urdu 中的语言学模式,以指导数据集开发;(3) 提出针对 Roman Urdu 语法和语义变异性优化的自定义注意力变换模型,通过5折交叉验证进行评估;(4) 使用t检验验证性能提升的统计显著性。所提出的模型XLM-R在交叉验证得分达到0.78,优于基线SVM(0.75)和BiLSTM(0.76),分别实现了4%和2.63%的提升。
引用
@article{arxiv.2506.21583,
title = {Hope Speech Detection in code-mixed Roman Urdu tweets: A Positive Turn in Natural Language Processing},
author = {Muhammad Ahmad and Muhammad Waqas and Ameer Hamza and Ildar Batyrshin and Grigori Sidorov},
journal= {arXiv preprint arXiv:2506.21583},
year = {2026}
}
备注
We are withdrawing this preprint because it contains initial experimental results and an early version of the manuscript. We are currently improving the methodology, conducting additional experiments, and refining the analysis. A substantially revised version will be submitted in the future