中文
相关论文

相关论文: Multilingual Twitter Corpus and Baselines for Eval…

200 篇论文

Many methods have been used to recognize author personality traits from text, typically combining linguistic feature engineering with shallow learning models, e.g. linear regression or Support Vector Machines. This work uses…

计算与语言 · 计算机科学 2016-10-17 Fei Liu , Julien Perez , Scott Nowson

Despite impressive advancements in multilingual corpora collection and model training, developing large-scale deployments of multilingual models still presents a significant challenge. This is particularly true for language tasks that are…

The fast spread of hate speech on social media impacts the Internet environment and our society by increasing prejudice and hurting people. Detecting hate speech has aroused broad attention in the field of natural language processing.…

计算与语言 · 计算机科学 2023-07-13 Junyu Lu , Hongfei Lin , Xiaokun Zhang , Zhaoqing Li , Tongyue Zhang , Linlin Zong , Fenglong Ma , Bo Xu

Social media platforms provide users the freedom of expression and a medium to exchange information and express diverse opinions. Unfortunately, this has also resulted in the growth of abusive content with the purpose of discriminating…

计算与语言 · 计算机科学 2021-07-01 Sohail Akhtar , Valerio Basile , Viviana Patti

Due to the wide adoption of social media platforms like Facebook, Twitter, etc., there is an emerging need of detecting online posts that can go against the community acceptance standards. The hostility detection task has been well explored…

计算与语言 · 计算机科学 2021-01-14 Arkadipta De , Venkatesh E , Kaushal Kumar Maurya , Maunendra Sankar Desarkar

Bias in textual data can lead to skewed interpretations and outcomes when the data is used. These biases could perpetuate stereotypes, discrimination, or other forms of unfair treatment. An algorithm trained on biased data may end up making…

计算与语言 · 计算机科学 2023-08-30 Shaina Raza , Muskan Garg , Deepak John Reji , Syed Raza Bashir , Chen Ding

Specific lexical choices in narrative text reflect both the writer's attitudes towards people in the narrative and influence the audience's reactions. Prior work has examined descriptions of people in English using contextual affective…

计算与语言 · 计算机科学 2021-04-09 Chan Young Park , Xinru Yan , Anjalie Field , Yulia Tsvetkov

Hateful and Toxic content has become a significant concern in today's world due to an exponential rise in social media. The increase in hate speech and harmful content motivated researchers to dedicate substantial efforts to the challenging…

计算与语言 · 计算机科学 2021-01-25 Suman Dowlagar , Radhika Mamidi

With language models becoming increasingly ubiquitous, it has become essential to address their inequitable treatment of diverse demographic groups and factors. Most research on evaluating and mitigating fairness harms has been concentrated…

计算与语言 · 计算机科学 2023-03-01 Krithika Ramesh , Sunayana Sitaram , Monojit Choudhury

The detection of offensive, hateful and profane language has become a critical challenge since many users in social networks are exposed to cyberbullying activities on a daily basis. In this paper, we present an analysis of combining…

计算与语言 · 计算机科学 2021-12-10 Sherzod Hakimov , Ralph Ewerth

Content moderation on social media platforms shapes the dynamics of online discourse, influencing whose voices are amplified and whose are suppressed. Recent studies have raised concerns about the fairness of content moderation practices,…

计算与语言 · 计算机科学 2024-06-25 Rebecca Dorn , Lee Kezar , Fred Morstatter , Kristina Lerman

Automatic hate speech detection in online social networks is an important open problem in Natural Language Processing (NLP). Hate speech is a multidimensional issue, strongly dependant on language and cultural factors. Despite its…

计算与语言 · 计算机科学 2021-05-03 Aymé Arango , Jorge Pérez , Barbara Poblete

Large-scale web-scraped text corpora used to train general-purpose AI models often contain harmful demographic-targeted social biases, creating a regulatory need for data auditing and developing scalable bias-detection methods. Although…

计算与语言 · 计算机科学 2026-04-10 Ayan Majumdar , Feihao Chen , Jinghui Li , Xiaozhen Wang

Recent advances in deep learning techniques have enabled machines to generate cohesive open-ended text when prompted with a sequence of words as context. While these models now empower many downstream applications from conversation bots to…

计算与语言 · 计算机科学 2021-01-29 Jwala Dhamala , Tony Sun , Varun Kumar , Satyapriya Krishna , Yada Pruksachatkun , Kai-Wei Chang , Rahul Gupta

The widespread use of social media necessitates reliable and efficient detection of offensive content to mitigate harmful effects. Although sophisticated models perform well on individual datasets, they often fail to generalize due to…

计算与语言 · 计算机科学 2024-10-08 Huy Nghiem , Hal Daumé

In order to study online hate speech, the availability of datasets containing the linguistic phenomena of interest are of crucial importance. However, when it comes to specific target groups, for example teenagers, collecting such data may…

计算与语言 · 计算机科学 2020-05-06 Alessio Palmero Aprosio , Stefano Menini , Sara Tonelli

Despite the growing reliance on fairness benchmarks to evaluate language models, the datasets that underpin these benchmarks remain critically underexamined. This survey addresses that overlooked foundation by offering a comprehensive…

计算与语言 · 计算机科学 2025-09-23 Jiale Zhang , Zichong Wang , Avash Palikhe , Zhipeng Yin , Wenbin Zhang

Code-mixed discourse combines multiple languages in a single text. It is commonly used in informal discourse in countries with several official languages, but also in many other countries in combination with English or neighboring…

计算与语言 · 计算机科学 2025-04-16 Anjali Yadav , Tanya Garg , Matej Klemen , Matej Ulcar , Basant Agarwal , Marko Robnik Sikonja

Twitter has become a pivotal platform for conducting information operations (IOs), particularly during high-stakes political events. In this study, we analyze over a million tweets about the 2024 U.S. presidential election to explore an…

社会与信息网络 · 计算机科学 2025-01-17 Bowen Yi

Sentiment analysis (SA) systems are widely deployed in many of the world's languages, and there is well-documented evidence of demographic bias in these systems. In languages beyond English, scarcer training data is often supplemented with…

计算与语言 · 计算机科学 2023-05-23 Seraphina Goldfarb-Tarrant , Björn Ross , Adam Lopez