中文
相关论文

相关论文: When Does Demographic Information Help? Data and M…

200 篇论文

Statistical language models conventionally implement representation learning based on the contextual distribution of words or other formal units, whereas any information related to the logographic features of written text are often ignored,…

计算与语言 · 计算机科学 2022-11-07 Zijian Jin , Duygu Ataman

Automatic hate speech detection is hampered by the scarcity of labeled datasetd, leading to poor generalization. We employ pretrained language models (LMs) to alleviate this data bottleneck. We utilize the GPT LM for generating large…

计算与语言 · 计算机科学 2021-09-03 Tomer Wullach , Amir Adler , Einat Minkov

Hate speech frequently appears on social media platforms and urgently needs to be effectively controlled. Alleviating the bias caused by hate speech can help resolve various ethical issues. Although existing research has constructed several…

计算与语言 · 计算机科学 2025-08-27 Hongyan Wu , Zhengming Chen , Zijian Li , Nankai Lin , Lianxi Wang , Shengyi Jiang , Aimin Yang

Public figures receive a disproportionate amount of abuse on social media, impacting their active participation in public life. Automated systems can identify abuse at scale but labelling training data is expensive, complex and potentially…

Hate speech detection is a common downstream application of natural language processing (NLP) in the real world. In spite of the increasing accuracy, current data-driven approaches could easily learn biases from the imbalanced data…

计算与语言 · 计算机科学 2022-09-22 Yi Cai , Arthur Zimek , Gerhard Wunder , Eirini Ntoutsi

Demographics, in particular, gender, age, and race, are a key predictor of human behavior. Despite the significant effect that demographics plays, most scientific studies using online social media do not consider this factor, mainly due to…

社会与信息网络 · 计算机科学 2016-03-08 Jisun An , Ingmar Weber

Human judgments are inherently subjective and are actively affected by personal traits such as gender and ethnicity. While Large Language Models (LLMs) are widely used to simulate human responses across diverse contexts, their ability to…

计算与语言 · 计算机科学 2025-02-18 Huaman Sun , Jiaxin Pei , Minje Choi , David Jurgens

Hate speech detection refers to the task of detecting hateful content that aims at denigrating an individual or a group based on their religion, gender, sexual orientation, or other characteristics. Due to the different policies of the…

计算与语言 · 计算机科学 2023-10-10 Paras Sheth , Tharindu Kumarage , Raha Moraffah , Aman Chadha , Huan Liu

Reliable automatic hate speech (HS) detection systems must adapt to the in-flow of diverse new data to curtail hate speech. However, hate speech detection systems commonly lack generalizability in identifying hate speech dissimilar to data…

计算与语言 · 计算机科学 2023-12-19 Shi Yin Hong , Susan Gauch

Label aggregation such as majority voting is commonly used to resolve annotator disagreement in dataset creation. However, this may disregard minority values and opinions. Recent studies indicate that learning from individual annotations…

计算与语言 · 计算机科学 2023-10-24 Xinpeng Wang , Barbara Plank

Critical decisions in hiring, college admissions, and credit lending are guided by predictions made in the presence of uncertainty. While uncertainty imparts errors across all demographic groups, this paper shows that the types of errors…

机器学习 · 统计学 2024-10-22 Claire Lazar Reich

The technical literature about data privacy largely consists of two complementary approaches: formal definitions of conditions sufficient for privacy preservation and attacks that demonstrate privacy breaches. Differential privacy is an…

密码学与安全 · 计算机科学 2025-02-06 Mark Bun , Marco Carmosino , Palak Jain , Gabriel Kaptchuk , Satchit Sivakumar

Large language models often achieve strong benchmark gains without corresponding improvements in broader capability. We hypothesize that this discrepancy arises from differences in training regimes induced by data distribution. To…

机器学习 · 计算机科学 2026-04-10 Hongjian Zou , Yidan Wang , Qi Ding , Yixuan Liao , Xiaoxin Chen

Online social media is rife with offensive and hateful comments, prompting the need for their automatic detection given the sheer amount of posts created every second. Creating high-quality human-labelled datasets for this task is difficult…

计算与语言 · 计算机科学 2023-08-01 João A. Leite , Carolina Scarton , Diego F. Silva

With the proliferation of social media, accurate detection of hate speech has become critical to ensure safety online. To combat nuanced forms of hate speech, it is important to identify and thoroughly explain hate speech to help users…

计算与语言 · 计算机科学 2023-11-23 Yongjin Yang , Joonkee Kim , Yujin Kim , Namgyu Ho , James Thorne , Se-young Yun

Discriminatory language and biases are often present in hate speech during conversations, which usually lead to negative impacts on targeted groups such as those based on race, gender, and religion. To tackle this issue, we propose an…

计算与语言 · 计算机科学 2023-07-21 Shaina Raza , Chen Ding , Deval Pandya

Face recognition algorithms, when used in the real world, can be very useful, but they can also be dangerous when biased toward certain demographics. So, it is essential to understand how these algorithms are trained and what factors affect…

计算机视觉与模式识别 · 计算机科学 2023-02-14 Manideep Kolla , Aravinth Savadamuthu

Crowdsourced annotations of data play a substantial role in the development of Artificial Intelligence (AI). It is broadly recognised that annotations of text data can contain annotator bias, where systematic disagreement in annotations can…

计算与语言 · 计算机科学 2024-10-22 Terne Sasha Thorn Jakobsen , Andreas Bjerre-Nielsen , Robert Böhm

Information on speaker characteristics can be useful as side information in improving speaker recognition accuracy. However, such information is often private. This paper investigates how privacy-preserving learning can improve a speaker…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Filip Granqvist , Matt Seigel , Rogier van Dalen , Áine Cahill , Stephen Shum , Matthias Paulik

Natural language processing research has begun to embrace the notion of annotator subjectivity, motivated by variations in labelling. This approach understands each annotator's view as valid, which can be highly suitable for tasks that…

计算与语言 · 计算机科学 2024-03-05 Amanda Cercas Curry , Gavin Abercrombie , Zeerak Talat