中文
相关论文

相关论文: When Does Demographic Information Help? Data and M…

200 篇论文

Although many fairness criteria have been proposed to ensure that machine learning algorithms do not exhibit or amplify our existing social biases, these algorithms are trained on datasets that can themselves be statistically biased. In…

机器学习 · 计算机科学 2023-05-04 Yiqiao Liao , Parinaz Naghizadeh

Individual and social biases undermine the effectiveness of human advisers by inducing judgment errors which can disadvantage protected groups. In this paper, we study the influence these biases can have in the pervasive problem of fake…

人机交互 · 计算机科学 2024-03-15 Axel Abels , Elias Fernandez Domingos , Ann Nowé , Tom Lenaerts

Large language models (LLMs) acquire general linguistic knowledge from massive-scale pretraining. However, pretraining data mainly comprised of web-crawled texts contain undesirable social biases which can be perpetuated or even amplified…

计算与语言 · 计算机科学 2025-09-04 Takuma Udagawa , Yang Zhao , Hiroshi Kanayama , Bishwaranjan Bhattacharjee

Language models (LMs) are pretrained on diverse data sources, including news, discussion forums, books, and online encyclopedias. A significant portion of this data includes opinions and perspectives which, on one hand, celebrate democracy…

计算与语言 · 计算机科学 2023-07-07 Shangbin Feng , Chan Young Park , Yuhan Liu , Yulia Tsvetkov

Training supervised machine learning systems with a fairness loss can improve prediction fairness across different demographic groups. However, doing so requires demographic annotations for training data, without which we cannot produce…

机器学习 · 计算机科学 2024-04-17 Carlos Aguirre , Mark Dredze

Computational social scientists often harness the Web as a "societal observatory" where data about human social behavior is collected. This data enables novel investigations of psychological, anthropological and sociological research…

计算机与社会 · 计算机科学 2016-03-15 Fariba Karimi , Claudia Wagner , Florian Lemmerich , Mohsen Jadidi , Markus Strohmaier

Affect preferences vary with user demographics, and tapping into demographic information provides important cues about the users' language preferences. In this paper, we utilize the user demographics, and propose EmpathBERT, a…

机器学习 · 计算机科学 2021-02-02 Bhanu Prakash Reddy Guda , Aparna Garimella , Niyati Chhaya

To tackle the rising phenomenon of hate speech, efforts have been made towards data curation and analysis. When it comes to analysis of bias, previous work has focused predominantly on race. In our work, we further investigate bias in hate…

计算与语言 · 计算机科学 2022-05-19 Antonis Maronikolakis , Philip Baader , Hinrich Schütze

Prediction models can improve efficiency by automating decisions such as the approval of loan applications. However, they may inherit bias against protected groups from the data they are trained on. This paper adds counterfactual…

机器学习 · 计算机科学 2024-05-03 Nicholas Tenev

Crowd counting has achieved significant progress by training regressors to predict instance positions. In heavily crowded scenarios, however, regressors are challenged by uncontrollable annotation variance, which causes density map bias and…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Mingyue Guo , Li Yuan , Zhaoyi Yan , Binghui Chen , Yaowei Wang , Qixiang Ye

Hate speech online targets individuals or groups based on identity attributes and spreads rapidly, posing serious social risks. Memes, which combine images and text, have emerged as a nuanced vehicle for disseminating hate speech, often…

多智能体系统 · 计算机科学 2026-03-26 Rui Xing , Qi Chai , Jie Ma , Jing Tao , Pinghui Wang , Shuming Zhang , Xinping Wang , Hao Wang

Crowdsourcing has emerged as a popular approach for collecting annotated data to train supervised machine learning models. However, annotator bias can lead to defective annotations. Though there are a few works investigating individual…

人机交互 · 计算机科学 2021-10-18 Haochen Liu , Joseph Thekinen , Sinem Mollaoglu , Da Tang , Ji Yang , Youlong Cheng , Hui Liu , Jiliang Tang

Algorithms are widely applied to detect hate speech and abusive language in social media. We investigated whether the human-annotated data used to train these algorithms are biased. We utilized a publicly available annotated Twitter dataset…

计算与语言 · 计算机科学 2020-05-29 Jae Yeon Kim , Carlos Ortiz , Sarah Nam , Sarah Santiago , Vivek Datta

In this short article, I leverage the National Crime Victimization Survey from 1992 to 2022 to examine how income, education, employment, and key demographic factors shape the type of crime victims experience (violent vs property). Using…

物理与社会 · 物理学 2025-06-06 Sydney Anuyah

Inferring information related to users enables to highly improve the quality of many mobile services. For example, knowing the demographic characteristics of a user allows a service to display more accurate information. According to the…

社会与信息网络 · 计算机科学 2018-03-13 Arielle Moro , Benoît Garbinato , Valérie Chavez-Demoulin

Social contexts -- such as families, schools, and neighborhoods -- shape life outcomes. The key question is not simply whether they matter, but rather for whom and under what conditions. Here, we argue that prediction gaps -- differences in…

社会与信息网络 · 计算机科学 2025-07-01 Javier Garcia-Bernardo , Eva Jaspers , Weverthon Machado , Samuel Plach , Erik Jan van Leeuwen

Approaches for mitigating bias in supervised models are designed to reduce models' dependence on specific sensitive features of the input data, e.g., mentioned social groups. However, in the case of hate speech detection, it is not always…

计算与语言 · 计算机科学 2020-10-27 Aida Mostafazadeh Davani , Ali Omrani , Brendan Kennedy , Mohammad Atari , Xiang Ren , Morteza Dehghani

It is tempting to think that machines are less prone to unfairness and prejudice. However, machine learning approaches compute their outputs based on data. While biases can enter at any stage of the development pipeline, models are…

计算机视觉与模式识别 · 计算机科学 2020-12-07 Patrick Esser , Robin Rombach , Björn Ommer

Generative classifiers offer potential advantages over their discriminative counterparts, namely in the areas of data efficiency, robustness to data shift and adversarial examples, and zero-shot learning (Ng and Jordan,2002; Yogatama et…

计算与语言 · 计算机科学 2019-10-02 Xiaoan Ding , Kevin Gimpel

NLP models often rely on human-labeled data for training and evaluation. Many approaches crowdsource this data from a large number of annotators with varying skills, backgrounds, and motivations, resulting in conflicting annotations. These…

计算与语言 · 计算机科学 2025-07-28 Jonathan Ivey , Susan Gauch , David Jurgens