中文
相关论文

相关论文: Token-Level Entropy Reveals Demographic Disparitie…

200 篇论文

During training, Large Language Models (LLMs) learn social regularities that can lead to gender bias in downstream applications. Most mitigation efforts focus on reducing bias in generated outputs, typically evaluated on structured…

Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-specified stereotypes. We introduce StereoTales, a multilingual…

Nationality identification unlocks important demographic information, with many applications in biomedical and sociological research. Existing name-based nationality classifiers use name substrings as features and are trained on small,…

社会与信息网络 · 计算机科学 2017-08-29 Junting Ye , Shuchu Han , Yifan Hu , Baris Coskun , Meizhu Liu , Hong Qin , Steven Skiena

The increasing use of Large Language Models (LLMs) in a large variety of domains has sparked worries about how easily they can perpetuate stereotypes and contribute to the generation of biased content. With a focus on gender and…

计算与语言 · 计算机科学 2025-07-28 Gioele Giachino , Marco Rondina , Antonio Vetrò , Riccardo Coppola , Juan Carlos De Martin

Diffusion models do not recover semantic structure uniformly over time. Instead, samples transition from semantic ambiguity to class commitment within a narrow regime. Recent theoretical work attributes this transition to dynamical…

Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model…

计算与语言 · 计算机科学 2025-03-03 Matthias Orlikowski , Paul Röttger , Philipp Cimiano , Dirk Hovy

Studying the ways in which language is gendered has long been an area of interest in sociolinguistics. Studies have explored, for example, the speech of male and female characters in film and the language used to describe male and female…

计算与语言 · 计算机科学 2019-06-13 Alexander Hoyle , Wolf-Sonkin , Hanna Wallach , Isabelle Augenstein , Ryan Cotterell

Word vector representations are well developed tools for various NLP and Machine Learning tasks and are known to retain significant semantic and syntactic structure of languages. But they are prone to carrying and amplifying bias which can…

计算与语言 · 计算机科学 2019-01-24 Sunipa Dev , Jeff Phillips

Gender-inclusive NLP research has documented the harmful limitations of gender binary-centric large language models (LLM), such as the inability to correctly use gender-diverse English neopronouns (e.g., xe, zir, fae). While data scarcity…

As pretrained large language models replace task-specific decoders in speech recognition, a critical question arises: do their text-derived priors make recognition fairer or more biased across demographic groups? We evaluate nine models…

计算与语言 · 计算机科学 2026-04-24 Srishti Ginjala , Eric Fosler-Lussier , Christopher W. Myers , Srinivasan Parthasarathy

Computational social scientists often harness the Web as a "societal observatory" where data about human social behavior is collected. This data enables novel investigations of psychological, anthropological and sociological research…

计算机与社会 · 计算机科学 2016-03-15 Fariba Karimi , Claudia Wagner , Florian Lemmerich , Mohsen Jadidi , Markus Strohmaier

We introduce a data-centric hypothesis-testing framework to quantify the influence of sequentially correlated literary properties--such as thematic continuity--on textual classification tasks. Our method models label sequences as stochastic…

计算与语言 · 计算机科学 2025-04-25 Gideon Yoffe , Nachum Dershowitz , Ariel Vishne , Barak Sober

Entropies must correspond to mean values for them to be measurable. The Shannon entropy corresponds to the weighted arithmetic mean, whereas the Renyi entropy corresponds to the exponential mean. These means refer to code lengths, which are…

统计力学 · 物理学 2011-10-25 B. H. Lavenda

Gender bias in artificial intelligence has become an important issue, particularly in the context of language models used in communication-oriented applications. This study examines the extent to which Large Language Models (LLMs) exhibit…

计算与语言 · 计算机科学 2024-11-18 Michael Döll , Markus Döhring , Andreas Müller

We consider the entropy of sums of independent discrete random variables, in analogy with Shannon's Entropy Power Inequality, where equality holds for normals. In our case, infinite divisibility suggests that equality should hold for…

信息论 · 计算机科学 2010-10-21 Oliver Johnson , Yaming Yu

We present our 7th place solution to the Gendered Pronoun Resolution challenge, which uses BERT without fine-tuning and a novel augmentation strategy designed for contextual embedding token-level tasks. Our method anonymizes the referent by…

计算与语言 · 计算机科学 2019-06-12 Bo Liu

We examine whether large language models (LLMs) exhibit race- and gender-based name discrimination in hiring decisions, similar to classic findings in the social sciences (Bertrand and Mullainathan, 2004). We design a series of templatic…

计算与语言 · 计算机科学 2024-06-18 Haozhe An , Christabel Acquaye , Colin Wang , Zongxia Li , Rachel Rudinger

In this paper, we pose the question: do people talk about women and men in different ways? We introduce two datasets and a novel integration of approaches for automatically inferring gender associations from language, discovering coherent…

计算与语言 · 计算机科学 2019-09-04 Serina Chang , Kathleen McKeown

The translation of written language has been known since the 3rd century BC; however, its necessity has become increasingly common in the information age. Today, many translators exist, based on encoder-decoder deep architectures,…

计算与语言 · 计算机科学 2025-11-18 Ronit D. Gross , Yanir Harel , Ido Kanter

Large language models (LLMs) increasingly influence global digital ecosystems, yet their potential to perpetuate social and cultural biases remains poorly understood in underrepresented contexts. This study presents a systematic analysis of…

计算与语言 · 计算机科学 2026-03-10 Ashish Pandey , Tek Raj Chhetri