English
Related papers

Related papers: Token-Level Entropy Reveals Demographic Disparitie…

200 papers

To answer questions about racial inequality and fairness, we often need a way to infer race and ethnicity from names. One way to infer race and ethnicity from names is by relying on the Census Bureau's list of popular last names. The list,…

Applications · Statistics 2023-07-13 Rajashekar Chintalapati , Suriyan Laohaprapanon , Gaurav Sood

Document classification is the detection specific content of interest in text documents. In contrast to the data-driven machine learning classifiers, knowledge-based classifiers can be constructed based on domain specific knowledge, which…

Computation and Language · Computer Science 2022-06-07 AtMa P. O. Chan

Annotators' sociodemographic backgrounds (i.e., the individual compositions of their gender, age, educational background, etc.) have a strong impact on their decisions when working on subjective NLP tasks, such as toxic language detection.…

Computation and Language · Computer Science 2024-02-09 Tilman Beck , Hendrik Schuff , Anne Lauscher , Iryna Gurevych

Self-supervised speaker embeddings are widely used in speaker verification systems, but prior work has shown that they often encode sensitive demographic attributes, raising fairness and privacy concerns. This paper investigates the extent…

Tokenization is associated with many poorly understood shortcomings in language models (LMs), yet remains an important component for long sequence scaling purposes. This work studies how tokenization impacts model performance by analyzing…

Computation and Language · Computer Science 2025-04-15 Buu Phan , Brandon Amos , Itai Gat , Marton Havasi , Matthew Muckley , Karen Ullrich

This paper investigates the impact of using first names in Large Language Models (LLMs) and Vision Language Models (VLMs), particularly when prompted with ethical decision-making tasks. We propose an approach that appends first names to…

Computation and Language · Computer Science 2024-08-12 Lorenzo Berlincioni , Luca Cultrera , Federico Becattini , Marco Bertini , Alberto Del Bimbo

We propose a compression-based version of the empirical entropy of a finite string over a finite alphabet. Whereas previously one considers the naked entropy of (possibly higher order) Markov processes, we consider the sum of the…

Information Theory · Computer Science 2011-04-05 Paul M. B. Vitányi

Large Language Models (LLMs) exhibit socio-economic biases that can propagate into downstream tasks. While prior studies have questioned whether intrinsic bias in LLMs affects fairness at the downstream task level, this work empirically…

Computation and Language · Computer Science 2025-09-23 'Mina Arzaghi' , 'Alireza Dehghanpour Farashah' , 'Florian Carichon' , ' Golnoosh Farnadi'

Semisupervised text classification has become a major focus of research over the past few years. Hitherto, most of the research has been based on supervised learning, but its main drawback is the unavailability of labeled data samples in…

Machine Learning · Computer Science 2021-11-17 Shivani Malhotra , Vinay Kumar , Alpana Agarwal

Transgender and non-binary (TGNB) individuals disproportionately experience discrimination and exclusion from daily life. Given the recent popularity and adoption of language generation technologies, the potential to further marginalize…

Computation and Language · Computer Science 2023-06-05 Anaelia Ovalle , Palash Goyal , Jwala Dhamala , Zachary Jaggers , Kai-Wei Chang , Aram Galstyan , Richard Zemel , Rahul Gupta

Text-to-image models are now easy to use and ubiquitous. However, prior work has found that they are prone to recapitulating harmful Western stereotypes. For example, requesting that a model generate an "African person and their house," may…

Computers and Society · Computer Science 2024-05-10 Joshua N. Williams , Molly FitzMorris , Osman Aka , Sarah Laszlo

Disparities in authorship and citations across gender can have substantial adverse consequences not just on the disadvantaged genders, but also on the field of study as a whole. Measuring gender gaps is a crucial step towards addressing…

Digital Libraries · Computer Science 2020-09-07 Saif M. Mohammad

The learning of the deep networks largely relies on the data with human-annotated labels. In some label insufficient situations, the performance degrades on the decision boundary with high data density. A common solution is to directly…

Computer Vision and Pattern Recognition · Computer Science 2020-03-30 Shuhao Cui , Shuhui Wang , Junbao Zhuo , Liang Li , Qingming Huang , Qi Tian

We find that the way we choose to represent data labels can have a profound effect on the quality of trained models. For example, training an image classifier to regress audio labels rather than traditional categorical probabilities…

Machine Learning · Computer Science 2021-04-07 Boyuan Chen , Yu Li , Sunand Raghupathi , Hod Lipson

Voice-based interfaces are widely used; however, achieving fair Wake-up Word detection across diverse speaker populations remains a critical challenge due to persistent demographic biases. This study evaluates the effectiveness of…

Computation and Language · Computer Science 2026-04-08 Fernando López , Paula Delgado-Santos , Pablo Gómez , David Solans , Jordi Luque

General-purpose language models are trained to produce varied natural language outputs, but for some tasks, like annotation or classification, we need more specific output formats. LLM systems increasingly support structured output, which…

Computation and Language · Computer Science 2025-08-04 Sil Hamilton , David Mimno

We show how generalized Gibbs-Shannon entropies can provide new insights on the statistical properties of texts. The universal distribution of word frequencies (Zipf's law) implies that the generalized entropies, computed at the word level,…

Physics and Society · Physics 2017-02-15 Eduardo G. Altmann , Laercio Dias , Martin Gerlach

We study the effect of tokenization on gender bias in machine translation, an aspect that has been largely overlooked in previous works. Specifically, we focus on the interactions between the frequency of gendered profession names in…

Computation and Language · Computer Science 2023-10-03 Bar Iluz , Tomasz Limisiewicz , Gabriel Stanovsky , David Mareček

Text-to-image generative models have achieved unprecedented success in generating high-quality images based on natural language descriptions. However, it is shown that these models tend to favor specific social groups when prompted with…

Computation and Language · Computer Science 2022-10-28 Hritik Bansal , Da Yin , Masoud Monajatipoor , Kai-Wei Chang

Automated depression detection often relies on static aggregation of conversational signals, potentially obscuring clinically meaningful behavioral dynamics. We investigated whether entropy-driven temporal biomarkers improve depression…

Other Quantitative Biology · Quantitative Biology 2026-05-01 Himadri S Samanta