English
Related papers

Related papers: Token-Level Entropy Reveals Demographic Disparitie…

200 papers

Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based proxies to detect…

Computation and Language · Computer Science 2026-04-03 Bhaskara Hanuma Vedula , Darshan Anghan , Ishita Goyal , Ponnurangam Kumaraguru , Abhijnan Chakraborty

Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge. Existing evaluation methods, largely based on task performance or short-context behavior, provide…

Computation and Language · Computer Science 2026-05-26 Kumiko Tanaka-Ishii

We employ an audit design to investigate biases in state-of-the-art large language models, including GPT-4. In our study, we prompt the models for advice involving a named individual across a variety of scenarios, such as during car…

Computation and Language · Computer Science 2025-01-27 Alejandro Salinas , Amit Haim , Julian Nyarko

Personality have been found to predict many life outcomes, and there have been huge interests on automatic personality recognition from a speaker's utterance. Previously, we achieved accuracies between 37%-44% for three-way classification…

Sound · Computer Science 2018-02-06 Guozhen An , Rivka Levitan

Dialectal data are characterized by linguistic variation that appears small to humans but has a significant impact on the performance of models. This dialect gap has been related to various factors (e.g., data size, economic and social…

Computation and Language · Computer Science 2025-09-25 Vani Kanjirangat , Tanja Samardžić , Ljiljana Dolamic , Fabio Rinaldi

Bias has been a constant in face recognition models. Over the years, researchers have looked at it from both the model and the data point of view. However, their approach to mitigation of data bias was limited and lacked insight on the real…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Pedro C. Neto , Naser Damer , Jaime S. Cardoso , Ana F. Sequeira

Understanding commonsense knowledge is crucial in the field of Natural Language Processing (NLP). However, the presence of demographic terms in commonsense knowledge poses a potential risk of compromising the performance of NLP models. This…

Computation and Language · Computer Science 2024-06-12 JinKyu Lee , Jihie Kim

Enriching datasets with demographic information, such as gender, race, and age from names, is a critical task in fields like healthcare, public policy, and social sciences. Such demographic insights allow for more precise and effective…

Computation and Language · Computer Science 2024-09-19 Khaled AlNuaimi , Gautier Marti , Mathieu Ravaut , Abdulla AlKetbi , Andreas Henschel , Raed Jaradat

Large language model (LLM) tokenizers act as structured compressors: by mapping text to discrete token sequences, they determine token count (and thus compute and context usage) and the statistical structure seen by downstream models.…

Information Theory · Computer Science 2026-01-15 Mete Erdogan , Abhiram Gorle , Shubham Chandak , Mert Pilanci , Tsachy Weissman

This paper introduces an objective metric for evaluating a parsing scheme. It is based on Shannon's original work with letter sequences, which can be extended to part-of-speech tag sequences. It is shown that this regular language is an…

cmp-lg · Computer Science 2008-02-03 Caroline Lyon , Stephen Brown

The problem of Shannon entropy estimation in countable infinite alphabets is addressed from the study and use of convergence results of the entropy functional, which is known to be discontinuous with respect to the total variation distance…

Information Theory · Computer Science 2018-04-03 Jorge F. Silva

As access to high-quality, domain-specific data grows increasingly scarce, multi-epoch training has become a practical strategy for adapting large language models (LLMs). However, autoregressive models often suffer from performance…

Computation and Language · Computer Science 2025-12-30 Jiapeng Wang , Yiwen Hu , Yanzipeng Gao , Haoyu Wang , Shuo Wang , Hongyu Lu , Jiaxin Mao , Wayne Xin Zhao , Junyi Li , Xiao Zhang

The characteristics of social partners have long been hypothesized as influential in guiding group interactions. Understanding how demographic cues impact networks of creative collaborators is critical for elevating creative performances…

Social and Information Networks · Computer Science 2021-04-30 Raiyan Abdul Baten , Richard Aslin , Gourab Ghoshal , Mohammed Ehsan Hoque

Large Language Model (LLM) outputs often vary across user sociodemographic attributes, leading to disparities in factual accuracy, utility, and safety, even for objective questions where demographic information is irrelevant. Unlike prior…

Computation and Language · Computer Science 2026-01-15 Miao Zhang , Kelly Chen , Md Mehrab Tanjim , Rumi Chunara

State-of-the-art language generation models can degenerate when applied to open-ended generation problems such as text completion, story generation, or dialog modeling. This degeneration usually shows up in the form of incoherence, lack of…

Computation and Language · Computer Science 2023-02-15 Kushal Arora , Timothy J. O'Donnell , Doina Precup , Jason Weston , Jackie C. K. Cheung

Predicting upcoming words is a core mechanism of language comprehension and may be quantified using Shannon entropy. There is currently no empirical consensus on how many human responses are required to obtain stable and unbiased entropy…

Computation and Language · Computer Science 2026-02-05 Estrella Pivel-Villanueva , Elisabeth Frederike Sterner , Franziska Knolle

Tokenization inefficiency imposes structural disadvantages on morphologically complex, low-resource languages, inflating compute resources and depressing accuracy. We evaluate 10 large language models (LLMs) on AfriMMLU (9,000 MCQA items; 5…

Computation and Language · Computer Science 2026-03-25 Jessica M. Lundin , Ada Zhang , Nihal Karim , Hamza Louzan , Victor Wei , David Adelani , Cody Carroll

The metaphor of a potential epigenetic differentiation landscape broadly suggests that during differentiation a stem cell follows the steepest descending gradient toward a stable equilibrium state which represents the final cell type. It…

Cell Behavior · Quantitative Biology 2020-09-22 K. Wiesner , J. Teles , M. Hartnor , C. Peterson

Generative Large Language Models (LLMs) infer user's demographic information from subtle cues in the conversation -- a phenomenon called implicit personalization. Prior work has shown that such inferences can lead to lower quality responses…

Computation and Language · Computer Science 2025-09-17 Vera Neplenbroek , Arianna Bisazza , Raquel Fernández

We analyze the extent to which internal representations of language models (LMs) identify and distinguish mentions of named entities, focusing on the many-to-many correspondence between entities and their mentions. We first formulate two…

Computation and Language · Computer Science 2025-07-22 Masaki Sakata , Benjamin Heinzerling , Sho Yokoi , Takumi Ito , Kentaro Inui