English

Token-Level Entropy Reveals Demographic Disparities in Language Models

Computation and Language 2026-05-08 v3 Computer Vision and Pattern Recognition

Abstract

We ask whether demographic identity, signaled by a name alone, systematically reshapes the generative distribution of a language model. Measuring full-vocabulary Shannon entropy at temperature zero across six open-weight base models and 5,760 implicit sentence-completion prompts (e.g., "Tanisha walked into the office on a Monday morning and"), we find that Black-associated names produce higher first-token entropy than White-associated names across all six architectures - opposite to the output-level homogeneity bias documented under explicit demographic prompting (Lee et al., 2024) - and Black-associated names always produce greater entropy above identity-neutral baselines than White-associated names (ΔΔ>0\Delta\Delta > 0 in all six models). Women-associated names co-occur with lower first-token entropy (DL-pooled β^=0.041,p=.019\hat\beta = -0.041, p = .019) and more homogeneous outputs (α^=+0.024,p<.001\hat\alpha = +0.024, p < .001) than men-associated names - a pattern convergent with homogeneity bias; race and gender effects are additive. Instruction tuning does not attenuate the race gap (matched-format DL-pooled β^=+0.153\hat{\beta}=+0.153). Running the same templates with explicit group labels instead of names yields null race effects in 10 of 12 models where implicit probing is significant - establishing that probing methodology is a primary determinant of which distributional structure is recovered.

Keywords

Cite

@article{arxiv.2501.19337,
  title  = {Token-Level Entropy Reveals Demographic Disparities in Language Models},
  author = {Messi H. J. Lee},
  journal= {arXiv preprint arXiv:2501.19337},
  year   = {2026}
}

Comments

9 pages

R2 v1 2026-06-28T21:28:06.485Z