English

Bias Association Discovery Framework for Open-Ended LLM Generations

Computation and Language 2026-01-21 v2

Abstract

Social biases embedded in Large Language Models (LLMs) raise critical concerns, resulting in representational harms -- unfair or distorted portrayals of demographic groups -- that may be expressed in subtle ways through generated language. Existing evaluation methods often depend on predefined identity-concept associations, limiting their ability to surface new or unexpected forms of bias. In this work, we present the Bias Association Discovery Framework (BADF), a systematic approach for extracting both known and previously unrecognized associations between demographic identities and descriptive concepts from open-ended LLM outputs. Through comprehensive experiments spanning multiple models and diverse real-world contexts, BADF enables robust mapping and analysis of the varied concepts that characterize demographic identities. Our findings advance the understanding of biases in open-ended generation and provide a scalable tool for identifying and analyzing bias associations in LLMs.

Keywords

Cite

@article{arxiv.2508.01412,
  title  = {Bias Association Discovery Framework for Open-Ended LLM Generations},
  author = {Jinhao Pan and Chahat Raj and Ziwei Zhu},
  journal= {arXiv preprint arXiv:2508.01412},
  year   = {2026}
}

Comments

AAAI 2026

R2 v1 2026-07-01T04:31:08.636Z