This study investigates how large language models, in particular LLaMA 3.2-3B, construct narratives about Black and white women in short stories generated in Portuguese. From 2100 texts, we applied computational methods to group semantically similar stories, allowing a selection for qualitative analysis. Three main discursive representations emerge: social overcoming, ancestral mythification and subjective self-realization. The analysis uncovers how grammatically coherent, seemingly neutral texts materialize a crystallized, colonially structured framing of the female body, reinforcing historical inequalities. The study proposes an integrated approach, that combines machine learning techniques with qualitative, manual discourse analysis.
@article{arxiv.2509.02834,
title = {Clustering Discourses: Racial Biases in Short Stories about Women Generated by Large Language Models},
author = {Gustavo Bonil and João Gondim and Marina dos Santos and Simone Hashiguti and Helena Maia and Nadia Silva and Helio Pedrini and Sandra Avila},
journal= {arXiv preprint arXiv:2509.02834},
year = {2026}
}
Comments
12 pages, 3 figures. Accepted at STIL @ BRACIS 2025