English

The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs

Audio and Speech Processing 2026-03-19 v1 Computation and Language Sound

Abstract

Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This introduces speaker identity dependent variation in responses. We present a large-scale intersectional evaluation of accent and gender bias in three SpeechLLMs using 2,880 controlled interactions across six English accents and two gender presentations, keeping linguistic content constant through voice cloning. Using pointwise LLM-judge ratings, pairwise comparisons, and Best-Worst Scaling with human validation, we detect consistent disparities. Eastern European-accented speech receives lower helpfulness scores, particularly for female-presenting voices. The bias is implicit: responses remain polite but differ in helpfulness. While LLM judges capture the directional trend of these biases, human evaluators exhibit significantly higher sensitivity, uncovering sharper intersectional disparities.

Keywords

Cite

@article{arxiv.2603.16941,
  title  = {The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs},
  author = {Shree Harsha Bokkahalli Satish and Christoph Minixhofer and Maria Teleki and James Caverlee and Ondřej Klejch and Peter Bell and Gustav Eje Henter and Éva Székely},
  journal= {arXiv preprint arXiv:2603.16941},
  year   = {2026}
}

Comments

5 pages, 3 figures, 1 table, Submitted to Interspeech 2026

R2 v1 2026-07-01T11:24:50.256Z