English

Debiasing Multilingual LLMs in Cross-lingual Latent Space

Computation and Language 2025-08-26 v1 Artificial Intelligence Machine Learning

Abstract

Debiasing techniques such as SentDebias aim to reduce bias in large language models (LLMs). Previous studies have evaluated their cross-lingual transferability by directly applying these methods to LLM representations, revealing their limited effectiveness across languages. In this work, we therefore propose to perform debiasing in a joint latent space rather than directly on LLM representations. We construct a well-aligned cross-lingual latent space using an autoencoder trained on parallel TED talk scripts. Our experiments with Aya-expanse and two debiasing techniques across four languages (English, French, German, Dutch) demonstrate that a) autoencoders effectively construct a well-aligned cross-lingual latent space, and b) applying debiasing techniques in the learned cross-lingual latent space significantly improves both the overall debiasing performance and cross-lingual transferability.

Keywords

Cite

@article{arxiv.2508.17948,
  title  = {Debiasing Multilingual LLMs in Cross-lingual Latent Space},
  author = {Qiwei Peng and Guimin Hu and Yekun Chai and Anders Søgaard},
  journal= {arXiv preprint arXiv:2508.17948},
  year   = {2025}
}

Comments

EMNLP 2025 Main

R2 v1 2026-07-01T05:04:29.395Z