Multilingual Bias Detection and Mitigation for Indian Languages
Computation and Language
2023-12-27 v1
Abstract
Lack of diverse perspectives causes neutrality bias in Wikipedia content leading to millions of worldwide readers getting exposed by potentially inaccurate information. Hence, neutrality bias detection and mitigation is a critical problem. Although previous studies have proposed effective solutions for English, no work exists for Indian languages. First, we contribute two large datasets, mWikiBias and mWNC, covering 8 languages, for the bias detection and mitigation tasks respectively. Next, we investigate the effectiveness of popular multilingual Transformer-based models for the two tasks by modeling detection as a binary classification problem and mitigation as a style transfer problem. We make the code and data publicly available.
Cite
@article{arxiv.2312.15181,
title = {Multilingual Bias Detection and Mitigation for Indian Languages},
author = {Ankita Maity and Anubhav Sharma and Rudra Dhar and Tushar Abhishek and Manish Gupta and Vasudeva Varma},
journal= {arXiv preprint arXiv:2312.15181},
year = {2023}
}