The collection of speech data carried out in Sociolinguistics has the potential to enhance large language models due to its quality and representativeness. In this paper, we examine the ethical considerations associated with the gathering and dissemination of such data. Additionally, we outline strategies for addressing the sensitivity of speech data, as it may facilitate the identification of informants who contributed with their speech.
@article{arxiv.2411.07512,
title = {\'Etica para LLMs: o compartilhamento de dados sociolingu\'isticos},
author = {Marta Deysiane Alves Faria Sousa and Raquel Meister Ko. Freitag and Túlio Sousa de Gois},
journal= {arXiv preprint arXiv:2411.07512},
year = {2025}
}
Comments
in Portuguese language. Paper accepted to LAAI-Ethics 2024