English

Curating Grounded Synthetic Data with Global Perspectives for Equitable AI

Computation and Language 2024-06-19 v2

Abstract

The development of robust AI models relies heavily on the quality and variety of training data available. In fields where data scarcity is prevalent, synthetic data generation offers a vital solution. In this paper, we introduce a novel approach to creating synthetic datasets, grounded in real-world diversity and enriched through strategic diversification. We synthesize data using a comprehensive collection of news articles spanning 12 languages and originating from 125 countries, to ensure a breadth of linguistic and cultural representations. Through enforced topic diversification, translation, and summarization, the resulting dataset accurately mirrors real-world complexities and addresses the issue of underrepresentation in traditional datasets. This methodology, applied initially to Named Entity Recognition (NER), serves as a model for numerous AI disciplines where data diversification is critical for generalizability. Preliminary results demonstrate substantial improvements in performance on traditional NER benchmarks, by up to 7.3%, highlighting the effectiveness of our synthetic data in mimicking the rich, varied nuances of global data sources. This paper outlines the strategies employed for synthesizing diverse datasets and provides such a curated dataset for NER.

Keywords

Cite

@article{arxiv.2406.10258,
  title  = {Curating Grounded Synthetic Data with Global Perspectives for Equitable AI},
  author = {Elin Törnquist and Robert Alexander Caulk},
  journal= {arXiv preprint arXiv:2406.10258},
  year   = {2024}
}
R2 v1 2026-06-28T17:06:34.404Z