English

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

Machine Learning 2026-05-19 v2 Computation and Language

Abstract

Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, these datasets often lack the sociotechnical nuance required for sensitive domains. We introduce NodeSynth, an evidence-grounded methodology that generates socially relevant synthetic queries by leveraging a fine-tuned taxonomy generator (TaG) anchored in real-world evidence. Evaluated against four mainstream LLMs (e.g., Claude 4.5 Haiku), NodeSynth elicited failure rates up to five times higher than human-authored benchmarks. Ablation studies confirm that our granular taxonomic expansion significantly drives these failure rates, while independent validation reveals critical deficiencies in prominent guard models (e.g., Llama-Guard-3). We open-source our end-to-end research prototype and datasets to enable scalable, high-stakes model evaluation and targeted safety interventions (https://github.com/google-research/nodesynth).

Keywords

Cite

@article{arxiv.2605.14381,
  title  = {NodeSynth: Socially Aligned Synthetic Data for AI Evaluation},
  author = {Qazi Mamunur Rashid and Xuan Yang and Zhengzhe Yang and Yanzhou Pan and Erin van Liemt and Darlene Neal and Kshitij Pancholi and Jamila Smith-Loud},
  journal= {arXiv preprint arXiv:2605.14381},
  year   = {2026}
}