English

AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI

Artificial Intelligence 2025-11-27 v1 Computers and Society Machine Learning

Abstract

The rapid evolution of generative AI necessitates robust safety evaluations. However, current safety datasets are predominantly English-centric, failing to capture specific risks in non-English, socio-cultural contexts such as Korean, and are often limited to the text modality. To address this gap, we introduce AssurAI, a new quality-controlled Korean multimodal dataset for evaluating the safety of generative AI. First, we define a taxonomy of 35 distinct AI risk factors, adapted from established frameworks by a multidisciplinary expert group to cover both universal harms and relevance to the Korean socio-cultural context. Second, leveraging this taxonomy, we construct and release AssurAI, a large-scale Korean multimodal dataset comprising 11,480 instances across text, image, video, and audio. Third, we apply the rigorous quality control process used to ensure data integrity, featuring a two-phase construction (i.e., expert-led seeding and crowdsourced scaling), triple independent annotation, and an iterative expert red-teaming loop. Our pilot study validates AssurAI's effectiveness in assessing the safety of recent LLMs. We release AssurAI to the public to facilitate the development of safer and more reliable generative AI systems for the Korean community.

Keywords

Cite

@article{arxiv.2511.20686,
  title  = {AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI},
  author = {Chae-Gyun Lim and Seung-Ho Han and EunYoung Byun and Jeongyun Han and Soohyun Cho and Eojin Joo and Heehyeon Kim and Sieun Kim and Juhoon Lee and Hyunsoo Lee and Dongkun Lee and Jonghwan Hyeon and Yechan Hwang and Young-Jun Lee and Kyeongryul Lee and Minhyeong An and Hyunjun Ahn and Jeongwoo Son and Junho Park and Donggyu Yoon and Taehyung Kim and Jeemin Kim and Dasom Choi and Kwangyoung Lee and Hyunseung Lim and Yeohyun Jung and Jongok Hong and Sooyohn Nam and Joonyoung Park and Sungmin Na and Yubin Choi and Jeanne Choi and Yoojin Hong and Sueun Jang and Youngseok Seo and Somin Park and Seoungung Jo and Wonhye Chae and Yeeun Jo and Eunyoung Kim and Joyce Jiyoung Whang and HwaJung Hong and Joseph Seering and Uichin Lee and Juho Kim and Sunna Choi and Seokyeon Ko and Taeho Kim and Kyunghoon Kim and Myungsik Ha and So Jung Lee and Jemin Hwang and JoonHo Kwak and Ho-Jin Choi},
  journal= {arXiv preprint arXiv:2511.20686},
  year   = {2025}
}

Comments

16 pages, HuggingFace: https://huggingface.co/datasets/TTA01/AssurAI