English
Related papers

Related papers: KITE: A Benchmark for Evaluating Korean Instructio…

200 papers

Instruction Tuning on Large Language Models is an essential process for model to function well and achieve high performance in specific tasks. Accordingly, in mainstream languages such as English, instruction-based datasets are being…

Computation and Language · Computer Science 2024-03-26 Dongjun Jang , Sungjoo Byun , Hyemi Jo , Hyopil Shin

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given to their evaluation…

Computation and Language · Computer Science 2024-03-21 Guijin Son , Hanwool Lee , Suwan Kim , Huiseo Kim , Jaecheol Lee , Je Won Yeom , Jihyu Jung , Jung Woo Kim , Songseong Kim

The Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models (LLMs), yet it has certain limitations. Notably, the disconnect between quantitative improvements on the overly academic leaderboard benchmarks…

Computation and Language · Computer Science 2025-03-05 Hyeonwoo Kim , Dahyun Kim , Jihoo Kim , Sukyung Lee , Yungi Kim , Chanjun Park

Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and tasks in languages…

Computation and Language · Computer Science 2024-10-14 Yeeun Kim , Young Rok Choi , Eunkyung Choi , Jinhwan Choi , Hai Jin Park , Wonseok Hwang

This paper introduces the Open Ko-LLM Leaderboard and the Ko-H5 Benchmark as vital tools for evaluating Large Language Models (LLMs) in Korean. Incorporating private test sets while mirroring the English Open LLM Leaderboard, we establish a…

Computation and Language · Computer Science 2024-08-20 Chanjun Park , Hyeonwoo Kim , Dahyun Kim , Seonghwan Cho , Sanghoon Kim , Sukyung Lee , Yungi Kim , Hwalsuk Lee

Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge. Because many existing Korean…

Computation and Language · Computer Science 2024-07-08 Eunsu Kim , Juyoung Suk , Philhoon Oh , Haneul Yoo , James Thorne , Alice Oh

Large language models have exhibited significant enhancements in performance across various tasks. However, the complexity of their evaluation increases as these models generate more fluent and coherent content. Current multilingual…

Computation and Language · Computer Science 2024-12-11 Xiaonan Wang , Jinyoung Yeo , Joon-Ho Lim , Hansaem Kim

Recent advancements in Korean large language models (LLMs) have driven numerous benchmarks and evaluation methods, yet inconsistent protocols cause up to 10 p.p performance gaps across institutions. Overcoming these reproducibility gaps…

Computational Engineering, Finance, and Science · Computer Science 2026-02-16 Hanwool Lee , Dasol Choi , Sooyong Kim , Ilgyun Jeong , Sangwon Baek , Guijin Son , Inseon Hwang , Naeun Lee , Seunghyeok Hong

Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain scarce for non-English languages, with Korean being one such underexplored case. In this…

Computation and Language · Computer Science 2026-04-23 Jinyoung Kim , Hyeongsoo Lim , Eunseo Seo , Minho Jang , Keunwoo Choi , Seungyoun Shin , Ji Won Yoon

Knowledge Tracing (KT) is a research field that aims to estimate a student's knowledge state through learning interactions-a crucial component of Intelligent Tutoring Systems (ITSs). Despite significant advancements, no current KT models…

Computers and Society · Computer Science 2024-12-13 Yongwan Cho , Rabia Emhamed AlMamlook , Tasnim Gharaibeh

Benchmarks play a significant role in the current evaluation of Large Language Models (LLMs), yet they often overlook the models' abilities to capture the nuances of human language, primarily focusing on evaluating embedded knowledge and…

Computation and Language · Computer Science 2024-10-18 Dojun Park , Jiwoo Lee , Hyeyun Jeong , Seohyun Park , Sungeun Lee

Large language models (LLMs) demonstrate exceptional performance on complex reasoning tasks. However, despite their strong reasoning capabilities in high-resource languages (e.g., English and Chinese), a significant performance gap persists…

Computation and Language · Computer Science 2025-02-03 Hyunwoo Ko , Guijin Son , Dasol Choi

We introduce Korean Language Understanding Evaluation (KLUE) benchmark. KLUE is a collection of 8 Korean natural language understanding (NLU) tasks, including Topic Classification, SemanticTextual Similarity, Natural Language Inference,…

The effective assessment of the instruction-following ability of large language models (LLMs) is of paramount importance. A model that cannot adhere to human instructions might be not able to provide reliable and helpful responses. In…

Computation and Language · Computer Science 2023-11-17 Yimin Jing , Renren Jin , Jiahao Hu , Huishi Qiu , Xiaohua Wang , Peng Wang , Deyi Xiong

This paper conducts a longitudinal study over eleven months to address the limitations of prior research on the Open Ko-LLM Leaderboard, which have relied on empirical studies with restricted observation periods of only five months. By…

Computation and Language · Computer Science 2025-03-05 Chanjun Park , Hyeonwoo Kim

We introduce KoBALT (Korean Benchmark for Advanced Linguistic Tasks), a comprehensive linguistically-motivated benchmark comprising 700 multiple-choice questions spanning 24 phenomena across five linguistic domains: syntax, semantics,…

Computation and Language · Computer Science 2025-05-23 Hyopil Shin , Sangah Lee , Dongjun Jang , Wooseok Song , Jaeyoon Kim , Chaeyoung Oh , Hyemi Jo , Youngchae Ahn , Sihyun Oh , Hyohyeong Chang , Sunkyoung Kim , Jinsik Lee

We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. While prior Korean benchmarks are translated from existing English benchmarks, KMMLU is…

Computation and Language · Computer Science 2024-06-07 Guijin Son , Hanwool Lee , Sungdong Kim , Seungone Kim , Niklas Muennighoff , Taekyoon Choi , Cheonbok Park , Kang Min Yoo , Stella Biderman

This work presents the first large-scale investigation into constructing a fully open bilingual large language model (LLM) for a non-English language, specifically Korean, trained predominantly on synthetic data. We introduce KORMo-10B, a…

Computation and Language · Computer Science 2025-10-13 Minjun Kim , Hyeonseok Lim , Hangyeol Yoo , Inho Won , Seungwoo Song , Minkyung Cho , Junhun Yuk , Changsu Choi , Dongjae Shin , Huige Lee , Hoyun Song , Alice Oh , Kyungtae Lim

A well-formulated benchmark plays a critical role in spurring advancements in the natural language processing (NLP) field, as it allows objective and precise evaluation of diverse models. As modern language models (LMs) have become more…

Computation and Language · Computer Science 2022-04-12 Dohyeong Kim , Myeongjun Jang , Deuk Sin Kwon , Eric Davis
‹ Prev 1 2 3 10 Next ›