English
Related papers

Related papers: Open Ko-LLM Leaderboard2: Bridging Foundational an…

200 papers

The proliferation of open-source Large Language Models (LLMs) from various institutions has highlighted the urgent need for comprehensive evaluation methods. However, current evaluation platforms, such as the widely recognized HuggingFace…

Computation and Language · Computer Science 2024-11-01 Fanghua Ye , Mingming Yang , Jianhui Pang , Longyue Wang , Derek F. Wong , Emine Yilmaz , Shuming Shi , Zhaopeng Tu

We introduce the Korean Canonical Legal Benchmark (KCL), a benchmark designed to assess language models' legal reasoning capabilities independently of domain-specific knowledge. KCL provides question-level supporting precedents, enabling a…

Computation and Language · Computer Science 2026-01-06 Hongseok Oh , Wonseok Hwang , Kyoung-Woon On

Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial before deploying them…

Developing a text readability assessment model specifically for texts in a foreign English Language Training (ELT) curriculum has never had much attention in the field of Natural Language Processing. Hence, most developed models show…

Computation and Language · Computer Science 2020-12-14 Bruce W. Lee , Jason Lee

As large language models (LLMs) become key advisors in various domains, their cultural sensitivity and reasoning skills are crucial in multicultural environments. We introduce Nunchi-Bench, a benchmark designed to evaluate LLMs' cultural…

Computation and Language · Computer Science 2025-07-08 Kyuhee Kim , Sangah Lee

Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation methods, as models exhibit varying capabilities across different…

Computation and Language · Computer Science 2025-06-04 Anna Sokol , Elizabeth Daly , Michael Hind , David Piorkowski , Xiangliang Zhang , Nuno Moniz , Nitesh Chawla

When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' growing capabilities, however there has been no similar focus…

Machine Learning · Computer Science 2025-02-06 Joshua Vendrow , Edward Vendrow , Sara Beery , Aleksander Madry

The rapid advancement of large language models (LLMs) has highlighted the need for robust evaluation frameworks that assess their core capabilities, such as reasoning, knowledge, and commonsense, leading to the inception of certain…

Computation and Language · Computer Science 2024-10-10 Dahyun Kim , Sukyung Lee , Yungi Kim , Attapol Rutherford , Chanjun Park

We introduce HyperCLOVA X, a family of large language models (LLMs) tailored to the Korean language and culture, along with competitive capabilities in English, math, and coding. HyperCLOVA X was trained on a balanced mix of Korean,…

Computation and Language · Computer Science 2024-04-16 Kang Min Yoo , Jaegeun Han , Sookyo In , Heewon Jeon , Jisu Jeong , Jaewook Kang , Hyunwook Kim , Kyung-Min Kim , Munhyong Kim , Sungju Kim , Donghyun Kwak , Hanock Kwak , Se Jung Kwon , Bado Lee , Dongsoo Lee , Gichang Lee , Jooho Lee , Baeseong Park , Seongjin Shin , Joonsang Yu , Seolki Baek , Sumin Byeon , Eungsup Cho , Dooseok Choe , Jeesung Han , Youngkyun Jin , Hyein Jun , Jaeseung Jung , Chanwoong Kim , Jinhong Kim , Jinuk Kim , Dokyeong Lee , Dongwook Park , Jeong Min Sohn , Sujung Han , Jiae Heo , Sungju Hong , Mina Jeon , Hyunhoon Jung , Jungeun Jung , Wangkyo Jung , Chungjoon Kim , Hyeri Kim , Jonghyun Kim , Min Young Kim , Soeun Lee , Joonhee Park , Jieun Shin , Sojin Yang , Jungsoon Yoon , Hwaran Lee , Sanghwan Bae , Jeehwan Cha , Karl Gylleus , Donghoon Ham , Mihak Hong , Youngki Hong , Yunki Hong , Dahyun Jang , Hyojun Jeon , Yujin Jeon , Yeji Jeong , Myunggeun Ji , Yeguk Jin , Chansong Jo , Shinyoung Joo , Seunghwan Jung , Adrian Jungmyung Kim , Byoung Hoon Kim , Hyomin Kim , Jungwhan Kim , Minkyoung Kim , Minseung Kim , Sungdong Kim , Yonghee Kim , Youngjun Kim , Youngkwan Kim , Donghyeon Ko , Dughyun Lee , Ha Young Lee , Jaehong Lee , Jieun Lee , Jonghyun Lee , Jongjin Lee , Min Young Lee , Yehbin Lee , Taehong Min , Yuri Min , Kiyoon Moon , Hyangnam Oh , Jaesun Park , Kyuyon Park , Younghun Park , Hanbae Seo , Seunghyun Seo , Mihyun Sim , Gyubin Son , Matt Yeo , Kyung Hoon Yeom , Wonjoon Yoo , Myungin You , Doheon Ahn , Homin Ahn , Joohee Ahn , Seongmin Ahn , Chanwoo An , Hyeryun An , Junho An , Sang-Min An , Boram Byun , Eunbin Byun , Jongho Cha , Minji Chang , Seunggyu Chang , Haesong Cho , Youngdo Cho , Dalnim Choi , Daseul Choi , Hyoseok Choi , Minseong Choi , Sangho Choi , Seongjae Choi , Wooyong Choi , Sewhan Chun , Dong Young Go , Chiheon Ham , Danbi Han , Jaemin Han , Moonyoung Hong , Sung Bum Hong , Dong-Hyun Hwang , Seongchan Hwang , Jinbae Im , Hyuk Jin Jang , Jaehyung Jang , Jaeni Jang , Sihyeon Jang , Sungwon Jang , Joonha Jeon , Daun Jeong , Joonhyun Jeong , Kyeongseok Jeong , Mini Jeong , Sol Jin , Hanbyeol Jo , Hanju Jo , Minjung Jo , Chaeyoon Jung , Hyungsik Jung , Jaeuk Jung , Ju Hwan Jung , Kwangsun Jung , Seungjae Jung , Soonwon Ka , Donghan Kang , Soyoung Kang , Taeho Kil , Areum Kim , Beomyoung Kim , Byeongwook Kim , Daehee Kim , Dong-Gyun Kim , Donggook Kim , Donghyun Kim , Euna Kim , Eunchul Kim , Geewook Kim , Gyu Ri Kim , Hanbyul Kim , Heesu Kim , Isaac Kim , Jeonghoon Kim , Jihye Kim , Joonghoon Kim , Minjae Kim , Minsub Kim , Pil Hwan Kim , Sammy Kim , Seokhun Kim , Seonghyeon Kim , Soojin Kim , Soong Kim , Soyoon Kim , Sunyoung Kim , Taeho Kim , Wonho Kim , Yoonsik Kim , You Jin Kim , Yuri Kim , Beomseok Kwon , Ohsung Kwon , Yoo-Hwan Kwon , Anna Lee , Byungwook Lee , Changho Lee , Daun Lee , Dongjae Lee , Ha-Ram Lee , Hodong Lee , Hwiyeong Lee , Hyunmi Lee , Injae Lee , Jaeung Lee , Jeongsang Lee , Jisoo Lee , Jongsoo Lee , Joongjae Lee , Juhan Lee , Jung Hyun Lee , Junghoon Lee , Junwoo Lee , Se Yun Lee , Sujin Lee , Sungjae Lee , Sungwoo Lee , Wonjae Lee , Zoo Hyun Lee , Jong Kun Lim , Kun Lim , Taemin Lim , Nuri Na , Jeongyeon Nam , Kyeong-Min Nam , Yeonseog Noh , Biro Oh , Jung-Sik Oh , Solgil Oh , Yeontaek Oh , Boyoun Park , Cheonbok Park , Dongju Park , Hyeonjin Park , Hyun Tae Park , Hyunjung Park , Jihye Park , Jooseok Park , Junghwan Park , Jungsoo Park , Miru Park , Sang Hee Park , Seunghyun Park , Soyoung Park , Taerim Park , Wonkyeong Park , Hyunjoon Ryu , Jeonghun Ryu , Nahyeon Ryu , Soonshin Seo , Suk Min Seo , Yoonjeong Shim , Kyuyong Shin , Wonkwang Shin , Hyun Sim , Woongseob Sim , Hyejin Soh , Bokyong Son , Hyunjun Son , Seulah Son , Chi-Yun Song , Chiyoung Song , Ka Yeon Song , Minchul Song , Seungmin Song , Jisung Wang , Yonggoo Yeo , Myeong Yeon Yi , Moon Bin Yim , Taehwan Yoo , Youngjoon Yoo , Sungmin Yoon , Young Jin Yoon , Hangyeol Yu , Ui Seon Yu , Xingdong Zuo , Jeongin Bae , Joungeun Bae , Hyunsoo Cho , Seonghyun Cho , Yongjin Cho , Taekyoon Choi , Yera Choi , Jiwan Chung , Zhenghui Han , Byeongho Heo , Euisuk Hong , Taebaek Hwang , Seonyeol Im , Sumin Jegal , Sumin Jeon , Yelim Jeong , Yonghyun Jeong , Can Jiang , Juyong Jiang , Jiho Jin , Ara Jo , Younghyun Jo , Hoyoun Jung , Juyoung Jung , Seunghyeong Kang , Dae Hee Kim , Ginam Kim , Hangyeol Kim , Heeseung Kim , Hyojin Kim , Hyojun Kim , Hyun-Ah Kim , Jeehye Kim , Jin-Hwa Kim , Jiseon Kim , Jonghak Kim , Jung Yoon Kim , Rak Yeong Kim , Seongjin Kim , Seoyoon Kim , Sewon Kim , Sooyoung Kim , Sukyoung Kim , Taeyong Kim , Naeun Ko , Bonseung Koo , Heeyoung Kwak , Haena Kwon , Youngjin Kwon , Boram Lee , Bruce W. Lee , Dagyeong Lee , Erin Lee , Euijin Lee , Ha Gyeong Lee , Hyojin Lee , Hyunjeong Lee , Jeeyoon Lee , Jeonghyun Lee , Jongheok Lee , Joonhyung Lee , Junhyuk Lee , Mingu Lee , Nayeon Lee , Sangkyu Lee , Se Young Lee , Seulgi Lee , Seung Jin Lee , Suhyeon Lee , Yeonjae Lee , Yesol Lee , Youngbeom Lee , Yujin Lee , Shaodong Li , Tianyu Liu , Seong-Eun Moon , Taehong Moon , Max-Lasse Nihlenramstroem , Wonseok Oh , Yuri Oh , Hongbeen Park , Hyekyung Park , Jaeho Park , Nohil Park , Sangjin Park , Jiwon Ryu , Miru Ryu , Simo Ryu , Ahreum Seo , Hee Seo , Kangdeok Seo , Jamin Shin , Seungyoun Shin , Heetae Sin , Jiangping Wang , Lei Wang , Ning Xiang , Longxiang Xiao , Jing Xu , Seonyeong Yi , Haanju Yoo , Haneul Yoo , Hwanhee Yoo , Liang Yu , Youngjae Yu , Weijie Yuan , Bo Zeng , Qian Zhou , Kyunghyun Cho , Jung-Woo Ha , Joonsuk Park , Jihyun Hwang , Hyoung Jo Kwon , Soonyong Kwon , Jungyeon Lee , Seungho Lee , Seonghyeon Lim , Hyunkyung Noh , Seungho Choi , Sang-Woo Lee , Jung Hwa Lim , Nako Sung

Large Language Models (LLMs) achieve remarkable performance across various tasks, but their tendency to produce hallucinations limits reliable adoption. Benchmarks such as TruthfulQA have been developed to measure truthfulness, yet they are…

Computation and Language · Computer Science 2025-09-09 Lorenzo Alfred Nery , Ronald Dawson Catignas , Thomas James Tiam-Lee

Large language models have exhibited significant enhancements in performance across various tasks. However, the complexity of their evaluation increases as these models generate more fluent and coherent content. Current multilingual…

Computation and Language · Computer Science 2024-12-11 Xiaonan Wang , Jinyoung Yeo , Joon-Ho Lim , Hansaem Kim

We present Ko-MuSR, the first benchmark to comprehensively evaluate multistep, soft reasoning in long Korean narratives while minimizing data contamination. Built following MuSR, Ko-MuSR features fully Korean narratives, reasoning chains,…

Computation and Language · Computer Science 2025-10-29 Chanwoo Park , Suyoung Park , JiA Kang , Jongyeon Park , Sangho Kim , Hyunji M. Park , Sumin Bae , Mingyu Kang , Jaejin Lee

As language models are often deployed as chatbot assistants, it becomes a virtue for models to engage in conversations in a user's first language. While these models are trained on a wide range of languages, a comprehensive evaluation of…

Computation and Language · Computer Science 2024-06-18 Seongbo Jang , Seonghyeon Lee , Hwanjo Yu

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-linguistic reasoning abilities. This dual limitation makes it…

We introduce KFinEval-Pilot, a benchmark suite specifically designed to evaluate large language models (LLMs) in the Korean financial domain. Addressing the limitations of existing English-centric benchmarks, KFinEval-Pilot comprises over…

We introduce GECKO, a bilingual large language model (LLM) optimized for Korean and English, along with programming languages. GECKO is pretrained on the balanced, high-quality corpus of Korean and English employing LLaMA architecture. In…

Computation and Language · Computer Science 2024-05-27 Sungwoo Oh , Donggyu Kim

This paper investigates the application of large language models (LLMs) to financial tasks. We fine-tuned foundation models using the Open FinLLM Leaderboard as a benchmark. Building on Qwen2.5 and Deepseek-R1, we employed techniques…

Computation and Language · Computer Science 2025-04-18 Varun Rao , Youran Sun , Mahendra Kumar , Tejas Mutneja , Agastya Mukherjee , Haizhao Yang

The critique capacity of Large Language Models (LLMs) is essential for reasoning abilities, which can provide necessary suggestions (e.g., detailed analysis and constructive feedback). Therefore, how to evaluate the critique capacity of…

In current benchmarks for evaluating large language models (LLMs), there are issues such as evaluation content restriction, untimely updates, and lack of optimization guidance. In this paper, we propose a new paradigm for the measurement of…

Computation and Language · Computer Science 2024-07-11 Jin Liu , Qingquan Li , Wenlong Du

We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams natively written in Korean, covering nine disciplines and nine…

Computation and Language · Computer Science 2026-04-20 Nahyun Lee , Guijin Son , Hyunwoo Ko , Chanyoung Kim , JunYoung An , Kyubeen Han , Il-Youp Kwak