中文
相关论文

相关论文: KMMLU: Measuring Massive Multitask Language Unders…

200 篇论文

We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams natively written in Korean, covering nine disciplines and nine…

计算与语言 · 计算机科学 2026-04-20 Nahyun Lee , Guijin Son , Hyunwoo Ko , Chanyoung Kim , JunYoung An , Kyubeen Han , Il-Youp Kwak

The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applicability in real-world scenarios. In this paper, we introduce…

计算与语言 · 计算机科学 2025-07-21 Seokhee Hong , Sunkyoung Kim , Guijin Son , Soyeon Kim , Yeonjung Hong , Jinsik Lee

Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguistic landscape, which combines Traditional Chinese script with…

计算与语言 · 计算机科学 2025-05-06 Chuxue Cao , Zhenghao Zhu , Junqi Zhu , Guoying Lu , Siyu Peng , Juntao Dai , Weijie Shi , Sirui Han , Yike Guo

Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and tasks in languages…

计算与语言 · 计算机科学 2024-10-14 Yeeun Kim , Young Rok Choi , Eunkyung Choi , Jinhwan Choi , Hai Jin Park , Wonseok Hwang

As the capabilities of large language models (LLMs) continue to advance, evaluating their performance becomes increasingly crucial and challenging. This paper aims to bridge this gap by introducing CMMLU, a comprehensive Chinese benchmark…

计算与语言 · 计算机科学 2024-01-19 Haonan Li , Yixuan Zhang , Fajri Koto , Yifei Yang , Hai Zhao , Yeyun Gong , Nan Duan , Timothy Baldwin

We present TMMLU+, a new benchmark designed for Traditional Chinese language understanding. TMMLU+ is a multi-choice question-answering dataset with 66 subjects from elementary to professional level. It is six times larger and boasts a more…

计算与语言 · 计算机科学 2024-07-12 Zhi-Rui Tam , Ya-Ting Pai , Yen-Wei Lee , Jun-Da Chen , Wei-Min Chu , Sega Cheng , Hong-Han Shuai

The Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models (LLMs), yet it has certain limitations. Notably, the disconnect between quantitative improvements on the overly academic leaderboard benchmarks…

计算与语言 · 计算机科学 2025-03-05 Hyeonwoo Kim , Dahyun Kim , Jihoo Kim , Sukyung Lee , Yungi Kim , Chanjun Park

We introduce Korean Language Understanding Evaluation (KLUE) benchmark. KLUE is a collection of 8 Korean natural language understanding (NLU) tasks, including Topic Classification, SemanticTextual Similarity, Natural Language Inference,…

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-linguistic reasoning abilities. This dual limitation makes it…

Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of Large Language Models (LLMs). While existing benchmarks employ automatic translation for multilingual evaluation, this approach is…

计算与语言 · 计算机科学 2024-10-04 Arda Yüksel , Abdullatif Köksal , Lütfi Kerem Şenel , Anna Korhonen , Hinrich Schütze

Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain scarce for non-English languages, with Korean being one such underexplored case. In this…

计算与语言 · 计算机科学 2026-04-23 Jinyoung Kim , Hyeongsoo Lim , Eunseo Seo , Minho Jang , Keunwoo Choi , Seungyoun Shin , Ji Won Yoon

We introduce the $\underline{Ko}rean \underline{G}rammar \underline{E}valuation Bench\underline{M}ark (KoGEM)$, designed to assess the linguistic competence of LLMs and humans in Korean. KoGEM consists of 1.5k multiple-choice QA pairs…

计算与语言 · 计算机科学 2025-06-03 SungHo Kim , Nayeon Kim , Taehee Jeon , SangKeun Lee

Large Language Models (LLMs) are commonly trained on multilingual corpora that include Greek, yet reliable evaluation benchmarks for Greek-particularly those based on authentic, native-sourced content-remain limited. Existing datasets are…

Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge. Because many existing Korean…

计算与语言 · 计算机科学 2024-07-08 Eunsu Kim , Juyoung Suk , Philhoon Oh , Haneul Yoo , James Thorne , Alice Oh

This paper introduces the Open Ko-LLM Leaderboard and the Ko-H5 Benchmark as vital tools for evaluating Large Language Models (LLMs) in Korean. Incorporating private test sets while mirroring the English Open LLM Leaderboard, we establish a…

计算与语言 · 计算机科学 2024-08-20 Chanjun Park , Hyeonwoo Kim , Dahyun Kim , Seonghwan Cho , Sanghoon Kim , Sukyung Lee , Yungi Kim , Hwalsuk Lee

To create culturally inclusive vision-language models (VLMs), developing a benchmark that tests their ability to address culturally relevant questions is essential. Existing approaches typically rely on human annotators, making the process…

计算与语言 · 计算机科学 2025-06-02 ChaeHun Park , Yujin Baek , Jaeseok Kim , Yu-Jung Heo , Du-Seong Chang , Jaegul Choo

The instruction-following capabilities of large language models (LLMs) are pivotal for numerous applications, from conversational agents to complex reasoning systems. However, current evaluations predominantly focus on English models,…

计算与语言 · 计算机科学 2025-10-20 Dongjun Kim , Chanhee Park , Chanjun Park , Heuiseok Lim

Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding, especially in Korean, are scarce. We conduct a corpus-based analysis of Korean negation and show that LLM performance…

计算与语言 · 计算机科学 2026-01-09 Sungmok Jung , Yeonkyoung So , Joonhak Lee , Sangho Kim , Yelim Ahn , Jaejin Lee

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given to their evaluation…

计算与语言 · 计算机科学 2024-03-21 Guijin Son , Hanwool Lee , Suwan Kim , Huiseo Kim , Jaecheol Lee , Je Won Yeom , Jihyu Jung , Jung Woo Kim , Songseong Kim

Despite having a population of twenty million, Kazakhstan's culture and language remain underrepresented in the field of natural language processing. Although large language models (LLMs) continue to advance worldwide, progress in Kazakh…

‹ 上一页 1 2 3 10 下一页 ›