中文
相关论文

相关论文: KAIO: A Collection of More Challenging Korean Ques…

200 篇论文

The Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models (LLMs), yet it has certain limitations. Notably, the disconnect between quantitative improvements on the overly academic leaderboard benchmarks…

计算与语言 · 计算机科学 2025-03-05 Hyeonwoo Kim , Dahyun Kim , Jihoo Kim , Sukyung Lee , Yungi Kim , Chanjun Park

We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. While prior Korean benchmarks are translated from existing English benchmarks, KMMLU is…

As large language models (LLMs) reach high scores on established mathematical benchmarks, such as GSM8K and MATH, the research community has turned to International Mathematical Olympiad (IMO) problems to push the evaluation frontier.…

人工智能 · 计算机科学 2025-09-10 Ziye Chen , Chengwei Qin , Yao Shu

A series of influential studies established that large language models cannot reliably solve even simple planning tasks. We show that the latest generation of frontier models overturns this conclusion. We evaluate three families of frontier…

人工智能 · 计算机科学 2026-05-18 Augusto B. Corrêa , André G. Pereira , Jendrik Seipp

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM reasoning. Whereas olympiad-style problems measure…

The instruction-following capabilities of large language models (LLMs) are pivotal for numerous applications, from conversational agents to complex reasoning systems. However, current evaluations predominantly focus on English models,…

计算与语言 · 计算机科学 2025-10-20 Dongjun Kim , Chanhee Park , Chanjun Park , Heuiseok Lim

This paper introduces the Open Ko-LLM Leaderboard and the Ko-H5 Benchmark as vital tools for evaluating Large Language Models (LLMs) in Korean. Incorporating private test sets while mirroring the English Open LLM Leaderboard, we establish a…

计算与语言 · 计算机科学 2024-08-20 Chanjun Park , Hyeonwoo Kim , Dahyun Kim , Seonghwan Cho , Sanghoon Kim , Sukyung Lee , Yungi Kim , Hwalsuk Lee

This study systematically evaluated the mathematical reasoning capabilities of Large Language Models (LLMs) using the 2026 Korean College Scholastic Ability Test (CSAT) Mathematics section, ensuring a completely contamination-free…

计算与语言 · 计算机科学 2025-12-02 Goun Pyeon , Inbum Heo , Jeesu Jung , Taewook Hwang , Hyuk Namgoong , Hyein Seo , Yerim Han , Eunbin Kim , Hyeonseok Kang , Sangkeun Jung

The evolving LLM landscape requires capabilities beyond simple text generation, prioritizing multi-step reasoning, long-context understanding, and agentic workflows. This shift challenges existing models in enterprise environments,…

计算与语言 · 计算机科学 2026-03-25 KT Tech innovation Group

This paper conducts a longitudinal study over eleven months to address the limitations of prior research on the Open Ko-LLM Leaderboard, which have relied on empirical studies with restricted observation periods of only five months. By…

计算与语言 · 计算机科学 2025-03-05 Chanjun Park , Hyeonwoo Kim

We introduce KFinEval-Pilot, a benchmark suite specifically designed to evaluate large language models (LLMs) in the Korean financial domain. Addressing the limitations of existing English-centric benchmarks, KFinEval-Pilot comprises over…

The performance of large language models (LLMs) on existing reasoning benchmarks has significantly improved over the past years. In response, we present JEEBench, a considerably more challenging benchmark dataset for evaluating the problem…

计算与语言 · 计算机科学 2023-10-24 Daman Arora , Himanshu Gaurav Singh , Mausam

Even though large language models are becoming increasingly capable, it is still unreasonable to expect them to excel at tasks that are under-represented on the Internet. Leveraging LLMs for specialized applications, particularly in niche…

机器学习 · 计算机科学 2025-08-18 Brendan R. Hogan , Will Brown , Adel Boyarsky , Anderson Schneider , Yuriy Nevmyvaka

Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant tasks. Existing suites have either saturated, heavily depend…

Large Language Models (LLMs) are commonly evaluated using human-crafted benchmarks, under the premise that higher scores implicitly reflect stronger human-like performance. However, there is growing concern that LLMs may ``game" these…

计算与语言 · 计算机科学 2024-12-16 Zhikai Lei , Tianyi Liang , Hanglei Hu , Jin Zhang , Yunhua Zhou , Yunfan Shao , Linyang Li , Chenchui Li , Changbo Wang , Hang Yan , Qipeng Guo

Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be addressed. This…

计算与语言 · 计算机科学 2024-02-27 Xiaotian Zhang , Chunyang Li , Yi Zong , Zhengyu Ying , Liang He , Xipeng Qiu

Large Language Models (LLMs) have shown impressive performance on a range of educational tasks, but are still understudied for their potential to solve mathematical problems. In this study, we compare three prominent LLMs, including GPT-4o,…

人工智能 · 计算机科学 2025-07-01 Ruonan Wang , Runxi Wang , Yunwen Shen , Chengfeng Wu , Qinglin Zhou , Rohitash Chandra

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given to their evaluation…

计算与语言 · 计算机科学 2024-03-21 Guijin Son , Hanwool Lee , Suwan Kim , Huiseo Kim , Jaecheol Lee , Je Won Yeom , Jihyu Jung , Jung Woo Kim , Songseong Kim

The rapid advancement of large language models (LLMs) has led to significant breakthroughs in automated mathematical reasoning and scientific discovery. Georgiev, G${\'o}$mez-Serrano, Tao, and Wagner [GGSTW+25] demonstrate that AI systems…

人工智能 · 计算机科学 2025-12-17 Yang Cao , Yubin Chen , Xuyang Guo , Zhao Song , Song Yue , Jiahao Zhang , Jiale Zhao

Recent frontier models employ long chain-of-thought reasoning to explore solution spaces in context and achieve stonger performance. While many works study distillation to build smaller yet capable models, most focus on English and little…

‹ 上一页 1 2 3 10 下一页 ›