English
Related papers

Related papers: Thunder-KoNUBench: A Corpus-Aligned Benchmark for …

200 papers

Large Language Models (LLMs) trained via Reinforcement Learning (RL) have recently achieved impressive results on reasoning benchmarks. Yet, growing evidence shows that these models often generate longer but ineffective chains of thought…

Machine Learning · Computer Science 2025-07-02 Jhouben Cuesta-Ramirez , Samuel Beaussant , Mehdi Mounsif

The rise of large language models (LLMs) has led to more diverse and higher-quality machine-generated text. However, their high expressive power makes it difficult to control outputs based on specific business instructions. In response,…

Computation and Language · Computer Science 2025-01-28 Kentaro Kurihara , Masato Mita , Peinan Zhang , Shota Sasaki , Ryosuke Ishigami , Naoaki Okazaki

We introduce HyperCLOVA X, a family of large language models (LLMs) tailored to the Korean language and culture, along with competitive capabilities in English, math, and coding. HyperCLOVA X was trained on a balanced mix of Korean,…

Computation and Language · Computer Science 2024-04-16 Kang Min Yoo , Jaegeun Han , Sookyo In , Heewon Jeon , Jisu Jeong , Jaewook Kang , Hyunwook Kim , Kyung-Min Kim , Munhyong Kim , Sungju Kim , Donghyun Kwak , Hanock Kwak , Se Jung Kwon , Bado Lee , Dongsoo Lee , Gichang Lee , Jooho Lee , Baeseong Park , Seongjin Shin , Joonsang Yu , Seolki Baek , Sumin Byeon , Eungsup Cho , Dooseok Choe , Jeesung Han , Youngkyun Jin , Hyein Jun , Jaeseung Jung , Chanwoong Kim , Jinhong Kim , Jinuk Kim , Dokyeong Lee , Dongwook Park , Jeong Min Sohn , Sujung Han , Jiae Heo , Sungju Hong , Mina Jeon , Hyunhoon Jung , Jungeun Jung , Wangkyo Jung , Chungjoon Kim , Hyeri Kim , Jonghyun Kim , Min Young Kim , Soeun Lee , Joonhee Park , Jieun Shin , Sojin Yang , Jungsoon Yoon , Hwaran Lee , Sanghwan Bae , Jeehwan Cha , Karl Gylleus , Donghoon Ham , Mihak Hong , Youngki Hong , Yunki Hong , Dahyun Jang , Hyojun Jeon , Yujin Jeon , Yeji Jeong , Myunggeun Ji , Yeguk Jin , Chansong Jo , Shinyoung Joo , Seunghwan Jung , Adrian Jungmyung Kim , Byoung Hoon Kim , Hyomin Kim , Jungwhan Kim , Minkyoung Kim , Minseung Kim , Sungdong Kim , Yonghee Kim , Youngjun Kim , Youngkwan Kim , Donghyeon Ko , Dughyun Lee , Ha Young Lee , Jaehong Lee , Jieun Lee , Jonghyun Lee , Jongjin Lee , Min Young Lee , Yehbin Lee , Taehong Min , Yuri Min , Kiyoon Moon , Hyangnam Oh , Jaesun Park , Kyuyon Park , Younghun Park , Hanbae Seo , Seunghyun Seo , Mihyun Sim , Gyubin Son , Matt Yeo , Kyung Hoon Yeom , Wonjoon Yoo , Myungin You , Doheon Ahn , Homin Ahn , Joohee Ahn , Seongmin Ahn , Chanwoo An , Hyeryun An , Junho An , Sang-Min An , Boram Byun , Eunbin Byun , Jongho Cha , Minji Chang , Seunggyu Chang , Haesong Cho , Youngdo Cho , Dalnim Choi , Daseul Choi , Hyoseok Choi , Minseong Choi , Sangho Choi , Seongjae Choi , Wooyong Choi , Sewhan Chun , Dong Young Go , Chiheon Ham , Danbi Han , Jaemin Han , Moonyoung Hong , Sung Bum Hong , Dong-Hyun Hwang , Seongchan Hwang , Jinbae Im , Hyuk Jin Jang , Jaehyung Jang , Jaeni Jang , Sihyeon Jang , Sungwon Jang , Joonha Jeon , Daun Jeong , Joonhyun Jeong , Kyeongseok Jeong , Mini Jeong , Sol Jin , Hanbyeol Jo , Hanju Jo , Minjung Jo , Chaeyoon Jung , Hyungsik Jung , Jaeuk Jung , Ju Hwan Jung , Kwangsun Jung , Seungjae Jung , Soonwon Ka , Donghan Kang , Soyoung Kang , Taeho Kil , Areum Kim , Beomyoung Kim , Byeongwook Kim , Daehee Kim , Dong-Gyun Kim , Donggook Kim , Donghyun Kim , Euna Kim , Eunchul Kim , Geewook Kim , Gyu Ri Kim , Hanbyul Kim , Heesu Kim , Isaac Kim , Jeonghoon Kim , Jihye Kim , Joonghoon Kim , Minjae Kim , Minsub Kim , Pil Hwan Kim , Sammy Kim , Seokhun Kim , Seonghyeon Kim , Soojin Kim , Soong Kim , Soyoon Kim , Sunyoung Kim , Taeho Kim , Wonho Kim , Yoonsik Kim , You Jin Kim , Yuri Kim , Beomseok Kwon , Ohsung Kwon , Yoo-Hwan Kwon , Anna Lee , Byungwook Lee , Changho Lee , Daun Lee , Dongjae Lee , Ha-Ram Lee , Hodong Lee , Hwiyeong Lee , Hyunmi Lee , Injae Lee , Jaeung Lee , Jeongsang Lee , Jisoo Lee , Jongsoo Lee , Joongjae Lee , Juhan Lee , Jung Hyun Lee , Junghoon Lee , Junwoo Lee , Se Yun Lee , Sujin Lee , Sungjae Lee , Sungwoo Lee , Wonjae Lee , Zoo Hyun Lee , Jong Kun Lim , Kun Lim , Taemin Lim , Nuri Na , Jeongyeon Nam , Kyeong-Min Nam , Yeonseog Noh , Biro Oh , Jung-Sik Oh , Solgil Oh , Yeontaek Oh , Boyoun Park , Cheonbok Park , Dongju Park , Hyeonjin Park , Hyun Tae Park , Hyunjung Park , Jihye Park , Jooseok Park , Junghwan Park , Jungsoo Park , Miru Park , Sang Hee Park , Seunghyun Park , Soyoung Park , Taerim Park , Wonkyeong Park , Hyunjoon Ryu , Jeonghun Ryu , Nahyeon Ryu , Soonshin Seo , Suk Min Seo , Yoonjeong Shim , Kyuyong Shin , Wonkwang Shin , Hyun Sim , Woongseob Sim , Hyejin Soh , Bokyong Son , Hyunjun Son , Seulah Son , Chi-Yun Song , Chiyoung Song , Ka Yeon Song , Minchul Song , Seungmin Song , Jisung Wang , Yonggoo Yeo , Myeong Yeon Yi , Moon Bin Yim , Taehwan Yoo , Youngjoon Yoo , Sungmin Yoon , Young Jin Yoon , Hangyeol Yu , Ui Seon Yu , Xingdong Zuo , Jeongin Bae , Joungeun Bae , Hyunsoo Cho , Seonghyun Cho , Yongjin Cho , Taekyoon Choi , Yera Choi , Jiwan Chung , Zhenghui Han , Byeongho Heo , Euisuk Hong , Taebaek Hwang , Seonyeol Im , Sumin Jegal , Sumin Jeon , Yelim Jeong , Yonghyun Jeong , Can Jiang , Juyong Jiang , Jiho Jin , Ara Jo , Younghyun Jo , Hoyoun Jung , Juyoung Jung , Seunghyeong Kang , Dae Hee Kim , Ginam Kim , Hangyeol Kim , Heeseung Kim , Hyojin Kim , Hyojun Kim , Hyun-Ah Kim , Jeehye Kim , Jin-Hwa Kim , Jiseon Kim , Jonghak Kim , Jung Yoon Kim , Rak Yeong Kim , Seongjin Kim , Seoyoon Kim , Sewon Kim , Sooyoung Kim , Sukyoung Kim , Taeyong Kim , Naeun Ko , Bonseung Koo , Heeyoung Kwak , Haena Kwon , Youngjin Kwon , Boram Lee , Bruce W. Lee , Dagyeong Lee , Erin Lee , Euijin Lee , Ha Gyeong Lee , Hyojin Lee , Hyunjeong Lee , Jeeyoon Lee , Jeonghyun Lee , Jongheok Lee , Joonhyung Lee , Junhyuk Lee , Mingu Lee , Nayeon Lee , Sangkyu Lee , Se Young Lee , Seulgi Lee , Seung Jin Lee , Suhyeon Lee , Yeonjae Lee , Yesol Lee , Youngbeom Lee , Yujin Lee , Shaodong Li , Tianyu Liu , Seong-Eun Moon , Taehong Moon , Max-Lasse Nihlenramstroem , Wonseok Oh , Yuri Oh , Hongbeen Park , Hyekyung Park , Jaeho Park , Nohil Park , Sangjin Park , Jiwon Ryu , Miru Ryu , Simo Ryu , Ahreum Seo , Hee Seo , Kangdeok Seo , Jamin Shin , Seungyoun Shin , Heetae Sin , Jiangping Wang , Lei Wang , Ning Xiang , Longxiang Xiao , Jing Xu , Seonyeong Yi , Haanju Yoo , Haneul Yoo , Hwanhee Yoo , Liang Yu , Youngjae Yu , Weijie Yuan , Bo Zeng , Qian Zhou , Kyunghyun Cho , Jung-Woo Ha , Joonsuk Park , Jihyun Hwang , Hyoung Jo Kwon , Soonyong Kwon , Jungyeon Lee , Seungho Lee , Seonghyeon Lim , Hyunkyung Noh , Seungho Choi , Sang-Woo Lee , Jung Hwa Lim , Nako Sung

The application of large language models (LLMs) has achieved remarkable success in various fields, but their effectiveness in specialized domains like the Chinese insurance industry remains underexplored. The complexity of insurance…

Computation and Language · Computer Science 2025-01-22 Jing Ding , Kai Feng , Binbin Lin , Jiarui Cai , Qiushi Wang , Yu Xie , Xiaojin Zhang , Zhongyu Wei , Wei Chen

We propose MMLU-SR, a novel dataset designed to measure the true comprehension abilities of Large Language Models (LLMs) by challenging their performance in question-answering tasks with modified terms. We reasoned that an agent that…

Computation and Language · Computer Science 2024-10-07 Wentian Wang , Sarthak Jain , Paul Kantor , Jacob Feldman , Lazaros Gallos , Hao Wang

Test-time scaling has significantly improved large language model performance, enabling deeper reasoning to solve complex problems. However, this increased reasoning capability also leads to excessive token generation and unnecessary…

Large Language Models (LLMs) have demonstrated remarkable capabilities in code understanding and generation. However, their effectiveness on non-code Software Engineering (SE) tasks remains underexplored. We present 'Software Engineering…

Software Engineering · Computer Science 2026-02-12 Fabian C. Peña , Steffen Herbold

Traditional Korean medicine (TKM) emphasizes individualized diagnosis and treatment. This uniqueness makes AI modeling difficult due to limited data and implicit processes. Large language models (LLMs) have demonstrated impressive medical…

Computation and Language · Computer Science 2023-12-19 Dongyeop Jang , Tae-Rim Yun , Choong-Yeol Lee , Young-Kyu Kwon , Chang-Eop Kim

Subword tokenization has become the de-facto standard for tokenization, although comparative evaluations of subword vocabulary quality across languages are scarce. Existing evaluation studies focus on the effect of a tokenization algorithm…

Computation and Language · Computer Science 2023-10-23 Lisa Beinborn , Yuval Pinter

Large Language Models (LLMs) have advanced machine translation but remain vulnerable to hallucinations. Unfortunately, existing MT benchmarks are not capable of exposing failures in multilingual LLMs. To disclose hallucination in…

Computation and Language · Computer Science 2025-10-29 Xinwei Wu , Heng Liu , Jiang Zhou , Xiaohu Zhao , Linlong Xu , Longyue Wang , Weihua Luo , Kaifu Zhang

Understanding and reasoning over text within visual contexts poses a significant challenge for Vision-Language Models (VLMs), given the complexity and diversity of real-world scenarios. To address this challenge, text-rich Visual Question…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Taebaek Hwang , Minseo Kim , Gisang Lee , Seonuk Kim , Hyunjun Eun

This study systematically evaluated the mathematical reasoning capabilities of Large Language Models (LLMs) using the 2026 Korean College Scholastic Ability Test (CSAT) Mathematics section, ensuring a completely contamination-free…

Computation and Language · Computer Science 2025-12-02 Goun Pyeon , Inbum Heo , Jeesu Jung , Taewook Hwang , Hyuk Namgoong , Hyein Seo , Yerim Han , Eunbin Kim , Hyeonseok Kang , Sangkeun Jung

Steerability, or the ability of large language models (LLMs) to adapt outputs to align with diverse community-specific norms, perspectives, and communication styles, is critical for real-world applications but remains under-evaluated. We…

Computation and Language · Computer Science 2025-06-05 Kai Chen , Zihao He , Taiwei Shi , Kristina Lerman

Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications. However, the computational demands of these complex models pose significant challenges,…

Large Language Models (LLMs) hold significant potential for advancing fact-checking by leveraging their capabilities in reasoning, evidence retrieval, and explanation generation. However, existing benchmarks fail to comprehensively evaluate…

Computation and Language · Computer Science 2025-06-17 Shuo Yang , Yuqin Dai , Guoqing Wang , Xinran Zheng , Jinfeng Xu , Jinze Li , Zhenzhe Ying , Weiqiang Wang , Edith C. H. Ngai

Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills and the limitations of existing metrics. LLM's performance in thousand-words level and…

Computation and Language · Computer Science 2026-04-22 Andrew Zhuoer Feng , Cunxiang Wang , Yu Luo , Lin Fan , Yilin Zhou , Zikang Wang , Xiaotao Gu , Jie Tang , Hongning Wang , Minlie Huang

Knowledge Editing (KE) has emerged as a promising paradigm for updating facts in Large Language Models (LLMs) without retraining. However, progress in Multilingual Knowledge Editing (MKE) is currently hindered by biased evaluation…

Computation and Language · Computer Science 2026-01-27 Yucheng Hu , Wei Zhou , Juesi Xiao

The advent of NMT has expanded the scope of translation beyond isolated sentences, enabling context to be preserved across paragraphs and documents. However, current evaluation metrics largely remain restricted to the sentence level and…

Computational Engineering, Finance, and Science · Computer Science 2026-04-23 Hyeokmin Lee , Youngkyu Kim , Byounghyun Yoo

Almost all frameworks for the manual or automatic evaluation of machine translation characterize the quality of an MT output with a single number. An exception is the Multidimensional Quality Metrics (MQM) framework which offers a…

Computation and Language · Computer Science 2024-03-20 Dojun Park , Sebastian Padó

Large Language Models (LLMs) have demonstrated significant potential and effectiveness across multiple application domains. To assess the performance of mainstream LLMs in public security tasks, this study aims to construct a specialized…

Artificial Intelligence · Computer Science 2024-03-22 Xin Tong , Bo Jin , Zhi Lin , Binjun Wang , Ting Yu , Qiang Cheng