English
Related papers

Related papers: EssayBench: Evaluating Large Language Models in Mu…

200 papers

Human feedback is crucial in the interactions between humans and Large Language Models (LLMs). However, existing research primarily focuses on benchmarking LLMs in single-turn dialogues. Even in benchmarks designed for multi-turn dialogues,…

Computation and Language · Computer Science 2025-02-18 Youquan Li , Miao Zheng , Fan Yang , Guosheng Dong , Bin Cui , Weipeng Chen , Zenan Zhou , Wentao Zhang

The emergence of Large Language Models (LLMs) presents transformative opportunities for education, generating numerous novel application scenarios. However, significant challenges remain: evaluation metrics vary substantially across…

Computers and Society · Computer Science 2025-08-01 Shou'ang Wei , Xinyun Wang , Shuzhen Bi , Jian Chen , Ruijia Li , Bo Jiang , Xin Lin , Min Zhang , Yu Song , BingDong Li , Aimin Zhou , Hao Hao

Software testing is a crucial phase in the software life cycle, helping identify potential risks and reduce maintenance costs. With the advancement of Large Language Models (LLMs), researchers have proposed an increasing number of LLM-based…

Software Engineering · Computer Science 2024-09-27 Quanjun Zhang , Ye Shang , Chunrong Fang , Siqi Gu , Jianyi Zhou , Zhenyu Chen

While Large Language Models (LLMs) are reshaping the paradigm of AI for Social Science (AI4SS), rigorously evaluating their capabilities in scholarly writing remains a major challenge. Existing benchmarks largely emphasize single-shot,…

Computation and Language · Computer Science 2026-02-18 Houping Yue , Zixiang Di , Mei Jiang , Bingdong Li , Hao Hao , Yu Song , Bo Jiang , Aimin Zhou

Most of the existing Large Language Model (LLM) benchmarks on scientific problem reasoning focus on problems grounded in high-school subjects and are confined to elementary algebraic operations. To systematically examine the reasoning…

Computation and Language · Computer Science 2024-07-01 Xiaoxuan Wang , Ziniu Hu , Pan Lu , Yanqiao Zhu , Jieyu Zhang , Satyen Subramaniam , Arjun R. Loomba , Shichang Zhang , Yizhou Sun , Wei Wang

Predictive analysis is a cornerstone of modern decision-making, with applications in various domains. Large Language Models (LLMs) have emerged as powerful tools in enabling nuanced, knowledge-intensive conversations, thus aiding in complex…

Computation and Language · Computer Science 2025-05-26 Qin Chen , Yuanyi Ren , Xiaojun Ma , Yuyang Shi

Recent advances in Large Language Models (LLMs) have highlighted the need for robust, comprehensive, and challenging benchmarks. Yet, research on evaluating their Emotional Intelligence (EI) is considerably limited. Existing benchmarks have…

Computation and Language · Computer Science 2024-07-18 Sahand Sabour , Siyang Liu , Zheyuan Zhang , June M. Liu , Jinfeng Zhou , Alvionna S. Sunaryo , Juanzi Li , Tatia M. C. Lee , Rada Mihalcea , Minlie Huang

Large language models (LLMs) show increasing potential in education, yet benchmarks for non-English languages in specialized domains remain scarce. We introduce MedBench-IT, the first comprehensive benchmark for evaluating LLMs on Italian…

Computation and Language · Computer Science 2025-09-10 Ruggero Marino Lazzaroni , Alessandro Angioi , Michelangelo Puliga , Davide Sanna , Roberto Marras

Recent advances in large language models (LLMs) and vision-language models (LVLMs) have shown promise across many tasks, yet their scientific reasoning capabilities remain untested, particularly in multimodal settings. We present…

Machine Learning · Computer Science 2025-06-03 Xinwu Ye , Chengfan Li , Siming Chen , Wei Wei , Xiangru Tang

Evaluating Large Language Models (LLMs) is crucial for understanding their capabilities and limitations across various applications, including natural language processing and code generation. Existing benchmarks like MMLU, C-Eval, and…

Cryptography and Security · Computer Science 2025-01-07 Pengfei Jing , Mengyun Tang , Xiaorong Shi , Xing Zheng , Sen Nie , Shi Wu , Yong Yang , Xiapu Luo

Large language models have recently made tremendous progress in a variety of aspects, e.g., cross-task generalization, instruction following. Comprehensively evaluating the capability of large language models in multiple tasks is of great…

Computation and Language · Computer Science 2023-05-23 Chuang Liu , Renren Jin , Yuqi Ren , Linhao Yu , Tianyu Dong , Xiaohan Peng , Shuting Zhang , Jianxiang Peng , Peiyi Zhang , Qingqing Lyu , Xiaowen Su , Qun Liu , Deyi Xiong

Large Language Models (LLMs) are transforming scholarly tasks like search and summarization, but their reliability remains uncertain. Current evaluation metrics for testing LLM reliability are primarily automated approaches that prioritize…

Human-Computer Interaction · Computer Science 2026-02-25 Anna Martin-Boyle , William Humphreys , Martha Brown , Cara Leckey , Harmanpreet Kaur

The ability of Large Language Models (LLMs) to critique and refine their reasoning is crucial for their application in evaluation, feedback provision, and self-improvement. This paper introduces CriticBench, a comprehensive benchmark…

Computation and Language · Computer Science 2024-06-04 Zicheng Lin , Zhibin Gou , Tian Liang , Ruilin Luo , Haowei Liu , Yujiu Yang

Large language models (LLMs) have emerged as pivotal contributors in contemporary natural language processing and are increasingly being applied across a diverse range of industries. However, these large-scale probabilistic statistical…

Computation and Language · Computer Science 2024-10-10 Xun Liang , Shichao Song , Simin Niu , Zhiyu Li , Feiyu Xiong , Bo Tang , Yezhaohui Wang , Dawei He , Peng Cheng , Zhonghao Wang , Haiying Deng

Given the importance of ancient Chinese in capturing the essence of rich historical and cultural heritage, the rapid advancements in Large Language Models (LLMs) necessitate benchmarks that can effectively evaluate their understanding of…

Computation and Language · Computer Science 2024-03-12 Yuting Wei , Yuanxing Xu , Xinru Wei , Simin Yang , Yangfu Zhu , Yuqing Li , Di Liu , Bin Wu

New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of…

Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field. However, existing benchmarks fail…

Computation and Language · Computer Science 2026-02-16 Ziqian Zhang , Xingjian Hu , Yue Huang , Kai Zhang , Ruoxi Chen , Yixin Liu , Qingsong Wen , Kaidi Xu , Xiangliang Zhang , Neil Zhenqiang Gong , Lichao Sun

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first…

Computation and Language · Computer Science 2025-09-23 Fei Zhao , Chengqiang Lu , Yufan Shen , Qimeng Wang , Yicheng Qian , Haoxin Zhang , Yan Gao , Yi Wu , Yao Hu , Zhen Wu , Shangyu Xing , Xinyu Dai

Large language models (LLMs) have demonstrated several emergent behaviors with scale, including reasoning and fluency in long-form text generation. However, they continue to struggle with tasks requiring precise spatial and positional…

Machine Learning · Computer Science 2025-12-05 Kerry Luo , Michael Fu , Joshua Peguero , Husnain Malik , Anvay Patil , Joyce Lin , Megan Van Overborg , Ryan Sarmiento , Kevin Zhu

Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation methods, as models exhibit varying capabilities across different…

Computation and Language · Computer Science 2025-06-04 Anna Sokol , Elizabeth Daly , Michael Hind , David Piorkowski , Xiangliang Zhang , Nuno Moniz , Nitesh Chawla