English
Related papers

Related papers: Benchmarking Chinese Commonsense Reasoning with a …

200 papers

Multi-modal large language models(MLLMs) have achieved remarkable progress and demonstrated powerful knowledge comprehension and reasoning abilities. However, the mastery of domain-specific knowledge, which is essential for evaluating the…

Computation and Language · Computer Science 2024-05-09 Zheqi He , Xinya Wu , Pengfei Zhou , Richeng Xuan , Guang Liu , Xi Yang , Qiannan Zhu , Hua Huang

Large language models (LLMs) have obtained promising results in mathematical reasoning, which is a foundational skill for human intelligence. Most previous studies focus on improving and measuring the performance of LLMs based on textual…

Computation and Language · Computer Science 2024-11-04 Wentao Liu , Qianjun Pan , Yi Zhang , Zhuo Liu , Ji Wu , Jie Zhou , Aimin Zhou , Qin Chen , Bo Jiang , Liang He

We introduce CHARM, the first benchmark for comprehensively and in-depth evaluating the commonsense reasoning ability of large language models (LLMs) in Chinese, which covers both globally known and Chinese-specific commonsense. We…

Computation and Language · Computer Science 2024-12-11 Jiaxing Sun , Weiquan Huang , Jiang Wu , Chenya Gu , Wei Li , Songyang Zhang , Hang Yan , Conghui He

As the capabilities of large language models (LLMs) continue to advance, evaluating their performance becomes increasingly crucial and challenging. This paper aims to bridge this gap by introducing CMMLU, a comprehensive Chinese benchmark…

Computation and Language · Computer Science 2024-01-19 Haonan Li , Yixuan Zhang , Fajri Koto , Yifei Yang , Hai Zhao , Yeyun Gong , Nan Duan , Timothy Baldwin

Currently, long-chain reasoning remains a key challenge for large language models (LLMs) because natural texts lack sufficient explicit reasoning data. However, existing benchmarks suffer from limitations such as narrow coverage, short…

Computation and Language · Computer Science 2025-05-20 Weidong Zhan , Yue Wang , Nan Hu , Liming Xiao , Jingyuan Ma , Yuhang Qin , Zheng Li , Yixin Yang , Sirui Deng , Jinkun Ding , Wenhan Ma , Rui Li , Weilin Luo , Qun Liu , Zhifang Sui

Large language models (LLMs) have demonstrated strong reasoning abilities across specialized domains, motivating research into their application to legal reasoning. However, existing legal benchmarks often conflate factual recall with…

Artificial Intelligence · Computer Science 2025-11-21 Wenhan Yu , Xinbo Lin , Lanxin Ni , Jinhua Cheng , Lei Sha

Real-world decision-making often requires integrating and reasoning over information from multiple modalities. While recent multimodal large language models (MLLMs) have shown promise in such tasks, their ability to perform multi-hop…

Computation and Language · Computer Science 2025-06-02 Seunghee Kim , Changhyeon Kim , Taeuk Kim

As the capabilities of large multimodal models (LMMs) continue to advance, evaluating the performance of LMMs emerges as an increasing need. Additionally, there is an even larger gap in evaluating the advanced knowledge and reasoning…

Temporal reasoning is fundamental to human cognition and is crucial for various real-world applications. While recent advances in Large Language Models have demonstrated promising capabilities in temporal reasoning, existing benchmarks…

Computation and Language · Computer Science 2025-02-25 Zhenglin Wang , Jialong Wu , Pengfei LI , Yong Jiang , Deyu Zhou

In this study, we introduced a new benchmark consisting of a curated dataset and a defined evaluation process to assess the compositional reasoning capabilities of large language models within the chemistry domain. We designed and validated…

Computation and Language · Computer Science 2025-08-07 Mohammad Khodadad , Ali Shiraee Kasmaee , Mahdi Astaraki , Nicholas Sherck , Hamidreza Mahyar , Soheila Samiee

What a large language model (LLM) would respond in ethically relevant context? In this paper, we curate a large benchmark CMoralEval for morality evaluation of Chinese LLMs. The data sources of CMoralEval are two-fold: 1) a Chinese TV…

Computation and Language · Computer Science 2024-08-20 Linhao Yu , Yongqi Leng , Yufei Huang , Shang Wu , Haixin Liu , Xinmeng Ji , Jiahui Zhao , Jinwang Song , Tingting Cui , Xiaoqing Cheng , Tao Liu , Deyi Xiong

Recent advancements in large language models (LLMs) have transformed the field of question answering (QA). However, evaluating LLMs in the medical field is challenging due to the lack of standardized and comprehensive datasets. To address…

Computation and Language · Computer Science 2023-10-24 Junling Liu , Peilin Zhou , Yining Hua , Dading Chong , Zhongyu Tian , Andrew Liu , Helin Wang , Chenyu You , Zhenhua Guo , Lei Zhu , Michael Lingzhi Li

Recent advancements in reasoning-reinforced Large Language Models (LLMs) have shown remarkable capabilities in complex reasoning tasks. However, the mechanism underlying their utilization of different human reasoning skills remains poorly…

Computation and Language · Computer Science 2025-08-15 Nghia Trung Ngo , Franck Dernoncourt , Thien Huu Nguyen

Developing Large Language Models (LLMs) with robust long-context capabilities has been the recent research focus, resulting in the emergence of long-context LLMs proficient in Chinese. However, the evaluation of these models remains…

Computation and Language · Computer Science 2024-10-17 Zexuan Qiu , Jingjing Li , Shijue Huang , Xiaoqi Jiao , Wanjun Zhong , Irwin King

This paper introduces ConceptMath, a bilingual (English and Chinese), fine-grained benchmark that evaluates concept-wise mathematical reasoning of Large Language Models (LLMs). Unlike traditional benchmarks that evaluate general…

Computation and Language · Computer Science 2024-02-26 Yanan Wu , Jie Liu , Xingyuan Bu , Jiaheng Liu , Zhanhui Zhou , Yuanxing Zhang , Chenchen Zhang , Zhiqi Bai , Haibin Chen , Tiezheng Ge , Wanli Ouyang , Wenbo Su , Bo Zheng

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 110k college-level questions spanning 300…

Alignment has become a critical step for instruction-tuned Large Language Models (LLMs) to become helpful assistants. However, the effective evaluation of alignment for emerging Chinese LLMs is still largely unexplored. To fill in this gap,…

To thoroughly assess the mathematical reasoning abilities of Large Language Models (LLMs), we need to carefully curate evaluation datasets covering diverse mathematical concepts and mathematical problems at different difficulty levels. In…

Computation and Language · Computer Science 2024-09-09 Yan Liu , Renren Jin , Ling Shi , Zheng Yao , Deyi Xiong

Cross-modal reasoning (CMR), the intricate process of synthesizing and drawing inferences across divergent sensory modalities, is increasingly recognized as a crucial capability in the progression toward more sophisticated and…

Computation and Language · Computer Science 2024-10-01 Shengsheng Qian , Zuyi Zhou , Dizhan Xue , Bing Wang , Changsheng Xu

Large Language Models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks. However, Chinese LLMs face unique challenges, primarily due to the dominance of unstructured free text and the lack of…

Computation and Language · Computer Science 2025-10-08 Chengwei Wu , Jiapu Wang , Mingyang Gao , Xingrui Zhuo , Jipeng Guo , Runlin Lei , Haoran Luo , Tianyu Chen , Haoyi Zhou , Shirui Pan , Zechao Li
‹ Prev 1 2 3 10 Next ›