English
Related papers

Related papers: Benchmarking Chinese Commonsense Reasoning with a …

200 papers

Visual Language Models (VLMs) are powerful generative tools but often produce factually inaccurate outputs due to a lack of robust reasoning capabilities. While extensive research has been conducted on integrating external knowledge for…

Artificial Intelligence · Computer Science 2025-11-26 Shamima Hossain

Traditional Chinese Medicine (TCM), as an effective alternative medicine, has been receiving increasing attention. In recent years, the rapid development of large language models (LLMs) tailored for TCM has highlighted the urgent need for…

Computation and Language · Computer Science 2025-10-28 Jiacheng Xie , Yang Yu , Ziyang Zhang , Shuai Zeng , Jiaxuan He , Ayush Vasireddy , Xiaoting Tang , Congyu Guo , Lening Zhao , Congcong Jing , Guanghui An , Dong Xu

Existing humor datasets and evaluations predominantly focus on English, leaving limited resources for culturally nuanced humor in non-English languages like Chinese. To address this gap, we construct Chumor, the first Chinese humor…

Computation and Language · Computer Science 2024-12-24 Ruiqi He , Yushu He , Longju Bai , Jiarui Liu , Zhenjie Sun , Zenghao Tang , He Wang , Hanchen Xia , Rada Mihalcea , Naihao Deng

Large Language Models (LLMs) provide a possibility to make a great breakthrough in medicine. The establishment of a standardized medical benchmark becomes a fundamental cornerstone to measure progression. However, medical environments in…

Computation and Language · Computer Science 2024-04-05 Xidong Wang , Guiming Hardy Chen , Dingjie Song , Zhiyi Zhang , Zhihong Chen , Qingying Xiao , Feng Jiang , Jianquan Li , Xiang Wan , Benyou Wang , Haizhou Li

We present ACCORD, a framework and benchmark suite for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) through controlled, multi-hop counterfactuals. ACCORD introduces formal elements to…

Artificial Intelligence · Computer Science 2025-02-10 François Roewer-Després , Jinyue Feng , Zining Zhu , Frank Rudzicz

As large language models (LLMs) evolve into tool-using agents, the ability to browse the web in real-time has become a critical yardstick for measuring their reasoning and retrieval competence. Existing benchmarks such as BrowseComp…

Existing Multimodal Large Language Models (MLLMs) are predominantly trained and tested on consistent visual-textual inputs, leaving open the question of whether they can handle inconsistencies in real-world, layout-rich content. To bridge…

Computation and Language · Computer Science 2025-06-12 Qianqi Yan , Yue Fan , Hongquan Li , Shan Jiang , Yang Zhao , Xinze Guan , Ching-Chen Kuo , Xin Eric Wang

New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present C-Eval, the first comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of…

Computation and Language · Computer Science 2023-11-07 Yuzhen Huang , Yuzhuo Bai , Zhihao Zhu , Junlei Zhang , Jinghan Zhang , Tangjun Su , Junteng Liu , Chuancheng Lv , Yikai Zhang , Jiayi Lei , Yao Fu , Maosong Sun , Junxian He

While Large Language Models (LLMs) excel in question-answering (QA) tasks, their real reasoning abilities on multiple evidence retrieval and integration on Multi-hop QA tasks remain less explored. Firstly, LLMs sometimes generate answers…

Computation and Language · Computer Science 2024-10-16 Jian Wu , Linyi Yang , Zhen Wang , Manabu Okumura , Yue Zhang

With the rapid evolution of large language models (LLMs), there is a growing concern that they may pose risks or have negative social impacts. Therefore, evaluation of human values alignment is becoming increasingly important. Previous work…

Computation and Language · Computer Science 2023-07-20 Guohai Xu , Jiayi Liu , Ming Yan , Haotian Xu , Jinghui Si , Zhuoran Zhou , Peng Yi , Xing Gao , Jitao Sang , Rong Zhang , Ji Zhang , Chao Peng , Fei Huang , Jingren Zhou

Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applications demand high standards of accuracy, reliability, and…

Computation and Language · Computer Science 2024-11-27 Haitao Li , You Chen , Qingyao Ai , Yueyue Wu , Ruizhe Zhang , Yiqun Liu

As Large Language Models (LLMs) are increasingly popularized in the multilingual world, ensuring hallucination-free factuality becomes markedly crucial. However, existing benchmarks for evaluating the reliability of Multimodal Large…

Computation and Language · Computer Science 2026-01-28 Yexing Du , Kaiyuan Liu , Youcheng Pan , Zheng Chu , Bo Yang , Xiaocheng Feng , Ming Liu , Yang Xiang

Despite significant advancements, current large language models (LLMs) and vision-language models (LVLMs) continue to struggle with complex, multi-step, cross-modal common sense reasoning tasks, often exhibiting a lack of "deliberative…

Computation and Language · Computer Science 2025-08-06 Wenjie Luo , Ruocheng Li , Shanshan Zhu , Julian Perry

Logical reasoning is a fundamental aspect of human intelligence and an essential capability for multimodal large language models (MLLMs). Despite the significant advancement in multimodal reasoning, existing benchmarks fail to…

Artificial Intelligence · Computer Science 2025-05-28 Jiakang Yuan , Tianshuo Peng , Yilei Jiang , Yiting Lu , Renrui Zhang , Kaituo Feng , Chaoyou Fu , Tao Chen , Lei Bai , Bo Zhang , Xiangyu Yue

Large Language Models (LLMs) are increasingly integrated into real-world decision-making, including in the domain of public policy. Yet, their ability to comprehend and reason about policy-related content remains underexplored. To fill this…

Computation and Language · Computer Science 2026-04-15 Han Bao , Penghao Zhang , Yue Huang , Zhengqing Yuan , Yanchi Ru , Rui Su , Yujun Zhou , Xiangqi Wang , Kehan Guo , Nitesh V Chawla , Yanfang Ye , Xiangliang Zhang

The rapid development of multimodal large language models (MLLMs) raises the question of how they compare to human performance. While existing datasets often feature synthetic or overly simplistic tasks, some models have already surpassed…

Computation and Language · Computer Science 2025-10-16 Zichen Zhu , Yang Xu , Lu Chen , Jingkai Yang , Yichuan Ma , Yiming Sun , Hailin Wen , Jiaqi Liu , Jinyu Cai , Yingzi Ma , Situo Zhang , Zihan Zhao , Liangtai Sun , Kai Yu

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench,…

Computation and Language · Computer Science 2024-10-07 Zetian Ouyang , Yishuai Qiu , Linlin Wang , Gerard de Melo , Ya Zhang , Yanfeng Wang , Liang He

The advancement of large language models (LLMs) has enhanced the ability to generalize across a wide range of unseen natural language processing (NLP) tasks through instruction-following. Yet, their effectiveness often diminishes in…

The recently unprecedented advancements in Large Language Models (LLMs) have propelled the medical community by establishing advanced medical-domain models. However, due to the limited collection of medical datasets, there are only a few…

Computation and Language · Computer Science 2024-06-10 Ping Yu , Kaitao Song , Fengchen He , Ming Chen , Jianfeng Lu

We present FinMMR, a novel bilingual multimodal benchmark tailored to evaluate the reasoning capabilities of multimodal large language models (MLLMs) in financial numerical reasoning tasks. Compared to existing benchmarks, our work…