English
Related papers

Related papers: Chinese SimpleQA: A Chinese Factuality Evaluation …

200 papers

The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models' abilities. These benchmarks consist of tens of thousands of examples making evaluation of LLMs very…

Computation and Language · Computer Science 2024-05-28 Felipe Maia Polo , Lucas Weber , Leshem Choshen , Yuekai Sun , Gongjun Xu , Mikhail Yurochkin

In this paper, we present Edu-Values, the first Chinese education values evaluation benchmark that includes seven core values: professional philosophy, teachers' professional ethics, education laws and regulations, cultural literacy,…

Computation and Language · Computer Science 2026-03-24 Peiyi Zhang , Yazhou Zhang , Bo Wang , Lu Rong , Prayag Tiwari , Jing Qin

As the capabilities of large multimodal models (LMMs) continue to advance, evaluating the performance of LMMs emerges as an increasing need. Additionally, there is an even larger gap in evaluating the advanced knowledge and reasoning…

What a large language model (LLM) would respond in ethically relevant context? In this paper, we curate a large benchmark CMoralEval for morality evaluation of Chinese LLMs. The data sources of CMoralEval are two-fold: 1) a Chinese TV…

Computation and Language · Computer Science 2024-08-20 Linhao Yu , Yongqi Leng , Yufei Huang , Shang Wu , Haixin Liu , Xinmeng Ji , Jiahui Zhao , Jinwang Song , Tingting Cui , Xiaoqing Cheng , Tao Liu , Deyi Xiong

The emergence of Large Language Models (LLMs) in the medical domain has stressed a compelling need for standard datasets to evaluate their question-answering (QA) performance. Although there have been several benchmark datasets for medical…

Computation and Language · Computer Science 2025-03-18 Qian Zhang , Panfeng Chen , Jiali Li , Linkun Feng , Shuyu Liu , Heng Zhao , Mei Chen , Hui Li , Yanhao Wang

Large Language Models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks. However, Chinese LLMs face unique challenges, primarily due to the dominance of unstructured free text and the lack of…

Computation and Language · Computer Science 2025-10-08 Chengwei Wu , Jiapu Wang , Mingyang Gao , Xingrui Zhuo , Jipeng Guo , Runlin Lei , Haoran Luo , Tianyu Chen , Haoyi Zhou , Shirui Pan , Zechao Li

Large language models (LLMs) have achieved remarkable performance on various NLP tasks, yet their potential in more challenging and domain-specific task, such as finance, has not been fully explored. In this paper, we present CFinBench: a…

Computation and Language · Computer Science 2024-07-03 Ying Nie , Binwei Yan , Tianyu Guo , Hao Liu , Haoyu Wang , Wei He , Binfan Zheng , Weihao Wang , Qiang Li , Weijian Sun , Yunhe Wang , Dacheng Tao

Language acquisition is vital to revealing the nature of human language intelligence and has recently emerged as a promising perspective for improving the interpretability of large language models (LLMs). However, it is ethically and…

Computation and Language · Computer Science 2025-11-20 Qihao Yang , Xuelin Wang , Jiale Chen , Xuelian Dong , Yuxin Hao , Tianyong Hao

This paper introduces AQA-Bench, a novel benchmark to assess the sequential reasoning capabilities of large language models (LLMs) in algorithmic contexts, such as depth-first search (DFS). The key feature of our evaluation benchmark lies…

Computation and Language · Computer Science 2025-06-23 Siwei Yang , Bingchen Zhao , Cihang Xie

The increased use of large language models (LLMs) across a variety of real-world applications calls for automatic tools to check the factual accuracy of their outputs, as LLMs often hallucinate. This is difficult as it requires assessing…

Computation and Language · Computer Science 2025-10-30 Hasan Iqbal , Yuxia Wang , Minghan Wang , Georgi Georgiev , Jiahui Geng , Iryna Gurevych , Preslav Nakov

In this thesis, we investigated the relevance, faithfulness, and succinctness aspects of Long Form Question Answering (LFQA). LFQA aims to generate an in-depth, paragraph-length answer for a given question, to help bridge the gap between…

Computation and Language · Computer Science 2022-11-16 Dan Su

The success of ChatGPT validates the potential of large language models (LLMs) in artificial general intelligence (AGI). Subsequently, the release of LLMs has sparked the open-source community's interest in instruction-tuning, which is…

Computation and Language · Computer Science 2023-10-23 Qingyi Si , Tong Wang , Zheng Lin , Xu Zhang , Yanan Cao , Weiping Wang

While Large Language Models (LLMs) excel in various general domains, they exhibit notable gaps in the highly specialized, knowledge-intensive, and legally regulated Chinese tax domain. Consequently, while tax-related benchmarks are gaining…

Computation and Language · Computer Science 2026-04-23 Gang Hu , Yating Chen , Haiyan Ding , Wang Gao , Jiajia Huang , Min Peng , Qianqian Xie , Kun Yue

Large language models (LLMs) excel at many general-purpose natural language processing tasks. However, their ability to perform deep reasoning and mathematical analysis, particularly for complex tasks as required in cryptography, remains…

Cryptography and Security · Computer Science 2025-12-03 Mayar Elfares , Pascal Reisert , Tilman Dietz , Manpa Barman , Ahmed Zaki , Ralf Küsters , Andreas Bulling

In recent years, large language models (LLMs) have achieved remarkable advancements in multimodal processing, including end-to-end speech-based language models that enable natural interactions and perform specific tasks in task-oriented…

Computation and Language · Computer Science 2025-08-15 Enzhi Wang , Qicheng Li , Shiwan Zhao , Aobo Kong , Jiaming Zhou , Xi Yang , Yequan Wang , Yonghua Lin , Yong Qin

In this paper, we introduce a novel psychological benchmark, CPsyExam, constructed from questions sourced from Chinese language examinations. CPsyExam is designed to prioritize psychological knowledge and case analysis separately,…

Computation and Language · Computer Science 2024-12-11 Jiahao Zhao , Jingwei Zhu , Minghuan Tan , Min Yang , Renhao Li , Di Yang , Chenhao Zhang , Guancheng Ye , Chengming Li , Xiping Hu , Derek F. Wong

The rapid expansion of context length in large language models (LLMs) has outpaced existing evaluation benchmarks. Current long-context benchmarks often trade off scalability and realism: synthetic tasks underrepresent real-world…

Computation and Language · Computer Science 2026-01-07 Ziyang Chen , Xing Wu , Junlong Jia , Chaochen Gao , Qi Fu , Debing Zhang , Songlin Hu

Recent literature has shown that large language models (LLMs) are generally excellent few-shot reasoners to solve text reasoning tasks. However, the capability of LLMs on table reasoning tasks is yet to be explored. In this paper, we aim at…

Computation and Language · Computer Science 2023-01-24 Wenhu Chen

Recent NLP tasks have benefited a lot from pre-trained language models (LM) since they are able to encode knowledge of various aspects. However, current LM evaluations focus on downstream performance, hence lack to comprehensively inspect…

Computation and Language · Computer Science 2020-12-01 Zhiruo Wang , Renfen Hu

Smart cities need the involvement of their residents to enhance quality of life. Conversational query-answering is an emerging approach for user engagement. There is an increasing demand of an advanced conversational question-answering that…

Computation and Language · Computer Science 2024-04-16 Pardis Moradbeiki , Nasser Ghadiri