English
Related papers

Related papers: ANGO: A Next-Level Evaluation Benchmark For Genera…

200 papers

Large language models (LLMs), renowned for their impressive capabilities in various tasks, have significantly advanced artificial intelligence. Yet, these advancements have raised growing concerns about privacy and security implications. To…

Artificial Intelligence · Computer Science 2024-03-28 Yuqi Yang , Xiaowen Huang , Jitao Sang

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex problems. However, existing benchmarks for evaluating LLMs'…

Retrieval-Augmented Generation (RAG) is a promising approach for mitigating the hallucination of large language models (LLMs). However, existing research lacks rigorous evaluation of the impact of retrieval-augmented generation on different…

Computation and Language · Computer Science 2023-12-21 Jiawei Chen , Hongyu Lin , Xianpei Han , Le Sun

Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We present MedBench v4, a nationwide, cloud-based benchmarking…

Code-mixing is increasingly prevalent in interactions between humans and large language models, yet existing work often reduces it to a translation or convertibility problem, making it difficult to assess whether a model's switching…

Computation and Language · Computer Science 2026-01-26 Qingyan Yang , Tongxi Wang , Yunsheng Luo

Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applications demand high standards of accuracy, reliability, and…

Computation and Language · Computer Science 2024-11-27 Haitao Li , You Chen , Qingyao Ai , Yueyue Wu , Ruizhe Zhang , Yiqun Liu

Large language models (LLMs) are increasingly deployed in cost-sensitive and on-device scenarios, and safety guardrails have advanced mainly in English. However, real-world Chinese malicious queries typically conceal intent via homophones,…

Computation and Language · Computer Science 2026-01-06 Zhenhong Zhou , Shilinlu Yan , Chuanpu Liu , Qiankun Li , Kun Wang , Zhigang Zeng

Recent NLP tasks have benefited a lot from pre-trained language models (LM) since they are able to encode knowledge of various aspects. However, current LM evaluations focus on downstream performance, hence lack to comprehensively inspect…

Computation and Language · Computer Science 2020-12-01 Zhiruo Wang , Renfen Hu

Sign language research has achieved significant progress due to the advances in large language models (LLMs). However, the intrinsic ability of LLMs to understand sign language, especially in multimodal contexts, remains underexplored. To…

Computation and Language · Computer Science 2026-04-27 Rui Zhao , Xuewen Zhong , Xiaoyun Zheng , Jinsong Su , Yidong Chen

Large Language Models (LLMs) have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition (ASR) is becoming a mainstream paradigm. Building upon this momentum, our research delves…

Recent strides in Large Language Models (LLMs) have saturated many Natural Language Processing (NLP) benchmarks, emphasizing the need for more challenging ones to properly assess LLM capabilities. However, domain-specific and multilingual…

Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow strictly from stated premises. Many existing logical-reasoning benchmarks are generated by…

We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation targets three main…

Sound · Computer Science 2025-05-07 Bin Wang , Xunlong Zou , Geyu Lin , Shuo Sun , Zhuohan Liu , Wenyu Zhang , Zhengyuan Liu , AiTi Aw , Nancy F. Chen

In this paper, we introduce a novel psychological benchmark, CPsyExam, constructed from questions sourced from Chinese language examinations. CPsyExam is designed to prioritize psychological knowledge and case analysis separately,…

Computation and Language · Computer Science 2024-12-11 Jiahao Zhao , Jingwei Zhu , Minghuan Tan , Min Yang , Renhao Li , Di Yang , Chenhao Zhang , Guancheng Ye , Chengming Li , Xiping Hu , Derek F. Wong

The effective assessment of the instruction-following ability of large language models (LLMs) is of paramount importance. A model that cannot adhere to human instructions might be not able to provide reliable and helpful responses. In…

Computation and Language · Computer Science 2023-11-17 Yimin Jing , Renren Jin , Jiahao Hu , Huishi Qiu , Xiaohua Wang , Peng Wang , Deyi Xiong

We introduce ClarQ-LLM, an evaluation framework consisting of bilingual English-Chinese conversation tasks, conversational agents and evaluation metrics, designed to serve as a strong benchmark for assessing agents' ability to ask…

Computation and Language · Computer Science 2024-09-17 Yujian Gan , Changling Li , Jinxia Xie , Luou Wen , Matthew Purver , Massimo Poesio

Large language models (LLMs) have substantially advanced machine translation (MT), yet their effectiveness in translating web novels remains unclear. Existing benchmarks rely on surface-level metrics that fail to capture the distinctive…

Computation and Language · Computer Science 2025-10-14 Enze Zhang , Jiaying Wang , Mengxi Xiao , Jifei Liu , Ziyan Kuang , Rui Dong , Eric Dong , Sophia Ananiadou , Min Peng , Qianqian Xie

The recent advances in natural language processing (NLP), have led to a new trend of applying large language models (LLMs) to real-world scenarios. While the latest LLMs are astonishingly fluent when interacting with humans, they suffer…

Computation and Language · Computer Science 2023-10-27 Tong Xiang , Liangzhi Li , Wangyue Li , Mingbai Bai , Lu Wei , Bowen Wang , Noa Garcia

Metaphors are common in everyday language, and the identification and understanding of metaphors are facilitated by models to achieve a better understanding of the text. Metaphors are mainly identified and generated by pre-trained models in…

Computation and Language · Computer Science 2024-08-20 Jie Wang , Jin Wang , Xuejie Zhang

Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural language processing and understanding. Benchmark evaluations are crucial for assessing the…

Computation and Language · Computer Science 2026-05-12 Wenbo Zhang , Hengrui Cai , Wenyu Chen