English
Related papers

Related papers: EssayBench: Evaluating Large Language Models in Mu…

200 papers

Large Language Models (LLMs) have demonstrated significant potential and effectiveness across multiple application domains. To assess the performance of mainstream LLMs in public security tasks, this study aims to construct a specialized…

Artificial Intelligence · Computer Science 2024-03-22 Xin Tong , Bo Jin , Zhi Lin , Binjun Wang , Ting Yu , Qiang Cheng

Multi-modal large language models(MLLMs) have achieved remarkable progress and demonstrated powerful knowledge comprehension and reasoning abilities. However, the mastery of domain-specific knowledge, which is essential for evaluating the…

Computation and Language · Computer Science 2024-05-09 Zheqi He , Xinya Wu , Pengfei Zhou , Richeng Xuan , Guang Liu , Xi Yang , Qiannan Zhu , Hua Huang

Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applications demand high standards of accuracy, reliability, and…

Computation and Language · Computer Science 2024-11-27 Haitao Li , You Chen , Qingyao Ai , Yueyue Wu , Ruizhe Zhang , Yiqun Liu

Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with specialized subdomains…

Artificial Intelligence · Computer Science 2026-01-16 Zixun Lan , Maochun Xu , Yifan Ren , Rui Wu , Jianghui Zhou , Xueyang Cheng , Jianan Ding Ding , Xinheng Wang , Mingmin Chi , Fei Ma

With the rapid development of Large language models (LLMs), understanding the capabilities of LLMs in identifying unsafe content has become increasingly important. While previous works have introduced several benchmarks to evaluate the…

Computation and Language · Computer Science 2025-04-15 Hengxiang Zhang , Hongfu Gao , Qiang Hu , Guanhua Chen , Lili Yang , Bingyi Jing , Hongxin Wei , Bing Wang , Haifeng Bai , Lei Yang

Evaluation frameworks for text summarization have evolved in terms of both domain coverage and metrics. However, existing benchmarks still lack domain-specific assessment criteria, remain predominantly English-centric, and face challenges…

Computation and Language · Computer Science 2025-06-03 Hyangsuk Min , Yuho Lee , Minjeong Ban , Jiaqi Deng , Nicole Hee-Yeon Kim , Taewon Yun , Hang Su , Jason Cai , Hwanjun Song

In the burgeoning field of large language models (LLMs), the assessment of fundamental knowledge remains a critical challenge, particularly for models tailored to Chinese language and culture. This paper introduces FoundaBench, a pioneering…

Computation and Language · Computer Science 2024-04-30 Wei Li , Ren Ma , Jiang Wu , Chenya Gu , Jiahui Peng , Jinyang Len , Songyang Zhang , Hang Yan , Dahua Lin , Conghui He

The application of large language models (LLMs) has achieved remarkable success in various fields, but their effectiveness in specialized domains like the Chinese insurance industry remains underexplored. The complexity of insurance…

Computation and Language · Computer Science 2025-01-22 Jing Ding , Kai Feng , Binbin Lin , Jiarui Cai , Qiushi Wang , Yu Xie , Xiaojin Zhang , Zhongyu Wei , Wei Chen

This paper introduces ChineseVideoBench, a pioneering benchmark specifically designed for evaluating Multimodal Large Language Models (MLLMs) in Chinese Video Question Answering. The growing demand for sophisticated video analysis…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yuxiang Nie , Han Wang , Yongjie Ye , Haiyang Yu , Weitao Jia , Tao Zeng , Hao Feng , Xiang Fei , Yang Li , Xiaohui Lv , Guozhi Tang , Jingqun Tang , Jinghui Lu , Zehui Dai , Jiacong Wang , Dingkang Yang , An-Lan Wang , Can Huang

The advancement of large language models (LLMs) has enhanced the ability to generalize across a wide range of unseen natural language processing (NLP) tasks through instruction-following. Yet, their effectiveness often diminishes in…

Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the overall efficiency and accuracy of content review. However,…

Computation and Language · Computer Science 2025-08-14 Kangwei Liu , Siyuan Cheng , Bozhong Tian , Xiaozhuan Liang , Yuyang Yin , Meng Han , Ningyu Zhang , Bryan Hooi , Xi Chen , Shumin Deng

Chinese Large Language Models (LLMs) have recently demonstrated impressive capabilities across various NLP benchmarks and real-world applications. However, the existing benchmarks for comprehensively evaluating these LLMs are still…

Computation and Language · Computer Science 2024-03-20 Chuang Liu , Renren Jin , Yuqi Ren , Deyi Xiong

The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluation platforms for assessing Chinese LLMs, many of these…

Computation and Language · Computer Science 2024-03-20 Chuang Liu , Linhao Yu , Jiaxuan Li , Renren Jin , Yufei Huang , Ling Shi , Junhui Zhang , Xinmeng Ji , Tingting Cui , Tao Liu , Jinwang Song , Hongying Zan , Sun Li , Deyi Xiong

Evaluating the alignment capabilities of large Vision-Language Models (VLMs) is essential for determining their effectiveness as helpful assistants. However, existing benchmarks primarily focus on basic abilities using nonverbal methods,…

Computation and Language · Computer Science 2025-06-05 Yuhang Wu , Wenmeng Yu , Yean Cheng , Yan Wang , Xiaohan Zhang , Jiazheng Xu , Ming Ding , Yuxiao Dong

Large Language Models (LLMs) have demonstrated remarkable capabilities in mathematical problem-solving. However, the transition from providing answers to generating high-quality educational questions presents significant challenges that…

Computation and Language · Computer Science 2025-08-15 Chengliang Zhou , Mei Wang , Ting Zhang , Qiannan Zhu , Jian Li , Hua Huang

Large Language Models (LLMs) have demonstrated remarkable instruction-following capabilities across various applications. However, their performance in multilingual settings lacks systematic investigation, with existing evaluations lacking…

Computation and Language · Computer Science 2025-11-04 Zhenyu Li , Kehai Chen , Yunfei Long , Xuefeng Bai , Yaoyin Zhang , Xuchen Wei , Juntao Li , Min Zhang

While the capabilities of Large Language Models (LLMs) have been studied in both Simplified and Traditional Chinese, it is yet unclear whether LLMs exhibit differential performance when prompted in these two variants of written Chinese.…

Computation and Language · Computer Science 2025-05-29 Hanjia Lyu , Jiebo Luo , Jian Kang , Allison Koenecke

Ancient Chinese text processing presents unique challenges for large language models (LLMs) due to its distinct linguistic features, complex structural constraints, and rich cultural context. While existing benchmarks have primarily focused…

Computation and Language · Computer Science 2025-03-21 Shangqing Zhao , Yuhao Zhou , Yupei Ren , Zhe Chen , Chenghao Jia , Fang Zhe , Zhaogaung Long , Shu Liu , Man Lan

Chinese idioms (Chengyu) are concise four-character expressions steeped in history and culture, whose literal translations often fail to capture their full meaning. This complexity makes them challenging for language models to interpret and…

Computation and Language · Computer Science 2025-06-24 Yicheng Fu , Zhemin Huang , Liuxin Yang , Yumeng Lu , Zhongdongming Dai

The rapid expansion of context length in large language models (LLMs) has outpaced existing evaluation benchmarks. Current long-context benchmarks often trade off scalability and realism: synthetic tasks underrepresent real-world…

Computation and Language · Computer Science 2026-01-07 Ziyang Chen , Xing Wu , Junlong Jia , Chaochen Gao , Qi Fu , Debing Zhang , Songlin Hu