English
Related papers

Related papers: E-EVAL: A Comprehensive Chinese K-12 Education Eva…

200 papers

The advancement of large language models (LLMs) has enhanced the ability to generalize across a wide range of unseen natural language processing (NLP) tasks through instruction-following. Yet, their effectiveness often diminishes in…

Chinese Large Language Models (LLMs) have recently demonstrated impressive capabilities across various NLP benchmarks and real-world applications. However, the existing benchmarks for comprehensively evaluating these LLMs are still…

Computation and Language · Computer Science 2024-03-20 Chuang Liu , Renren Jin , Yuqi Ren , Deyi Xiong

Large language models (LLMs) excel in high-resource languages but struggle with low-resource languages (LRLs), particularly those spoken by minority communities in China, such as Tibetan, Uyghur, Kazakh, and Mongolian. To systematically…

Computation and Language · Computer Science 2025-06-03 Chen Zhang , Mingxu Tao , Zhiyuan Liao , Yansong Feng

Multimodal large language models have demonstrated remarkable reasoning capabilities in various visual tasks. However, their abilities in K12 scenarios are still systematically underexplored. Previous studies suffer from various limitations…

Artificial Intelligence · Computer Science 2025-06-03 Chong Li , Chenglin Zhu , Tao Zhang , Mingan Lin , Zenan Zhou , Jian Xie

Large Language Models (LLMs) have transformed artificial intelligence, offering profound opportunities for educational applications. However, their ability to provide fine-grained educational feedback for K-12 English writing remains…

Computation and Language · Computer Science 2025-12-01 Jingheng Ye , Shen Wang , Jiaqi Chen , Hebin Wang , Deqing Zou , Yanyu Zhu , Jiwei Tang , Hai-Tao Zheng , Ruitong Liu , Haoyang Li , Yanfeng Wang , Qingsong Wen

Recently, there has been growing interest in extending the context length of large language models (LLMs), aiming to effectively process long inputs of one turn or conversations with more extensive histories. While proprietary models such…

Computation and Language · Computer Science 2023-10-05 Chenxin An , Shansan Gong , Ming Zhong , Xingjian Zhao , Mukai Li , Jun Zhang , Lingpeng Kong , Xipeng Qiu

Large language models (LLMs) have demonstrated exceptional capabilities in general domains, yet their application in highly specialized and culturally-rich fields like Traditional Chinese Medicine (TCM) requires rigorous and nuanced…

Computation and Language · Computer Science 2025-11-18 Tianai Huang , Jiayuan Chen , Lu Lu , Pengcheng Chen , Tianbin Li , Bing Han , Wenchao Tang , Jie Xu , Ming Li

Pretrained Language Models (PLMs) have achieved tremendous success in natural language understanding tasks. While different learning schemes -- fine-tuning, zero-shot, and few-shot learning -- have been widely explored and compared for…

Computation and Language · Computer Science 2021-09-30 Liang Xu , Xiaojing Lu , Chenyang Yuan , Xuanwei Zhang , Huilin Xu , Hu Yuan , Guoao Wei , Xiang Pan , Xin Tian , Libo Qin , Hu Hai

State-of-the-art large language models (LLMs) are now claiming remarkable supported context lengths of 256k or even more. In contrast, the average context lengths of mainstream benchmarks are insufficient (5k-21k), and they suffer from…

Computation and Language · Computer Science 2025-10-23 Tao Yuan , Xuefei Ning , Dong Zhou , Zhijie Yang , Shiyao Li , Minghui Zhuang , Zheyue Tan , Zhuyu Yao , Dahua Lin , Boxun Li , Guohao Dai , Shengen Yan , Yu Wang

Large language models (LLMs) have obtained promising results in mathematical reasoning, which is a foundational skill for human intelligence. Most previous studies focus on improving and measuring the performance of LLMs based on textual…

Computation and Language · Computer Science 2024-11-04 Wentao Liu , Qianjun Pan , Yi Zhang , Zhuo Liu , Ji Wu , Jie Zhou , Aimin Zhou , Qin Chen , Bo Jiang , Liang He

Large language models (LLMs) are increasingly applied in computer science education for tasks such as tutoring, content generation, and code assessment. However, systematic evaluations aligned with formal curricula and certification…

The need to evaluate instructional materials for K-12 science education has become increasingly important, as more educators use generative AI to create instructional materials. However, the review of instructional materials is…

Artificial Intelligence · Computer Science 2026-04-29 Zhaohui Li , Peng He , Zhiyuan Chen , Honglu Liu , Zeyuan Wang , Tingting Li , Jinjun Xiong

The rapid development of Large Language Models (LLMs) in vertical domains, including intellectual property (IP), lacks a specific evaluation benchmark for assessing their understanding, application, and reasoning abilities. To fill this…

Computation and Language · Computer Science 2024-06-19 Qiyao Wang , Jianguo Huang , Shule Lu , Yuan Lin , Kan Xu , Liang Yang , Hongfei Lin

Large language models (LLMs) excel in various NLP tasks and modern medicine, but their evaluation in traditional Chinese medicine (TCM) is underexplored. To address this, we introduce TCM3CEval, a benchmark assessing LLMs in TCM across…

Computation and Language · Computer Science 2025-03-11 Tianai Huang , Lu Lu , Jiayuan Chen , Lihao Liu , Junjun He , Yuping Zhao , Wenchao Tang , Jie Xu

With the continuous emergence of Chinese Large Language Models (LLMs), how to evaluate a model's capabilities has become an increasingly significant issue. The absence of a comprehensive Chinese benchmark that thoroughly assesses a model's…

Computation and Language · Computer Science 2023-10-17 Yanyang Li , Jianqiao Zhao , Duo Zheng , Zi-Yuan Hu , Zhi Chen , Xiaohui Su , Yongfeng Huang , Shijia Huang , Dahua Lin , Michael R. Lyu , Liwei Wang

With the emergence of more and more economy-specific LLMS, how to measure whether they can be safely invested in production becomes a problem. Previous research has primarily focused on evaluating the performance of LLMs within specific…

Artificial Intelligence · Computer Science 2024-08-21 Xinyu Liu , Ke Jin

Chinese essay writing and its evaluation are critical in educational contexts, yet the capabilities of Large Language Models (LLMs) in this domain remain largely underexplored. Existing benchmarks often rely on coarse-grained text quality…

Computation and Language · Computer Science 2025-06-04 Fan Gao , Dongyuan Li , Ding Xia , Fei Mi , Yasheng Wang , Lifeng Shang , Baojun Wang

Large language models (LLMs) have performed remarkably well in various natural language processing tasks by benchmarking, including in the Western medical domain. However, the professional evaluation benchmarks for LLMs have yet to be…

Computation and Language · Computer Science 2024-06-04 Wenjing Yue , Xiaoling Wang , Wei Zhu , Ming Guan , Huanran Zheng , Pengfei Wang , Changzhi Sun , Xin Ma

Text-to-Table aims to generate structured tables to convey the key information from unstructured documents. Existing text-to-table datasets are typically oriented English, limiting the research in non-English languages. Meanwhile, the…

Computation and Language · Computer Science 2024-05-21 Haoxiang Shi , Jiaan Wang , Jiarong Xu , Cen Wang , Tetsuya Sakai

What a large language model (LLM) would respond in ethically relevant context? In this paper, we curate a large benchmark CMoralEval for morality evaluation of Chinese LLMs. The data sources of CMoralEval are two-fold: 1) a Chinese TV…

Computation and Language · Computer Science 2024-08-20 Linhao Yu , Yongqi Leng , Yufei Huang , Shang Wu , Haixin Liu , Xinmeng Ji , Jiahui Zhao , Jinwang Song , Tingting Cui , Xiaoqing Cheng , Tao Liu , Deyi Xiong