CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark
Abstract
Realizing general-purpose language intelligence has been a longstanding goal for natural language processing, where standard evaluation benchmarks play a fundamental and guiding role. We argue that for general-purpose language intelligence evaluation, the benchmark itself needs to be comprehensive and systematic. To this end, we propose CUGE, a Chinese Language Understanding and Generation Evaluation benchmark with the following features: (1) Hierarchical benchmark framework, where datasets are principally selected and organized with a language capability-task-dataset hierarchy. (2) Multi-level scoring strategy, where different levels of model performance are provided based on the hierarchical framework. To facilitate CUGE, we provide a public leaderboard that can be customized to support flexible model judging criteria. Evaluation results on representative pre-trained language models indicate ample room for improvement towards general-purpose language intelligence. CUGE is publicly available at cuge.baai.ac.cn.
Cite
@article{arxiv.2112.13610,
title = {CUGE: A Chinese Language Understanding and Generation Evaluation Benchmark},
author = {Yuan Yao and Qingxiu Dong and Jian Guan and Boxi Cao and Zhengyan Zhang and Chaojun Xiao and Xiaozhi Wang and Fanchao Qi and Junwei Bao and Jinran Nie and Zheni Zeng and Yuxian Gu and Kun Zhou and Xuancheng Huang and Wenhao Li and Shuhuai Ren and Jinliang Lu and Chengqiang Xu and Huadong Wang and Guoyang Zeng and Zile Zhou and Jiajun Zhang and Juanzi Li and Minlie Huang and Rui Yan and Xiaodong He and Xiaojun Wan and Xin Zhao and Xu Sun and Yang Liu and Zhiyuan Liu and Xianpei Han and Erhong Yang and Zhifang Sui and Maosong Sun},
journal= {arXiv preprint arXiv:2112.13610},
year = {2022}
}
Comments
We add two new datasets, including grammatical error correction dataset YACLC from Beijing Language and Culture University, and reading comprehension dataset GCRC from Shanxi University, and also improve the description consistency of all datasets