中文
相关论文

相关论文: WYWEB: A NLP Evaluation Benchmark For Classical Ch…

200 篇论文

Large language models (LLMs) excel in high-resource languages but struggle with low-resource languages (LRLs), particularly those spoken by minority communities in China, such as Tibetan, Uyghur, Kazakh, and Mongolian. To systematically…

计算与语言 · 计算机科学 2025-06-03 Chen Zhang , Mingxu Tao , Zhiyuan Liao , Yansong Feng

The impressive progress in NLP techniques has been driven by the development of multi-task benchmarks such as GLUE and SuperGLUE. While these benchmarks focus on tasks for one or two input sentences, there has been exciting work in…

计算与语言 · 计算机科学 2025-10-20 G Thomas Hudson , Noura Al Moubayed

Deriving inference from heterogeneous inputs (such as images, text, and audio) is an important skill for humans to perform day-to-day tasks. A similar ability is desirable for the development of advanced Artificial Intelligence (AI)…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Shailaja Keyur Sampat , Mutsumi Nakamura , Shankar Kailas , Kartik Aggarwal , Mandy Zhou , Yezhou Yang , Chitta Baral

To measure advances in retrieval, test collections with relevance judgments that can faithfully distinguish systems are required. This paper presents NeuCLIRBench, an evaluation collection for cross-language and multilingual retrieval. The…

Despite the major advances in NLP, significant disparities in NLP system performance across languages still exist. Arguably, these are due to uneven resource allocation and sub-optimal incentives to work on less resourced languages. To…

Modeling discourse -- the linguistic phenomena that go beyond individual sentences, is a fundamental yet challenging aspect of natural language processing (NLP). However, existing evaluation benchmarks primarily focus on the evaluation of…

计算与语言 · 计算机科学 2023-07-25 Longyue Wang , Zefeng Du , Donghuai Liu , Deng Cai , Dian Yu , Haiyun Jiang , Yan Wang , Leyang Cui , Shuming Shi , Zhaopeng Tu

The prevalence of rapidly evolving slang, neologisms, and highly stylized expressions in informal user-generated text, particularly on Chinese social media, poses significant challenges for Machine Translation (MT) benchmarking.…

计算与语言 · 计算机科学 2026-02-02 Kaiyan Zhao , Zheyong Xie , Zhongtao Miao , Xinze Lyu , Yao Hu , Shaosheng Cao

We introduce SuperCLUE-Math6(SC-Math6), a new benchmark dataset to evaluate the mathematical reasoning abilities of Chinese language models. SC-Math6 is designed as an upgraded Chinese version of the GSM8K dataset with enhanced difficulty,…

计算与语言 · 计算机科学 2024-02-05 Liang Xu , Hang Xue , Lei Zhu , Kangkang Zhao

The critical field of psychology necessitates a comprehensive benchmark to enhance the evaluation and development of domain-specific Large Language Models (LLMs). Existing MMLU-type benchmarks, such as C-EVAL and CMMLU, include…

计算与语言 · 计算机科学 2024-06-18 Junlei Zhang , Hongliang He , Nirui Song , Zhanchao Zhou , Shuyuan He , Shuai Zhang , Huachuan Qiu , Anqi Li , Yong Dai , Lizhi Ma , Zhenzhong Lan

With the advancements of transformer-based architectures, we observe the rise of natural language preprocessing (NLPre) tools capable of solving preliminary NLP tasks (e.g. tokenisation, part-of-speech tagging, dependency parsing, or…

计算与语言 · 计算机科学 2024-03-28 Martyna Wiącek , Piotr Rybak , Łukasz Pszenny , Alina Wróblewska

Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with specialized subdomains…

人工智能 · 计算机科学 2026-01-16 Zixun Lan , Maochun Xu , Yifan Ren , Rui Wu , Jianghui Zhou , Xueyang Cheng , Jianan Ding Ding , Xinheng Wang , Mingmin Chi , Fei Ma

With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evaluations predominantly…

计算与语言 · 计算机科学 2026-02-16 Yangzhuo Li , Shengpeng Ji , Yifu Chen , Tianle Liang , Haorong Ying , Yule Wang , Junbo Li , Jun Fang , Zhou Zhao

Recent advances in mobile Graphical User Interface (GUI) agents highlight the growing need for comprehensive evaluation benchmarks. While new online benchmarks offer more realistic testing than offline ones, they tend to focus on the…

计算与语言 · 计算机科学 2026-01-30 Qinzhuo Wu , Zhizhuo Yang , Hanhao Li , Pengzhi Gao , Wei Liu , Jian Luan

Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are…

The explosion of high-performing conversational language models (LMs) has spurred a shift from classic natural language processing (NLP) benchmarks to expensive, time-consuming and noisy human evaluations - yet the relationship between…

We introduce \texttt{N-LTP}, an open-source neural language technology platform supporting six fundamental Chinese NLP tasks: {lexical analysis} (Chinese word segmentation, part-of-speech tagging, and named entity recognition), {syntactic…

计算与语言 · 计算机科学 2021-09-24 Wanxiang Che , Yunlong Feng , Libo Qin , Ting Liu

Multilingual NLP is often treated as a route to global inclusion, but linguistic coverage and cultural competence frequently diverge. This paper synthesizes over 50 papers spanning multilingual performance inequality, cross-lingual…

计算与语言 · 计算机科学 2026-05-05 Sina Bagheri Nezhad

In recent years, a series of Transformer-based models unlocked major improvements in general natural language understanding (NLU) tasks. Such a fast pace of research would not be possible without general NLU benchmarks, which allow for a…

计算与语言 · 计算机科学 2020-05-05 Piotr Rybak , Robert Mroczkowski , Janusz Tracz , Ireneusz Gawlik

Different from the traditional translation tasks, classical Chinese poetry translation requires both adequacy and fluency in translating culturally and historically significant content and linguistic poetic elegance. Large language models…

计算与语言 · 计算机科学 2024-12-31 Andong Chen , Lianzhang Lou , Kehai Chen , Xuefeng Bai , Yang Xiang , Muyun Yang , Tiejun Zhao , Min Zhang

We present Dolphin, a novel benchmark that addresses the need for a natural language generation (NLG) evaluation framework dedicated to the wide collection of Arabic languages and varieties. The proposed benchmark encompasses a broad range…

计算与语言 · 计算机科学 2023-10-25 El Moatez Billah Nagoudi , AbdelRahim Elmadany , Ahmed El-Shangiti , Muhammad Abdul-Mageed