English
Related papers

Related papers: FarsEval-PKBETS: A new diverse benchmark for evalu…

200 papers

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing…

Pre-trained Language Models (PLMs) are integral to many modern natural language processing (NLP) systems. Although multilingual models cover a wide range of languages, they often grapple with challenges like high inference costs and a lack…

Computation and Language · Computer Science 2024-07-19 Murtadha Ahmed , Saghir Alfasly , Bo Wen , Jamaal Qasem , Mohammed Ahmed , Yunfeng Liu

The growing use of large language models (LLMs) has raised concerns regarding their safety. While many studies have focused on English, the safety of LLMs in Arabic, with its linguistic and cultural complexities, remains under-explored.…

Computation and Language · Computer Science 2025-02-11 Yasser Ashraf , Yuxia Wang , Bin Gu , Preslav Nakov , Timothy Baldwin

The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluation platforms for assessing Chinese LLMs, many of these…

Computation and Language · Computer Science 2024-03-20 Chuang Liu , Linhao Yu , Jiaxuan Li , Renren Jin , Yufei Huang , Ling Shi , Junhui Zhang , Xinmeng Ji , Tingting Cui , Tao Liu , Jinwang Song , Hongying Zan , Sun Li , Deyi Xiong

Large language models (LLMs) excel in high-resource languages but face notable challenges in low-resource languages like Mongolian. This paper addresses these challenges by categorizing capabilities into language abilities (syntax and…

Computation and Language · Computer Science 2024-11-15 Mengyuan Zhang , Ruihui Wang , Bo Xia , Yuan Sun , Xiaobing Zhao

Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow strictly from stated premises. Many existing logical-reasoning benchmarks are generated by…

The reliance on translated or adapted datasets from English or multilingual resources introduces challenges regarding linguistic and cultural suitability. This study addresses the need for robust and culturally appropriate benchmarks by…

Large language models (LLMs) are becoming increasingly proficient in processing and generating multilingual texts, which allows them to address real-world problems more effectively. However, language understanding is a far more complex…

Computation and Language · Computer Science 2025-03-04 Sławomir Dadas , Małgorzata Grębowiec , Michał Perełkiewicz , Rafał Poświata

As millions of Muslims turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question arises: Can these AI systems reliably reason about Islamic law? We introduce IslamicLegalBench, the first benchmark evaluating…

Computation and Language · Computer Science 2026-02-26 Ezieddin Elmahjub , Junaid Qadir , Abdullah Mushtaq , Rafay Naeem , Ibrahim Ghaznavi , Waleed Iqbal

Large language models (LLMs) have achieved impressive results in high-resource languages like English, yet their effectiveness in low-resource and morphologically rich languages remains underexplored. In this paper, we present a…

Computation and Language · Computer Science 2026-02-13 Chengxuan Xia , Qianye Wu , Hongbin Guan , Sixuan Tian , Yilun Hao , Xiaoyu Wu

The rapid development of Large Language Models (LLMs) in vertical domains, including intellectual property (IP), lacks a specific evaluation benchmark for assessing their understanding, application, and reasoning abilities. To fill this…

Computation and Language · Computer Science 2024-06-19 Qiyao Wang , Jianguo Huang , Shule Lu , Yuan Lin , Kan Xu , Liang Yang , Hongfei Lin

Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more…

Large language models (LLMs) offer significant potential in enhancing psychiatric practice, from improving diagnostic accuracy to streamlining clinical documentation and therapeutic support. However, existing evaluation resources heavily…

Computation and Language · Computer Science 2025-11-25 Aya E. Fouda , Abdelrahamn A. Hassan , Radwa J. Hanafy , Mohammed E. Fouda

Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advances. While advances reported for English using PLMs are…

Computation and Language · Computer Science 2021-04-12 Amit Seker , Elron Bandel , Dan Bareket , Idan Brusilovsky , Refael Shaked Greenfeld , Reut Tsarfaty

Evaluating progress in large language models (LLMs) is often constrained by the challenge of verifying responses, limiting assessments to tasks like mathematics, programming, and short-form question-answering. However, many real-world…

Computation and Language · Computer Science 2026-05-19 Zhilin Wang , Jaehun Jung , Ximing Lu , Shizhe Diao , Ellie Evans , Jiaqi Zeng , Pavlo Molchanov , Yejin Choi , Jan Kautz , Yi Dong

Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the Open LLM Leaderboard aim to quantify these differences with several large benchmarks (sets of test items to which an LLM can respond either…

Computation and Language · Computer Science 2025-02-21 Alex Kipnis , Konstantinos Voudouris , Luca M. Schulze Buschoff , Eric Schulz

Large language models have been widely evaluated on tasks such as comprehension, summarization, code generation, etc. However, their performance on graduate-level, culturally grounded questions in the Indian context remains largely…

Computation and Language · Computer Science 2025-10-09 Ayush Maheshwari , Kaushal Sharma , Vivek Patel , Aditya Maheshwari

We introduce FaBERT, a Persian BERT-base model pre-trained on the HmBlogs corpus, encompassing both informal and formal Persian texts. FaBERT is designed to excel in traditional Natural Language Understanding (NLU) tasks, addressing the…

Computation and Language · Computer Science 2024-02-12 Mostafa Masumi , Seyed Soroush Majd , Mehrnoush Shamsfard , Hamid Beigy

Sentiment analysis aims to extract people's emotions and opinion from their comments on the web. It widely used in businesses to detect sentiment in social data, gauge brand reputation, and understand customers. Most of articles in this…

Computation and Language · Computer Science 2022-12-13 Ali Nazarizadeh , Touraj Banirostam , Minoo Sayyadpour

The biomedical domain has sparked a significant interest in the field of Natural Language Processing (NLP), which has seen substantial advancements with pre-trained language models (PLMs). However, comparing these models has proven…

‹ Prev 1 4 5 6 7 8 10 Next ›