English
Related papers

Related papers: PBBQ: A Persian Bias Benchmark Dataset Curated wit…

200 papers

With the widespread adoption of Large Language Models (LLMs) across various applications, it is empirical to ensure their fairness across all user communities. However, most LLMs are trained and evaluated on Western centric data, with…

Computation and Language · Computer Science 2025-09-30 Abdullah Hashmat , Muhammad Arham Mirza , Agha Ali Raza

As large language models (LLMs) become increasingly embedded in our daily lives, evaluating their quality and reliability across diverse contexts has become essential. While comprehensive benchmarks exist for assessing LLM performance in…

Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine,…

Large language models predominantly reflect Western cultures, largely due to the dominance of English-centric training data. This imbalance presents a significant challenge, as LLMs are increasingly used across diverse contexts without…

Computation and Language · Computer Science 2025-07-21 Erfan Moosavi Monazzah , Vahid Rahimzadeh , Yadollah Yaghoobzadeh , Azadeh Shakery , Mohammad Taher Pilehvar

This paper presents a comprehensive evaluation framework for aligning Persian Large Language Models (LLMs) with critical ethical dimensions, including safety, fairness, and social norms. It addresses the gaps in existing LLM evaluation…

Research on evaluating and analyzing large language models (LLMs) has been extensive for resource-rich languages such as English, yet their performance in languages such as Persian has received considerably less attention. This paper…

This paper presents a comprehensive evaluation framework for assessing the cultural competence of large language models (LLMs) in Persian. Existing Persian cultural benchmarks rely predominantly on multiple-choice formats and…

Computation and Language · Computer Science 2026-03-17 Reihaneh Iranmanesh , Saeedeh Davoudi , Pasha Abrishamchian , Ophir Frieder , Nazli Goharian

Generative large language models (LLMs) have been shown to exhibit harmful biases and stereotypes. While safety fine-tuning typically takes place in English, if at all, these models are being used by speakers of many different languages.…

Computation and Language · Computer Science 2024-07-18 Vera Neplenbroek , Arianna Bisazza , Raquel Fernández

Evaluating social biases in language models (LMs) is crucial for ensuring fairness and minimizing the reinforcement of harmful stereotypes in AI systems. Existing benchmarks, such as the Bias Benchmark for Question Answering (BBQ),…

Computation and Language · Computer Science 2025-08-12 Aditya Tomar , Nihar Ranjan Sahoo , Pushpak Bhattacharyya

Medical consumer question answering (CQA) is crucial for empowering patients by providing personalized and reliable health information. Despite recent advances in large language models (LLMs) for medical QA, consumer-oriented and…

Computation and Language · Computer Science 2025-05-27 Naghmeh Jamali , Milad Mohammadi , Danial Baledi , Zahra Rezvani , Hesham Faili

Large Language Models (LLMs), trained on extensive datasets using advanced deep learning architectures, have demonstrated remarkable performance across a wide range of language tasks, becoming a cornerstone of modern AI technologies.…

While emerging Persian NLP benchmarks have expanded into pragmatics and politeness, they rarely distinguish between memorized cultural facts and the ability to reason about implicit social norms. We introduce DivanBench, a diagnostic…

Computation and Language · Computer Science 2026-02-20 Alireza Sakhaeirad , Ali Ma'manpoosh , Arshia Hemmat

Holistically measuring societal biases of large language models is crucial for detecting and reducing ethical risks in highly capable AI models. In this work, we present a Chinese Bias Benchmark dataset that consists of over 100K questions…

Computation and Language · Computer Science 2023-06-29 Yufei Huang , Deyi Xiong

Evaluating Large Language Models (LLMs) is challenging due to their generative nature, necessitating precise evaluation methodologies. Additionally, non-English LLM evaluation lags behind English, resulting in the absence or weakness of…

Large language models (LLMs) struggle to navigate culturally specific communication norms, limiting their effectiveness in global contexts. We focus on Persian taarof, a social norm in Iranian interactions, which is a sophisticated system…

Computation and Language · Computer Science 2025-09-03 Nikta Gohari Sadr , Sahar Heidariasl , Karine Megerdoomian , Laleh Seyyed-Kalantari , Ali Emami

In recent years, multilingual Large Language Models (LLMs) have become an inseparable part of daily life, making it crucial for them to master the rules of conversational language in order to communicate effectively with users. While…

Computation and Language · Computer Science 2026-01-30 Ghazal Kalhor , Behnam Bahrak

With the development of large language models (LLMs), social biases in these LLMs have become a pressing issue. Although there are various benchmarks for social biases across languages, the extent to which Japanese LLMs exhibit social…

Computation and Language · Computer Science 2025-06-16 Hitomi Yanaka , Namgi Han , Ryoma Kumon , Jie Lu , Masashi Takeshita , Ryo Sekizawa , Taisei Kato , Hiromi Arai

The Bias Benchmark for Question Answering (BBQ) is designed to evaluate social biases of language models (LMs), but it is not simple to adapt this benchmark to cultural contexts other than the US because social biases depend heavily on the…

Computation and Language · Computer Science 2024-01-26 Jiho Jin , Jiseon Kim , Nayeon Lee , Haneul Yoo , Alice Oh , Hwaran Lee

Despite the widespread use of the Persian language by millions globally, limited efforts have been made in natural language processing for this language. The use of large language models as effective tools in various natural language…

Computation and Language · Computer Science 2023-12-27 Mohammad Amin Abbasi , Arash Ghafouri , Mahdi Firouzmandi , Hassan Naderi , Behrouz Minaei Bidgoli

Large Language Models (LLMs) inherently reflect the vast data distributions they encounter during their pre-training phase. As this data is predominantly sourced from the web, there is a high chance it will be skewed towards high-resourced…

Computation and Language · Computer Science 2025-09-03 Fakhraddin Alwajih , Abdellah El Mekki , Hamdy Mubarak , Majd Hawasly , Abubakr Mohamed , Muhammad Abdul-Mageed
‹ Prev 1 2 3 10 Next ›