English
Related papers

Related papers: KoBBQ: Korean Bias Benchmark for Question Answerin…

200 papers

Large Language Models (LLMs) often exhibit cultural biases due to training data dominated by high-resource languages like English and Chinese. This poses challenges for accurately representing and evaluating diverse cultural contexts,…

Computation and Language · Computer Science 2025-08-11 Zhong Ken Hew , Jia Xin Low , Sze Jue Yang , Chee Seng Chan

While language embeddings have been shown to have stereotyping biases, how these biases affect downstream question answering (QA) models remains unexplored. We present UNQOVER, a general framework to probe and quantify biases through…

Computation and Language · Computer Science 2020-10-13 Tao Li , Tushar Khot , Daniel Khashabi , Ashish Sabharwal , Vivek Srikumar

Existing benchmarks evaluating biases in large language models (LLMs) primarily rely on explicit cues, declaring protected attributes like religion, race, gender by name. However, real-world interactions often contain implicit biases,…

Computation and Language · Computer Science 2025-12-09 Aarushi Wagh , Saniya Srivastava

The pervasive influence of social biases in language data has sparked the need for benchmark datasets that capture and evaluate these biases in Large Language Models (LLMs). Existing efforts predominantly focus on English language and the…

Computation and Language · Computer Science 2024-04-04 Nihar Ranjan Sahoo , Pranamya Prashant Kulkarni , Narjis Asad , Arif Ahmad , Tanu Goyal , Aparna Garimella , Pushpak Bhattacharyya

Bias studies on multilingual models confirm the presence of gender-related stereotypes in masked models processing languages with high NLP resources. We expand on this line of research by introducing Filipino CrowS-Pairs and Filipino…

Computation and Language · Computer Science 2025-04-29 Lance Calvin Lim Gamboa , Mark Lee

Prior benchmarks for evaluating the domain-specific knowledge of large language models (LLMs) lack the scalability to handle complex academic tasks. To address this, we introduce \texttt{ScholarBench}, a benchmark centered on deep expert…

Computation and Language · Computer Science 2025-10-17 Dongwon Noh , Donghyeok Koh , Junghun Yuk , Gyuwan Kim , Jaeyong Lee , Kyungtae Lim , Cheoneum Park

Large language models (LLMs) are widely deployed for open-ended communication, yet most bias evaluations still rely on English, classification-style tasks. We introduce \corpusname, a new multilingual, debate-style benchmark designed to…

Computation and Language · Computer Science 2026-03-31 Muhammed Saeed , Muhammad Abdul-mageed , Shady Shehata

Large Language Models (LLMs) can exhibit latent biases towards specific nationalities even when explicit demographic markers are not present. In this work, we introduce a novel name-based benchmarking approach derived from the Bias…

Computation and Language · Computer Science 2025-07-24 Giulio Pelosio , Devesh Batra , Noémie Bovey , Robert Hankache , Cristovao Iglesias , Greig Cowan , Raad Khraishi

Robust, diverse, and challenging cultural knowledge benchmarks are essential for measuring our progress towards making LMs that are helpful across diverse cultures. We introduce CulturalBench: a set of 1,696 human-written and human-verified…

Cognitive distortion refers to negative thinking patterns that can lead to mental health issues like depression and anxiety in adolescents. Previous studies using natural language processing (NLP) have focused mainly on small-scale adult…

Computation and Language · Computer Science 2025-09-23 JunSeo Kim , HyeHyeon Kim

We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams natively written in Korean, covering nine disciplines and nine…

Computation and Language · Computer Science 2026-04-20 Nahyun Lee , Guijin Son , Hyunwoo Ko , Chanyoung Kim , JunYoung An , Kyubeen Han , Il-Youp Kwak

Recent advances in vision-language models (VLMs) have enabled accurate image-based geolocation, raising serious concerns about location privacy risks in everyday social media posts. However, current benchmarks remain coarse-grained,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Xiaonan Wang , Bo Shao , Hansaem Kim

Recent evaluations of Large language models (LLMs) audit social bias primarily through prompts that explicitly reference demographic attributes, overlooking whether models infer sensitive demographics from neutral questions. Such inference…

Computation and Language · Computer Science 2026-01-27 Srikant Panda , Hitesh Laxmichand Patel , Shahad Al-Khalifa , Amit Agarwal , Hend Al-Khalifa , Sharefah Al-Ghamdi

In efforts to keep up with the rapid progress and use of large language models, gender bias research is becoming more prevalent in NLP. Non-English bias research, however, is still in its infancy with most work focusing on English. In our…

Computation and Language · Computer Science 2023-06-19 Victor Steinborn , Antonis Maronikolakis , Hinrich Schütze

We introduce the Korean Canonical Legal Benchmark (KCL), a benchmark designed to assess language models' legal reasoning capabilities independently of domain-specific knowledge. KCL provides question-level supporting precedents, enabling a…

Computation and Language · Computer Science 2026-01-06 Hongseok Oh , Wonseok Hwang , Kyoung-Woon On

Question answering over knowledge bases (KBQA) has become a popular approach to help users extract information from knowledge bases. Although several systems exist, choosing one suitable for a particular application scenario is difficult.…

Computation and Language · Computer Science 2022-11-16 Khiem Vinh Tran , Hao Phu Phan , Khang Nguyen Duc Quach , Ngan Luu-Thuy Nguyen , Jun Jo , Thanh Tam Nguyen

Recent advancements in Large Language Models (LLMs) have positioned them as powerful tools for clinical decision-making, with rapidly expanding applications in healthcare. However, concerns about bias remain a significant challenge in the…

Artificial Intelligence · Computer Science 2024-10-23 Kenza Benkirane , Jackie Kay , Maria Perez-Ortiz

Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making contexts. While prior work has shown that LLMs exhibit cognitive biases behaviorally, whether these biases correspond to identifiable internal…

Artificial Intelligence · Computer Science 2026-04-03 Fan Huang , Songheng Zhang , Haewoon Kwak , Jisun An

Legal reasoning requires not only the application of legal rules but also an understanding of the context in which those rules operate. However, existing legal benchmarks primarily evaluate rule application under the assumption of fixed…

Computation and Language · Computer Science 2026-03-30 JiHyeok Jung , TaeYoung Yoon , HyunSouk Cho
‹ Prev 1 3 4 5 6 7 10 Next ›