中文
相关论文

相关论文: Dissecting Bias in LLMs: A Mechanistic Interpretab…

200 篇论文

Assessments of algorithmic bias in large language models (LLMs) are generally catered to uncovering systemic discrimination based on protected characteristics such as sex and ethnicity. However, there are over 180 documented cognitive…

人机交互 · 计算机科学 2023-08-30 Alaina N. Talboy , Elizabeth Fuller

Large Vision Language Models (LVLMs) have achieved remarkable progress in multimodal tasks, yet they also exhibit notable social biases. These biases often manifest as unintended associations between neutral concepts and sensitive human…

人工智能 · 计算机科学 2025-05-28 Zhengyang Ji , Yifan Jia , Shang Gao , Yutao Yue

Large language models (LLMs) have rapidly become indispensable tools for acquiring information and supporting human decision-making. However, ensuring that these models uphold fairness across varied contexts is critical to their safe and…

计算机与社会 · 计算机科学 2026-03-05 Xulang Zhang , Rui Mao , Erik Cambria

Large Language Models (LLMs) have demonstrated remarkable capabilities in executing tasks based on natural language queries. However, these models, trained on curated datasets, inherently embody biases ranging from racial to national and…

计算与语言 · 计算机科学 2024-07-29 Lynnette Hui Xian Ng , Iain Cruickshank , Roy Ka-Wei Lee

Large language models are becoming the go-to solution for the ever-growing number of tasks. However, with growing capacity, models are prone to rely on spurious correlations stemming from biases and stereotypes present in the training data.…

计算与语言 · 计算机科学 2024-05-30 Tomasz Limisiewicz , David Mareček , Tomáš Musil

Large Language Models (LLMs) have excelled at language understanding and generating human-level text. However, even with supervised training and human alignment, these LLMs are susceptible to adversarial attacks where malicious users can…

Large Language Models (LLMs) have shown remarkable capabilities in a multitude of Natural Language Processing (NLP) tasks. However, these models are still not immune to limitations such as social biases, especially gender bias. This work…

计算与语言 · 计算机科学 2024-10-15 Divij Bajaj , Yuanyuan Lei , Jonathan Tong , Ruihong Huang

Large Language Models (LLMs), such as GPT-4 and BERT, have rapidly gained traction in natural language processing (NLP) and are now integral to financial decision-making. However, their deployment introduces critical challenges,…

计算机与社会 · 计算机科学 2024-10-29 Hui Zhong , Songsheng Chen , Mian Liang

In traditional decision making processes, social biases of human decision makers can lead to unequal economic outcomes for underrepresented social groups, such as women, racial or ethnic minorities. Recently, the increasing popularity of…

综合经济学 · 经济学 2024-03-25 Jiafu An , Difang Huang , Chen Lin , Mingzhu Tai

Large language models (LLMs) are increasingly examined as both behavioral subjects and decision systems, yet it remains unclear whether observed cognitive biases reflect surface imitation or deeper probability shifts. Anchoring bias, a…

人工智能 · 计算机科学 2025-11-11 Felipe Valencia-Clavijo

The rapid growth of Large Language Models (LLMs) has put forward the study of biases as a crucial field. It is important to assess the influence of different types of biases embedded in LLMs to ensure fair use in sensitive fields. Although…

计算与语言 · 计算机科学 2024-12-16 Jayanta Sadhu , Maneesha Rani Saha , Rifat Shahriyar

The pervasive spread of misinformation and disinformation in social media underscores the critical importance of detecting media bias. While robust Large Language Models (LLMs) have emerged as foundational tools for bias prediction,…

计算机与社会 · 计算机科学 2024-12-11 Luyang Lin , Lingzhi Wang , Jinsong Guo , Kam-Fai Wong

As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In this paper, we argue that the predominant approach of simply…

计算机与社会 · 计算机科学 2026-01-13 Addison J. Wu , Ryan Liu , Xuechunzi Bai , Thomas L. Griffiths

Large Language Models (LLMs) are a transformational technology, fundamentally changing how people obtain information and interact with the world. As people become increasingly reliant on them for an enormous variety of tasks, a body of…

计算机与社会 · 计算机科学 2025-05-08 Nouar Aldahoul , Hazem Ibrahim , Matteo Varvello , Aaron Kaufman , Talal Rahwan , Yasir Zaki

We propose to measure political bias in LLMs by analyzing both the content and style of their generated content regarding political issues. Existing benchmarks and measures focus on gender and racial biases. However, political bias exists…

计算与语言 · 计算机科学 2024-03-29 Yejin Bang , Delong Chen , Nayeon Lee , Pascale Fung

Large language models (LLMs) have led to breakthroughs in language tasks, yet the internal mechanisms that enable their remarkable generalization and reasoning abilities remain opaque. This lack of transparency presents challenges such as…

计算与语言 · 计算机科学 2024-04-17 Haiyan Zhao , Fan Yang , Bo Shen , Himabindu Lakkaraju , Mengnan Du

Large Language Models (LLMs) have revolutionized natural language processing, yet concerns persist regarding their tendency to reflect or amplify social biases. This study introduces a novel evaluation framework to uncover gender biases in…

计算与语言 · 计算机科学 2026-03-10 Evan Chen , Run-Jun Zhan , Yan-Bai Lin , Hung-Hsuan Chen

Large Language models (LLMs), such as ChatGPT, have gained popularity in recent years with the advancement of Natural Language Processing (NLP), with use cases spanning many disciplines and daily lives as well. LLMs inherit explicit and…

计算与语言 · 计算机科学 2025-12-01 Fatima Kazi

With the rapid development of large language models (LLMs), they have significantly improved efficiency across a wide range of domains. However, recent studies have revealed that LLMs often exhibit gender bias, leading to serious social…

计算与语言 · 计算机科学 2025-06-17 Xiaoqing Cheng , Hongying Zan , Lulu Kong , Jinwang Song , Min Peng

Large language models (LLMs) have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque. Mechanistic interpretability (i.e., the systematic study of how neural networks…

计算与语言 · 计算机科学 2026-02-13 Usman Naseem