English
Related papers

Related papers: Characterizing Selective Refusal Bias in Large Lan…

200 papers

Generative large language models (LLMs) have been shown to exhibit harmful biases and stereotypes. While safety fine-tuning typically takes place in English, if at all, these models are being used by speakers of many different languages.…

Computation and Language · Computer Science 2024-07-18 Vera Neplenbroek , Arianna Bisazza , Raquel Fernández

The recent advances in large language models (LLMs) have revolutionized industries such as finance, marketing, and customer service by enabling sophisticated natural language processing tasks. However, the broad adoption of LLMs brings…

Computers and Society · Computer Science 2025-02-19 Berk Yilmaz , Huthaifa I. Ashqar

Large Language Models (LLMs) offer the potential to automate hiring by matching job descriptions with candidate resumes, streamlining recruitment processes, and reducing operational costs. However, biases inherent in these models may lead…

Computation and Language · Computer Science 2025-03-26 Hayate Iso , Pouya Pezeshkpour , Nikita Bhutani , Estevam Hruschka

Large language models (LLMs) have transformed the way we access information. These models are often tuned to refuse to comply with requests that are considered harmful and to produce responses that better align with the preferences of those…

Computation and Language · Computer Science 2025-08-12 Hannah Cyberey , David Evans

Large Language Models (LLMs) have gained widespread adoption across various domains, including chatbots and auto-task completion agents. However, these models are susceptible to safety vulnerabilities such as jailbreaking, prompt injection,…

Cryptography and Security · Computer Science 2024-09-10 Divyanshu Kumar , Anurakt Kumar , Sahil Agarwal , Prashanth Harshangi

Most prior safety research of large language models (LLMs) has focused on enhancing the alignment of LLMs to better suit the safety requirements of humans. However, internalizing such safeguard features into larger models brought challenges…

Computation and Language · Computer Science 2025-01-24 Ohjoon Kwon , Donghyeon Jeon , Nayoung Choi , Gyu-Hwung Cho , Changbong Kim , Hyunwoo Lee , Inho Kang , Sun Kim , Taiwoo Park

Generated texts from large language models (LLMs) have been shown to exhibit a variety of harmful, human-like biases against various demographics. These findings motivate research efforts aiming to understand and measure such effects. This…

Computation and Language · Computer Science 2025-07-25 Yuen Chen , Vethavikashini Chithrra Raghuram , Justus Mattern , Rada Mihalcea , Zhijing Jin

As Large Language Models (LLMs) continue to evolve, they are increasingly being employed in numerous studies to simulate societies and execute diverse social tasks. However, LLMs are susceptible to societal biases due to their exposure to…

Computation and Language · Computer Science 2024-10-04 Angana Borah , Rada Mihalcea

Large language models (LLMs), known for their capability in understanding and following instructions, are vulnerable to adversarial attacks. Researchers have found that current commercial LLMs either fail to be "harmless" by presenting…

Cryptography and Security · Computer Science 2023-10-05 Bocheng Chen , Advait Paliwal , Qiben Yan

Safety evaluations of large language models (LLMs) typically report binary outcomes, i.e. attack success rate (ASR), refusal rate, or harmful versus safe classification, which hide how risk changes between prompt and response. We present a…

Computation and Language · Computer Science 2026-05-21 Mengya Hu , Qiong Wei , Sandeep Atluri

Safety-aligned large language models (LLMs) often generate refusal responses to harmless queries due to the over-refusal problem. However, existing methods for mitigating over-refusal cannot maintain a low refusal ratio for harmless queries…

Computation and Language · Computer Science 2026-04-21 Yupeng Qi , Ziyu Lyu , Lixin Cui , Lu Bai , Feng Xia

We present a general approach towards controllable societal biases in natural language generation (NLG). Building upon the idea of adversarial triggers, we develop a method to induce societal biases in generated text when input prompts…

Computation and Language · Computer Science 2020-10-08 Emily Sheng , Kai-Wei Chang , Premkumar Natarajan , Nanyun Peng

The use of language technologies in high-stake settings is increasing in recent years, mostly motivated by the success of Large Language Models (LLMs). However, despite the great performance of LLMs, they are are susceptible to ethical…

Artificial Intelligence · Computer Science 2025-06-16 Alejandro Peña , Julian Fierrez , Aythami Morales , Gonzalo Mancera , Miguel Lopez , Ruben Tolosana

We present SGuard-v1, a lightweight safety guardrail for Large Language Models (LLMs), which comprises two specialized models to detect harmful content and screen adversarial prompts in human-AI conversational settings. The first component,…

Computation and Language · Computer Science 2025-11-18 JoonHo Lee , HyeonMin Cho , Jaewoong Yun , Hyunjae Lee , JunKyu Lee , Juree Seok

Studying the robustness of Large Language Models (LLMs) to unsafe behaviors is an important topic of research today. Building safety classification models or guard models, which are fine-tuned models for input/output safety classification…

Computation and Language · Computer Science 2025-07-30 Sowmya Vajjala

Recent advances in the capacity of large language models to generate human-like text have resulted in their increased adoption in user-facing settings. In parallel, these improvements have prompted a heated discourse around the risks of…

Computation and Language · Computer Science 2023-02-23 Sachin Kumar , Vidhisha Balachandran , Lucille Njoo , Antonios Anastasopoulos , Yulia Tsvetkov

The use of Large Language Models (LLMs) has proven to be a tool that could help in the automatic detection of sexism. Previous studies have shown that these models contain biases that do not accurately reflect reality, especially for…

Computation and Language · Computer Science 2025-08-26 Judith Tavarez-Rodríguez , Fernando Sánchez-Vega , A. Pastor López-Monroy

Large language models (LLMs) increasingly employ guardrails to enforce ethical, legal, and application-specific constraints on their outputs. While effective at mitigating harmful responses, these guardrails introduce a new class of…

Cryptography and Security · Computer Science 2025-11-07 Hongwei Yao , Yun Xia , Shuo Shao , Haoran Shi , Tong Qiao , Cong Wang

With the growing deployment of large language models (LLMs) across various applications, assessing the influence of gender biases embedded in LLMs becomes crucial. The topic of gender bias within the realm of natural language processing…

Computation and Language · Computer Science 2024-03-04 Jinman Zhao , Yitian Ding , Chen Jia , Yining Wang , Zifan Qian

Recent research has highlighted that assigning specific personas to large language models (LLMs) can significantly increase harmful content generation. However, limited attention has been given to persona-driven toxicity in non-Western…

Computers and Society · Computer Science 2026-05-28 Geng Liu , Li Feng , Carlo Alberto Bono , Songbo Yang , Mengxiao Zhu , Francesco Pierri