English
Related papers

Related papers: LionGuard: Building a Contextualized Moderation Cl…

200 papers

Large language models (LLMs) often fail to maintain safety in low-resource language varieties, such as code-mixed vernaculars and regional dialects. We introduce RabakBench, a multilingual safety benchmark and scalable pipeline localized to…

Computation and Language · Computer Science 2026-02-03 Gabriel Chua , Leanne Tan , Ziyu Ge , Roy Ka-Wei Lee

The widespread dissemination of hate speech, harassment, harmful and sexual content, and violence across websites and media platforms presents substantial challenges and provokes widespread concern among different sectors of society.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Nouar AlDahoul , Myles Joshua Toledo Tan , Harishwar Reddy Kasireddy , Yasir Zaki

Although region-specific large language models (LLMs) are increasingly developed, their safety remains underexplored, particularly in culturally diverse settings like Indonesia, where sensitivity to local norms is essential and highly…

Computation and Language · Computer Science 2025-06-04 Muhammad Falensi Azmi , Muhammad Dehan Al Kautsar , Alfan Farizki Wicaksono , Fajri Koto

Powered by remarkable advancements in Large Language Models (LLMs), Multimodal Large Language Models (MLLMs) demonstrate impressive capabilities in manifold tasks. However, the practical application scenarios of MLLMs are intricate,…

Computation and Language · Computer Science 2024-06-18 Tianle Gu , Zeyang Zhou , Kexin Huang , Dandan Liang , Yixu Wang , Haiquan Zhao , Yuanqi Yao , Xingge Qiao , Keqing Wang , Yujiu Yang , Yan Teng , Yu Qiao , Yingchun Wang

We introduce WildGuard -- an open, light-weight moderation tool for LLM safety that achieves three goals: (1) identifying malicious intent in user prompts, (2) detecting safety risks of model responses, and (3) determining model refusal…

Computation and Language · Computer Science 2024-12-11 Seungju Han , Kavel Rao , Allyson Ettinger , Liwei Jiang , Bill Yuchen Lin , Nathan Lambert , Yejin Choi , Nouha Dziri

Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on general safety and overlook cultural context. In this work, we introduce FanarGuard, a bilingual…

Computation and Language · Computer Science 2025-11-25 Masoomali Fatehkia , Enes Altinisik , Husrev Taha Sencar

As Large Language Models (LLMs) and generative AI become more widespread, the content safety risks associated with their use also increase. We find a notable deficiency in high-quality content safety datasets and benchmarks that…

Machine Learning · Computer Science 2024-09-12 Shaona Ghosh , Prasoon Varshney , Erick Galinkin , Christopher Parisien

Large language models (LLMs) now mediate many web-based mental-health, crisis, and other emotionally sensitive services, yet their psychosocial safety in these settings remains poorly understood and weakly evaluated. We present DialogGuard,…

Artificial Intelligence · Computer Science 2025-12-03 Han Luo , Guy Laban

As Large Language Models (LLMs) grow increasingly powerful, ensuring their safety and alignment with human values remains a critical challenge. Ideally, LLMs should provide informative responses while avoiding the disclosure of harmful or…

Computation and Language · Computer Science 2024-10-04 Lingrui Mei , Shenghua Liu , Yiwei Wang , Baolong Bi , Ruibin Yuan , Xueqi Cheng

Multi-modal Large Language Models (MLLMs) have achieved remarkable performance across a wide range of visual reasoning tasks, yet their vulnerability to safety risks remains a pressing concern. While prior research primarily focuses on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Ce Zhang , Jinxi He , Junyi He , Katia Sycara , Yaqi Xie

Vision-language models (VLMs) are essential for contextual understanding of both visual and textual information. However, their vulnerability to adversarially manipulated inputs presents significant risks, leading to compromised outputs and…

Machine Learning · Computer Science 2024-10-02 Xuefeng Du , Reshmi Ghosh , Robert Sim , Ahmed Salem , Vitor Carvalho , Emily Lawton , Yixuan Li , Jack W. Stokes

The trend towards large language models (LLMs) for guardrailing against undesired behaviors is increasing and has shown promise for censoring user inputs. However, increased latency, memory consumption, hosting expenses and non-structured…

Computation and Language · Computer Science 2025-04-30 James O' Neill , Santhosh Subramanian , Eric Lin , Vaikkunth Mugunthan

Despite the impressive capabilities of Large Language Models (LLMs) in various tasks, their vulnerability to unsafe prompts remains a critical issue. These prompts can lead LLMs to generate responses on illegal or sensitive topics, posing a…

Computation and Language · Computer Science 2024-07-10 Jinseok Kim , Jaewon Jung , Sangyeop Kim , Sohyung Park , Sungzoon Cho

The recent advancements in Large Language Models(LLMs) have had a significant impact on a wide range of fields, from general domains to specialized areas. However, these advancements have also significantly increased the potential for…

Cryptography and Security · Computer Science 2025-10-01 Taegyeong Lee , Jeonghwa Yoo , Hyoungseo Cho , Soo Yong Kim , Yunho Maeng

Current safety mechanisms for Large Language Models (LLMs) rely heavily on static, fine-tuned classifiers that suffer from adaptation rigidity, the inability to enforce new governance rules without expensive retraining. To address this, we…

Artificial Intelligence · Computer Science 2026-02-27 Umid Suleymanov , Rufiz Bayramov , Suad Gafarli , Seljan Musayeva , Taghi Mammadov , Aynur Akhundlu , Murat Kantarcioglu

With the growing deployment of large language models (LLMs) in real-world applications, establishing robust safety guardrails to moderate their inputs and outputs has become essential to ensure adherence to safety policies. Current…

Computation and Language · Computer Science 2026-03-04 Minseok Choi , Dongjin Kim , Seungbin Yang , Subin Kim , Youngjun Kwak , Juyoung Oh , Jaegul Choo , Jungmin Son

Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LLMs like ChaptGPT or Llama are often developed with content moderation guardrails that…

Human-Computer Interaction · Computer Science 2026-05-26 Jiwon Kim , Claire Wang , Taeung Yoon , Sabelle Huang , Koustuv Saha

Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerable to evolving harms, cross-lingual failures, and cultural misalignment. Moreover, most…

Computation and Language · Computer Science 2026-05-20 Tassallah Abdullahi , Macton Mgonzo , Mardiyyah Oduwole , Paul Okewunmi , Abraham Owodunni , Ritambhara Singh , Carsten Eickhoff

Social media platforms utilize Machine Learning (ML) and Artificial Intelligence (AI) powered recommendation algorithms to maximize user engagement, which can result in inadvertent exposure to harmful content. Current moderation efforts,…

Computation and Language · Computer Science 2025-05-30 Rajvardhan Oak , Muhammad Haroon , Claire Jo , Magdalena Wojcieszak , Anshuman Chhabra

Safeguarding large language models (LLMs) against unsafe or adversarial behavior is critical as they are increasingly deployed in conversational and agentic settings. Existing moderation tools often treat safety risks (e.g. toxicity, bias)…