English
Related papers

Related papers: SafeText: A Benchmark for Exploring Physical Safet…

200 papers

With the advent of text-to-image models and concerns about their misuse, developers are increasingly relying on image safety classifiers to moderate their generated unsafe images. Yet, the performance of current image safety classifiers…

Cryptography and Security · Computer Science 2025-09-12 Yiting Qu , Xinyue Shen , Yixin Wu , Michael Backes , Savvas Zannettou , Yang Zhang

Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research…

Conditional text generation has been a challenging task that is yet to see human-level performance from state-of-the-art models. In this work, we specifically focus on the Commongen benchmark, wherein the aim is to generate a plausible…

Computation and Language · Computer Science 2020-12-22 Yikang Li , Pulkit Goel , Varsha Kuppur Rajendra , Har Simrat Singh , Jonathan Francis , Kaixin Ma , Eric Nyberg , Alessandro Oltramari

Studying the robustness of Large Language Models (LLMs) to unsafe behaviors is an important topic of research today. Building safety classification models or guard models, which are fine-tuned models for input/output safety classification…

Computation and Language · Computer Science 2025-07-30 Sowmya Vajjala

With the rapid evolution of large language models (LLMs), new and hard-to-predict harmful capabilities are emerging. This requires developers to be able to identify risks through the evaluation of "dangerous capabilities" in order to…

Computation and Language · Computer Science 2023-09-06 Yuxia Wang , Haonan Li , Xudong Han , Preslav Nakov , Timothy Baldwin

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias,…

Artificial Intelligence · Computer Science 2025-06-23 Lina Berrayana , Sean Rooney , Luis Garcés-Erice , Ioana Giurgiu

Despite the impressive capabilities of Large Language Models (LLMs) in various tasks, their vulnerability to unsafe prompts remains a critical issue. These prompts can lead LLMs to generate responses on illegal or sensitive topics, posing a…

Computation and Language · Computer Science 2024-07-10 Jinseok Kim , Jaewon Jung , Sangyeop Kim , Sohyung Park , Sungzoon Cho

While large neural-based conversational models have become increasingly proficient dialogue agents, recent work has highlighted safety issues with these systems. For example, these systems can be goaded into generating toxic content, which…

Computation and Language · Computer Science 2023-10-24 Nicholas Meade , Spandana Gella , Devamanyu Hazarika , Prakhar Gupta , Di Jin , Siva Reddy , Yang Liu , Dilek Hakkani-Tür

Large Language Models (LLMs) are increasingly used for decision making in embodied agents, yet existing safety evaluations often rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these…

Artificial Intelligence · Computer Science 2025-05-27 Yejin Son , Minseo Kim , Sungwoong Kim , Seungju Han , Jian Kim , Dongju Jang , Youngjae Yu , Chanyoung Park

Large Vision-Language Models face growing safety challenges with multimodal inputs. This paper introduces the concept of Implicit Reasoning Safety, a vulnerability in LVLMs. Benign combined inputs trigger unsafe LVLM outputs due to flawed…

Artificial Intelligence · Computer Science 2025-08-13 Wei Cai , Jian Zhao , Yuchu Jiang , Tianle Zhang , Xuelong Li

The widespread adoption of Large Language Models (LLMs) raises concerns about the potential harmfulness of their responses. In this paper, we first investigate the harmfulness of responses from four general-purpose LLMs. Next, we propose…

Software Engineering · Computer Science 2026-05-11 Giordano d'Aloisio , David Williams , Giusy Annunziata , Zhiwei Fei , Antinisca Di Marco , Federica Sarro

Guidance on how to validate computational text-based measures of social constructs is fragmented. While researchers generally acknowledge the importance of validating text-based measures, they often lack a shared vocabulary and a unified…

Computation and Language · Computer Science 2024-06-11 Lukas Birkenmaier , Claudia Wagner , Clemens Lechner

Large Reasoning Models (LRMs) have significantly improved problem-solving through explicit Chain-of-Thought (CoT) reasoning. However, this capability creates a Safety-Helpfulness Paradox: the reasoning process itself can be misused to…

Artificial Intelligence · Computer Science 2026-01-27 Xin Gao , Shaohan Yu , Zerui Chen , Yueming Lyu , Weichen Yu , Guanghao Li , Jiyao Liu , Jianxiong Gao , Jian Liang , Ziwei Liu , Chenyang Si

Rapid deployment of vision-language models (VLMs) magnifies safety risks, yet most evaluations rely on artificial images. This study asks: How safe are current VLMs when confronted with meme images that ordinary users share? To investigate…

Computation and Language · Computer Science 2025-09-24 DongGeon Lee , Joonwon Jang , Jihae Jeong , Hwanjo Yu

Increasing use of large language models (LLMs) demand performant guardrails to ensure the safety of inputs and outputs of LLMs. When these safeguards are trained on imbalanced data, they can learn the societal biases. We present a…

Computation and Language · Computer Science 2024-10-23 Olivia Sturman , Aparna Joshi , Bhaktipriya Radharapu , Piyush Kumar , Renee Shelby

Genuine spatial reasoning relies on the capacity to construct and manipulate coherent internal spatial representations, often conceptualized as mental models, rather than merely processing surface linguistic associations. While large…

Artificial Intelligence · Computer Science 2026-03-04 Peiyao Jiang , Zequn Qin , Xi Li

Large pre-trained language models are often trained on large volumes of internet data, some of which may contain toxic or abusive language. Consequently, language models encode toxic information, which makes the real-world usage of these…

Computation and Language · Computer Science 2021-12-16 Andrew Wang , Mohit Sudhakar , Yangfeng Ji

This paper introduces RiskCards, a framework for structured assessment and documentation of risks associated with an application of language models. As with all language, text generated by language models can be harmful, or used to bring…

Computation and Language · Computer Science 2023-04-03 Leon Derczynski , Hannah Rose Kirk , Vidhisha Balachandran , Sachin Kumar , Yulia Tsvetkov , M. R. Leiser , Saif Mohammad

Some consider large-scale language models that can generate long and coherent pieces of text as dangerous, since they may be used in misinformation campaigns. Here we formulate large-scale language model output detection as a hypothesis…

Computation and Language · Computer Science 2020-02-11 Lav R. Varshney , Nitish Shirish Keskar , Richard Socher

The rapid growth in user generated content on social media has resulted in a significant rise in demand for automated content moderation. Various methods and frameworks have been proposed for the tasks of hate speech detection and toxic…

Computation and Language · Computer Science 2024-09-27 Elizaveta Korotkova , Isaac Chung