English
Related papers

Related papers: Building a Domain-specific Guardrail Model in Prod…

200 papers

There has been rapid development in generative AI tools across the education sector, which in turn is leading to increased adoption by teachers. However, this raises concerns regarding the safety and age-appropriateness of the AI-generated…

Computers and Society · Computer Science 2025-08-08 Hannah-Beth Clark , Laura Benton , Emma Searle , Margaux Dowland , Matthew Gregory , Will Gayne , John Roberts

Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-generation autonomous cars. In this context, safety is no longer about blocking harmful…

Artificial Intelligence · Computer Science 2026-05-20 Ravi Pandya , Madison Bland , Duy P. Nguyen , Changliu Liu , Jaime Fernández Fisac , Andrea Bajcsy

Reasoning-based language models have demonstrated strong performance across various domains, with the most notable gains seen in mathematical and coding tasks. Recent research has shown that reasoning also offers significant benefits for…

Artificial Intelligence · Computer Science 2025-05-27 Makesh Narsimhan Sreedhar , Traian Rebedea , Christopher Parisien

AI agents that interact with their environments through tools enable powerful applications, but in high-stakes business settings, unintended actions can cause unacceptable harm, such as privacy breaches and financial loss. Existing…

Software Engineering · Computer Science 2026-04-20 Yining Hong , Yining She , Eunsuk Kang , Christopher S. Timperley , Christian Kästner

Foundation models are increasingly deployed in socially sensitive domains such as education, mental health, and caregiving, where failures are often cumulative and context-dependent. Existing guardrail approaches -- ranging from…

Artificial Intelligence · Computer Science 2026-05-20 Rebecca Ramnauth , Drazen Brscic , Brian Scassellati

The AI era has ushered in Large Language Models (LLM) to the technological forefront, which has been much of the talk in 2023, and is likely to remain as such for many years to come. LLMs are the AI models that are the power house behind…

Cryptography and Security · Computer Science 2026-01-22 Anjanava Biswas , Wrick Talukdar

Generative AI models are capable of performing a wide variety of tasks that have traditionally required creativity and human understanding. During training, they learn patterns from existing data and can subsequently generate new content…

Agentic AI systems plan, use tools, maintain state, and produce multi-step trajectories with external effects. Those properties create a governance problem that differs materially from single-turn generative AI: important risks emerge dur-…

Artificial Intelligence · Computer Science 2026-04-08 Christopher Koch

AI safety is a rapidly growing area of research that seeks to prevent the harm and misuse of frontier AI technology, particularly with respect to generative AI (GenAI) tools that are capable of creating realistic and high-quality content…

Artificial Intelligence · Computer Science 2025-02-19 Pin-Yu Chen

A multi-modal guardrail must effectively filter image content based on user-defined policies, identifying material that may be hateful, reinforce harmful stereotypes, contain explicit material, or spread misinformation. Deploying such…

Machine Learning · Computer Science 2025-07-29 Cheng-Fu Yang , Thanh Tran , Christos Christodoulopoulos , Weitong Ruan , Rahul Gupta , Kai-Wei Chang

This paper explores the development of an ethical guardrail framework for AI systems, emphasizing the importance of customizable guardrails that align with diverse user values and underlying ethics. We address the challenges of AI ethics by…

Computers and Society · Computer Science 2024-11-25 Kristina Šekrst , Jeremy McHugh , Jonathan Rodriguez Cefalu

We introduce a lightweight yet highly effective safety guardrail framework for language models, demonstrating that small-scale language models can achieve, and even surpass, the performance of larger counterparts in content moderation…

Machine Learning · Computer Science 2025-07-14 Aleksei Ilin , Gor Matevosyan , Xueying Ma , Vladimir Eremin , Suhaa Dada , Muqun Li , Riyaaz Shaik , Haluk Noyan Tokgozoglu

The use of reward functions to structure AI learning and decision making is core to the current reinforcement learning paradigm; however, without careful design of reward functions, agents can learn to solve problems in ways that may be…

Artificial Intelligence · Computer Science 2025-01-22 Jonathan Keane , Sam Keyser , Jeremy Kedziora

Guardrail, an emerging mechanism designed to ensure that large language models (LLMs) align with human values by moderating harmful or toxic responses, requires a sociotechnical approach in their design. This paper addresses a critical…

Artificial Intelligence · Computer Science 2025-06-05 Jinwei Hu , Yi Dong , Xiaowei Huang

The advancement of autonomous systems -- from legged robots to self-driving vehicles and aircraft -- necessitates executing increasingly high-performance and dynamic motions without ever putting the system or its environment in harm's way.…

Robotics · Computer Science 2026-03-31 Andrew W. Singletary , Max H. Cohen , Tamas G. Molnar , Aaron D. Ames

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks…

Generative models are now capable of producing natural language text that is, in some cases, comparable in quality to the text produced by people. In the computing education context, these models are being used to generate code, code…

Human-Computer Interaction · Computer Science 2023-08-09 Cynthia Zastudil , Magdalena Rogalska , Christine Kapp , Jennifer Vaughn , Stephen MacNeil

Large language models (LLMs) have convincing performance in a variety of downstream tasks. However, these systems are prone to generating undesirable outputs such as harmful and biased text. In order to remedy such generations, the…

Computation and Language · Computer Science 2025-08-08 Manish Nagireddy , Inkit Padhi , Soumya Ghosh , Prasanna Sattigeri

To responsibly develop Generative AI (GenAI) products, it is critical to define the scope of acceptable inputs and outputs. What constitutes a "safe" response is an actively debated question. Academic work puts an outsized focus on…

Foundation Model (FM)-based agents are revolutionizing application development across various domains. However, their rapidly growing capabilities and autonomy have raised significant concerns about AI safety. Researchers are exploring…

Software Engineering · Computer Science 2025-01-28 Md Shamsujjoha , Qinghua Lu , Dehai Zhao , Liming Zhu
‹ Prev 1 2 3 10 Next ›