English
Related papers

Related papers: Basic Legibility Protocols Improve Trusted Monitor…

200 papers

Despite the impressive capabilities of Large Language Models (LLMs) in various tasks, their vulnerability to unsafe prompts remains a critical issue. These prompts can lead LLMs to generate responses on illegal or sensitive topics, posing a…

Computation and Language · Computer Science 2024-07-10 Jinseok Kim , Jaewon Jung , Sangyeop Kim , Sohyung Park , Sungzoon Cho

Vulnerability of Frontier language models to misuse and jailbreaks has prompted the development of safety measures like filters and alignment training in an effort to ensure safety through robustness to adversarially crafted prompts. We…

Cryptography and Security · Computer Science 2024-10-31 David Glukhov , Ziwen Han , Ilia Shumailov , Vardan Papyan , Nicolas Papernot

Large language models and other highly capable AI systems ease the burdens of deciding what to say or do, but this very ease can undermine the effectiveness of our actions in social contexts. We explain this apparent tension by introducing…

Computers and Society · Computer Science 2025-01-07 Zachary Wojtowicz , Simon DeDeo

The deployment of Large Language Models (LLMs) in robotic systems presents unique safety challenges, particularly in unpredictable environments. Although LLMs, leveraging zero-shot learning, enhance human-robot interaction and…

Robotics · Computer Science 2025-03-07 Ahmad Hafez , Alireza Naderi Akhormeh , Amr Hegazy , Amr Alanwar

Widespread adoption of autonomous cars will require greater confidence in their safety than is currently possible. Certified control is a new safety architecture whose goal is two-fold: to achieve a very high level of safety, and to provide…

As AI systems approach dangerous capability levels where inability safety cases become insufficient, we need alternative approaches to ensure safety. This paper presents a roadmap for constructing safety cases based on chain-of-thought…

Machine Learning · Computer Science 2025-10-23 Julian Schulz

Existing approaches to monitoring AI agents rely on supervised evaluation: human-written rules or LLM-based judges that check for known failure modes. However, novel misbehaviors may fall outside predefined categories entirely and LLM-based…

Artificial Intelligence · Computer Science 2026-04-14 Ziqian Zhong , Shashwat Saxena , Aditi Raghunathan

The rapid progress in Large Language Models (LLMs) could transform many fields, but their fast development creates significant challenges for oversight, ethical creation, and building user trust. This comprehensive review looks at key trust…

Computers and Society · Computer Science 2024-07-22 Md Meftahul Ferdaus , Mahdi Abdelguerfi , Elias Ioup , Kendall N. Niles , Ken Pathak , Steven Sloan

This work mainly addresses continuous-time multiagent consensus networks where an adverse attacker affects the convergence performances of said protocol. In particular, we develop a novel secure-by-design approach in which the presence of a…

Systems and Control · Electrical Eng. & Systems 2025-01-30 Marco Fabris , Daniel Zelazo

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

Cryptography and Security · Computer Science 2025-07-30 Petr Spelda , Vit Stritecky

Monitoring AIs at runtime can help us detect and stop harmful actions. In this paper, we study how to efficiently combine multiple runtime monitors into a single monitoring protocol. The protocol's objective is to maximize the probability…

Computers and Society · Computer Science 2025-10-22 Tim Tian Hua , James Baskerville , Henri Lemoine , Mia Hopman , Aryan Bhatt , Tyler Tracy

Backdoor mechanisms have traditionally been studied as security threats that compromise the integrity of machine learning models. However, the same mechanism -- the conditional activation of specific behaviors through input triggers -- can…

Cryptography and Security · Computer Science 2026-03-10 Yige Li , Wei Zhao , Zhe Li , Nay Myat Min , Hanxun Huang , Yunhan Zhao , Xingjun Ma , Yu-Gang Jiang , Jun Sun

As AI capabilities advance, we increasingly rely on powerful models to decompose complex tasks $\unicode{x2013}$ but what if the decomposer itself is malicious? Factored cognition protocols decompose complex tasks into simpler child tasks:…

Cryptography and Security · Computer Science 2025-12-18 Edward Lue Chee Lip , Anthony Channg , Diana Kim , Aaron Sandoval , Kevin Zhu

Trustworthy evaluations of dangerous capabilities are increasingly crucial for determining whether an AI system is safe to deploy. One empirically demonstrated threat is sandbagging - the strategic underperformance on evaluations by AI…

Cryptography and Security · Computer Science 2025-11-03 Chloe Li , Mary Phuong , Noah Y. Siegel

We introduce a novel framework for learning context-aware runtime monitors for AI-based control ensembles. Machine-learning (ML) controllers are increasingly deployed in (autonomous) cyber-physical systems because of their ability to solve…

Machine Learning · Computer Science 2026-04-03 Alejandro Luque-Cerpa , Mengyuan Wang , Emil Carlsson , Sanjit A. Seshia , Devdatt Dubhashi , Hazem Torfah

System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive operational context in agentic AI applications. These instructions may contain sensitive…

Cryptography and Security · Computer Science 2026-04-02 Anubhab Sahu , Diptisha Samanta , Reza Soosahabi

Large Language Models (LLMs) have transformed natural language processing (NLP) by enabling robust text generation and understanding. However, their deployment in sensitive domains like healthcare, finance, and legal services raises…

Artificial Intelligence · Computer Science 2024-12-09 Georgios Feretzakis , Vassilios S. Verykios

Learning-based methods provide a promising approach to solving highly non-linear control tasks that are often challenging for classical control methods. To ensure the satisfaction of a safety property, learning-based methods jointly learn a…

Machine Learning · Computer Science 2024-12-18 Emily Yu , Đorđe Žikelić , Thomas A. Henzinger

Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex tasks by engaging in extended reasoning before producing final answers. Beyond improving abilities, these detailed reasoning traces also create a new…

Computation and Language · Computer Science 2026-01-08 Shu Yang , Junchao Wu , Xilin Gong , Xuansheng Wu , Derek Wong , Ninghao Liu , Di Wang