English
Related papers

Related papers: StepShield: When, Not Whether to Intervene on Rogu…

200 papers

While trajectory prediction plays a critical role in enabling safe and effective path-planning in automated vehicles, standardized practices for evaluating such models remain underdeveloped. Recent efforts have aimed to unify dataset…

Machine Learning · Computer Science 2025-09-19 Julian F. Schumann , Anna Mészáros , Jens Kober , Arkady Zgonnikov

Existing AI agent safety benchmarks focus on generic criminal harm (cybercrime, harassment, weapon synthesis), leaving a systematic blind spot for a distinct and commercially consequential threat category: agents harming their own…

Cryptography and Security · Computer Science 2026-04-22 Dongcheng Zhang , Yiqing Jiang

The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehensive industry-standard benchmark for assessing AI-product…

Computers and Society · Computer Science 2025-04-22 Shaona Ghosh , Heather Frase , Adina Williams , Sarah Luger , Paul Röttger , Fazl Barez , Sean McGregor , Kenneth Fricklas , Mala Kumar , Quentin Feuillade--Montixi , Kurt Bollacker , Felix Friedrich , Ryan Tsang , Bertie Vidgen , Alicia Parrish , Chris Knotz , Eleonora Presani , Jonathan Bennion , Marisa Ferrara Boston , Mike Kuniavsky , Wiebke Hutiri , James Ezick , Malek Ben Salem , Rajat Sahay , Sujata Goswami , Usman Gohar , Ben Huang , Supheakmungkol Sarin , Elie Alhajjar , Canyu Chen , Roman Eng , Kashyap Ramanandula Manjusha , Virendra Mehta , Eileen Long , Murali Emani , Natan Vidra , Benjamin Rukundo , Abolfazl Shahbazi , Kongtao Chen , Rajat Ghosh , Vithursan Thangarasa , Pierre Peigné , Abhinav Singh , Max Bartolo , Satyapriya Krishna , Mubashara Akhtar , Rafael Gold , Cody Coleman , Luis Oala , Vassil Tashev , Joseph Marvin Imperial , Amy Russ , Sasidhar Kunapuli , Nicolas Miailhe , Julien Delaunay , Bhaktipriya Radharapu , Rajat Shinde , Tuesday , Debojyoti Dutta , Declan Grabb , Ananya Gangavarapu , Saurav Sahay , Agasthya Gangavarapu , Patrick Schramowski , Stephen Singam , Tom David , Xudong Han , Priyanka Mary Mammen , Tarunima Prabhakar , Venelin Kovatchev , Rebecca Weiss , Ahmed Ahmed , Kelvin N. Manyeki , Sandeep Madireddy , Foutse Khomh , Fedor Zhdanov , Joachim Baumann , Nina Vasan , Xianjun Yang , Carlos Mougn , Jibin Rajan Varghese , Hussain Chinoy , Seshakrishna Jitendar , Manil Maskey , Claire V. Hardgrove , Tianhao Li , Aakash Gupta , Emil Joswin , Yifan Mai , Shachi H Kumar , Cigdem Patlak , Kevin Lu , Vincent Alessi , Sree Bhargavi Balija , Chenhe Gu , Robert Sullivan , James Gealy , Matt Lavrisa , James Goel , Peter Mattson , Percy Liang , Joaquin Vanschoren

As autonomous agents become more powerful and widely used, it is becoming increasingly important to ensure they behave safely and stay aligned with system goals, especially in multi-agent settings. Current systems often rely on agents…

Multiagent Systems · Computer Science 2025-04-08 Sagar Tamang , Dibya Jyoti Bora

Large Language Model (LLM) agents increasingly operate across domains such as robotics, virtual assistants, and web automation. However, their stochastic decision-making introduces safety risks that are difficult to anticipate during…

Artificial Intelligence · Computer Science 2026-03-30 Haoyu Wang , Christopher M. Poskitt , Jiali Wei , Jun Sun

Agent skills extend LLM agents with privileged third-party capabilities such as filesystem access, credentials, network calls, and shell execution. Existing safety work catches malicious prompts and risky runtime actions, but the skill…

Cryptography and Security · Computer Science 2026-05-13 Yuhao Wu , Tung-Ling Li , Hongliang Liu

LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation protocols primarily focus on task success, overlooking a critical aspect of agent…

Artificial Intelligence · Computer Science 2026-05-29 Minyang Hu , Bo Yang , Zhinuo Zhou , Jiachen Liang , Guo Jiahao , Yiyang Yin , Xiongwei Han

Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines can be set more fundamentally -- at the level of value, evidence, and source…

Artificial Intelligence · Computer Science 2026-04-14 Seulki Lee

Graphical User Interface (GUI) agents have the potential to assist users in interacting with complex software (e.g., PowerPoint, Photoshop). While prior research has primarily focused on automating user actions through clicks and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Saelyne Yang , Jaesang Yu , Yi-Hao Peng , Kevin Qinghong Lin , Jae Won Cho , Yale Song , Juho Kim

Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as…

Cryptography and Security · Computer Science 2026-04-28 Wenjie Xiao , Xuehai Tang , Biyu Zhou , Songlin Hu , Jizhong Han

AI agents are increasingly embedded in real software systems, where they execute multi-step workflows through multi-turn dialogue, tool invocations, and intermediate decisions. These long execution histories, called agentic traces, make…

Software Engineering · Computer Science 2026-05-12 Reshabh K Sharma , Shraddha Barke , Benjamin Zorn

As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts that penalize valid…

Computation and Language · Computer Science 2026-04-29 Xinming Tu , Tianze Wang , Yingzhou , Lu , Kexin Huang , Yuanhao Qu , Sara Mostafavi

Agents built on LLMs are increasingly deployed across diverse domains, automating complex decision-making and task execution. However, their autonomy introduces safety risks, including security vulnerabilities, legal violations, and…

Artificial Intelligence · Computer Science 2025-08-01 Haoyu Wang , Christopher M. Poskitt , Jun Sun

AI companions powered by large language models (LLMs) are increasingly integrated into users' daily lives, offering emotional support and companionship. While existing safety systems focus on overt harms, they rarely address early-stage…

Human-Computer Interaction · Computer Science 2025-10-21 Ziv Ben-Zion , Paul Raffelhüschen , Max Zettl , Antonia Lüönd , Achim Burrer , Philipp Homan , Tobias R Spiller

With the rapidly increasing capabilities and adoption of code agents for AI-assisted coding, safety concerns, such as generating or executing risky code, have become significant barriers to the real-world deployment of these agents. To…

Software Engineering · Computer Science 2024-11-13 Chengquan Guo , Xun Liu , Chulin Xie , Andy Zhou , Yi Zeng , Zinan Lin , Dawn Song , Bo Li

The safety of autonomous AI agents is increasingly recognized as a critical open problem. As agents transition from passive text generators to active actors capable of executing shell commands, modifying files, calling APIs, and browsing…

Artificial Intelligence · Computer Science 2026-05-19 Ashwin Aravind

Cyberattacks have grown into a major risk for organizations, with common consequences being data theft, sabotage, and extortion. Since preventive measures do not suffice to repel attacks, timely detection of successful intruders is crucial…

Cryptography and Security · Computer Science 2023-12-21 Rafael Uetz , Marco Herzog , Louis Hackländer , Simon Schwarz , Martin Henze

Large language models are increasingly deployed as *deep agents* that plan, maintain persistent state, and invoke external tools, shifting safety failures from unsafe text to unsafe *trajectories*. We introduce **AgentFence**, an…

Cryptography and Security · Computer Science 2026-02-10 Sai Puppala , Ismail Hossain , Md Jahangir Alam , Yoonpyo Lee , Jay Yoo , Tanzim Ahad , Syed Bahauddin Alam , Sajedul Talukder

As large language models (LLMs) are increasingly deployed as agents, their integration into interactive environments and tool use introduce new safety challenges beyond those associated with the models themselves. However, the absence of…

Computation and Language · Computer Science 2025-05-21 Zhexin Zhang , Shiyao Cui , Yida Lu , Jingzhuo Zhou , Junxiao Yang , Hongning Wang , Minlie Huang

Artificial Intelligence (AI) agents can now orchestrate cyberattacks. This development is already increasing the speed and scale of cyber attacks, decreasing attack costs, and improving the operational autonomy of cyber capabilities. To…

Computers and Society · Computer Science 2026-05-22 Matt Mittelsteadt , Jam Kraprayoon , Robin Staes-Polet , Oskar Galeev , Jan Wehner , Christopher Covino , Shaun Ee