Computers and Society · Computer Science
Reliable agent engineering should integrate machine-compatible organizational principles
R. Patrick Xian, Garry A. Gabison, Ahmed Alaa, Christoph Riedl +1
2025-12-09
Computers and Society · Computer Science
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
Miles Q. Li, Benjamin C. M. Fung, Boyang Li, Heba Ismail +1
2026-05-19
Software Engineering · Computer Science
Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
Jia Yi Goh, Shaun Khoo, Nyx Iskandar, Gabriel Chua +2
2025-07-15
Artificial Intelligence · Computer Science
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
Jinhu Qi, Muzhi Li, Jiahong Liu, Yuqin Shu +8
2026-05-26
General Finance · Quantitative Finance
Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk
Zichen Chen, Jiaao Chen, Jianda Chen, Misha Sra
2025-06-03
Artificial Intelligence · Computer Science
Do LLMs estimate uncertainty well in instruction-following?
Juyeon Heo, Miao Xiong, Christina Heinze-Deml, Jaya Narain
2025-03-31
Computers and Society · Computer Science
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
Claudia Biancotti, Carolina Camassa, Andrea Coletta, Oliver Giudice +1
2025-02-26
Multiagent Systems · Computer Science
Risk Analysis Techniques for Governed LLM-based Multi-Agent Systems
Alistair Reid, Simon O'Callaghan, Liam Carroll, Tiberio Caetano
2025-08-11
Computers and Society · Computer Science
Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI
José Antonio Siqueira de Cerqueira, Mamia Agbese, Rebekah Rousi, Nannan Xi +2
2025-05-19
Computation and Language · Computer Science
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
Wenbin Hu, Huihao Jing, Haochen Shi, Haoran Li +1
2025-09-29
Computation and Language · Computer Science
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
Adi Simhi, Jonathan Herzig, Martin Tutek, Itay Itzhak +2
2026-03-04
Computers and Society · Computer Science
Institutional AI: A Governance Framework for Distributional AGI Safety
Federico Pierucci, Marcello Galisai, Marcantonio Syrnikov Bracale, Matteo Prandi +5
2026-01-21
Artificial Intelligence · Computer Science
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
Tomek Korbak, Mikita Balesni, Buck Shlegeris, Geoffrey Irving
2025-04-08
Artificial Intelligence · Computer Science
LM Agents May Fail to Act on Their Own Risk Knowledge
Yuzhi Tang, Tianxiao Li, Elizabeth Li, Chris J. Maddison +2
2025-08-20
Machine Learning · Computer Science
Evaluation and Benchmarking of LLM Agents: A Survey
Mahmoud Mohammadi, Yipeng Li, Jane Lo, Wendy Yip
2025-07-30
Cryptography and Security · Computer Science
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
Ishaan Domkundwar, Mukunda N S, Ishaan Bhola, Riddhik Kochhar
2025-03-04
Cryptography and Security · Computer Science
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
Julia Bazinska, Max Mathys, Francesco Casucci, Mateo Rojas-Carulla +3
2026-02-25
Computation and Language · Computer Science
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Zhexin Zhang, Shiyao Cui, Yida Lu, Jingzhuo Zhou +3
2025-05-21
Computation and Language · Computer Science
Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
Vamshi Krishna Bonagiri, Ponnurangam Kumaragurum, Khanh Nguyen, Benjamin Plaut
2026-02-03
Artificial Intelligence · Computer Science
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
Akshat Naik, Patrick Quinn, Guillermo Bosch, Emma Gouné +3
2025-10-02