Cryptography and Security · Computer Science
The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models
Zihui Wu, Haichang Gao, Jianping He, Ping Wang
2024-12-25
Computation and Language · Computer Science
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
Alberto Purpura, Sahil Wadhwa, Jesse Zymet, Akshay Gupta +4
2025-03-06
Cryptography and Security · Computer Science
From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs
Alsharif Abuadbba, Chris Hicks, Kristen Moore, Vasilios Mavroudis +3
2025-06-17
Multiagent Systems · Computer Science
Effective Red-Teaming of Policy-Adherent Agents
Itay Nakash, George Kour, Koren Lazar, Matan Vetzler +2
2025-08-26
Cryptography and Security · Computer Science
Autonomous Adversary: Red-Teaming in the age of LLM
Mohammad Mamun, Mohamed Gaber, Scott Buffett, Sherif Saad
2026-05-08
Computation and Language · Computer Science
Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations
Tarun Raheja, Nilay Pochhi, F. D. C. M. Curie
2024-12-18
Artificial Intelligence · Computer Science
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
Jianshuo Dong, Sheng Guo, Hao Wang, Xun Chen +5
2026-05-29
Computation and Language · Computer Science
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
Lingbo Mo, Boshi Wang, Muhao Chen, Huan Sun
2024-04-03
Machine Learning · Computer Science
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
Shuo Chen, Zhen Han, Bailan He, Zifeng Ding +4
2024-12-17
Cryptography and Security · Computer Science
Autonomous LLM Agents & CTFs: A Second Look
Youness Bouchari, Matteo Boffa, Marco Mellia, Idilio Drago +2
2026-05-22
Cryptography and Security · Computer Science
Toward Securing AI Agents Like Operating Systems
Lukas Pirch, Micha Horlboge, Patrick Großmann, Syeda Mahnur Asif +3
2026-05-15
Artificial Intelligence · Computer Science
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
Sanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang +3
2026-02-18
Cryptography and Security · Computer Science
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
Boyang Zhang, Yicong Tan, Yun Shen, Ahmed Salem +3
2024-07-31
Cryptography and Security · Computer Science
Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
Chenghao Du, Quanfeng Huang, Tingxuan Tang, Zihao Wang +2
2025-11-07
Cryptography and Security · Computer Science
Evaluation of Prompt Injection Defenses in Large Language Models
Priyal Deep, Shane Emmons, Amy Fox, Kyle Bacon +3
2026-05-14
Cryptography and Security · Computer Science
Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms
Tarek Gasmi, Ramzi Guesmi, Ines Belhadj, Jihene Bennaceur
2025-07-10
Computation and Language · Computer Science
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
Apurv Verma, Satyapriya Krishna, Sebastian Gehrmann, Madhavan Seshadri +6
2025-12-30
Multiagent Systems · Computer Science
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
Ishan Kavathekar, Hemang Jain, Ameya Rathod, Ponnurangam Kumaraguru +1
2025-11-10
Artificial Intelligence · Computer Science
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
Andy Zou, Maxwell Lin, Eliot Jones, Micha Nowak +13
2025-07-29
Cryptography and Security · Computer Science
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
Mateusz Dziemian, Maxwell Lin, Xiaohan Fu, Micha Nowak +27
2026-03-18
Computation and Language · Computer Science
More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
Weimin Xiong, Ke Wang, Yifan Song, Hanchao Liu +3
2025-06-30
Cryptography and Security · Computer Science
A Red Teaming Roadmap Towards System-Level Safety
Zifan Wang, Christina Q. Knight, Jeremy Kritz, Willow E. Primack +1
2025-06-10