中文
相关论文

相关论文: Building a Foundational Guardrail for General Agen…

200 篇论文

Autonomous agents powered by foundation models have seen widespread adoption across various real-world applications. However, they remain highly vulnerable to malicious instructions and attacks, which can result in severe consequences such…

机器学习 · 计算机科学 2025-12-01 Zhaorun Chen , Mintong Kang , Bo Li

With the rapid adoption of large language models (LLMs), conversational AI agents have become widely deployed across real-world applications. To enhance safety, these agents are often equipped with guardrails that moderate harmful content.…

密码学与安全 · 计算机科学 2026-04-14 Ziqing Yang , Yixin Wu , Rui Wen , Michael Backes , Yang Zhang

This paper introduces a multi-agent framework guided by Large Language Models (LLMs) to assist in the early stages of engineering design, a phase often characterized by vast parameter spaces and inherent uncertainty. Operating under a…

人工智能 · 计算机科学 2026-04-21 Varun Kumar , George Em Karniadakis

The adoption of Generative AI (GenAI) in applications inevitably comes with the expansion of the attack surface, combining new security threats along with the traditional ones. Consequently, numerous research and industrial initiatives aim…

密码学与安全 · 计算机科学 2025-08-22 Itay Hazan , Idan Habler , Ron Bitton , Itsik Mantin

As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that exploit extended user-agent-environment interactions to pursue malicious objectives…

密码学与安全 · 计算机科学 2026-05-06 Yuhui Wang , Tanqiu Jiang , Jiacheng Liang , Charles Fleming , Ting Wang

Large language model (LLM) alignment algorithms typically consist of post-training over preference pairs. While such algorithms are widely used to enable safety guardrails and align LLMs with general human preferences, we show that…

机器学习 · 计算机科学 2026-05-13 John T. Halloran

Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon actions where a single misstep, such as accessing files or entering credentials, can cause…

计算与语言 · 计算机科学 2026-03-04 Aradhye Agarwal , Gurdit Siyan , Yash Pandya , Joykirat Singh , Akshay Nambi , Ahmed Awadallah

We propose Guardian-FC, a novel two-layer framework for privacy preserving federated computing that unifies safety enforcement across diverse privacy preserving mechanisms, including cryptographic back-ends like fully homomorphic encryption…

密码学与安全 · 计算机科学 2025-06-26 Narasimha Raghavan Veeraragavan , Jan Franz Nygård

As large language models increasingly deployed into agentic systems, existing methods face critical gaps in observing, assessing, and mitigating deployment-specific risks. We present a comprehensive, observability-driven workflow: we…

计算与语言 · 计算机科学 2026-04-28 Ilham Wicaksono , Zekun Wu , Rahul Patel , Theo King , Adriano Koshiyama , Philip Treleaven

Deploying guardrails for custom policies remains challenging, as generic safety models fail to capture task-specific requirements, while prompting LLMs suffers from inconsistent boundary-case performance and high inference costs. Training…

计算与语言 · 计算机科学 2026-04-29 Arnon Mazza , Elad Levi

The integration of foundation models (FMs) into robotics has accelerated real-world deployment, while introducing new safety challenges arising from open-ended semantic reasoning and embodied physical action. These challenges require safety…

系统与控制 · 电气工程与系统科学 2026-02-05 Joonkyung Kim , Wenxi Chen , Davood Soleymanzadeh , Yi Ding , Xiangbo Gao , Zhengzhong Tu , Ruqi Zhang , Fan Fei , Sushant Veer , Yiwei Lyu , Minghui Zheng , Yan Gu

LLM-based multi-agent systems are increasingly deployed on long-horizon tasks, but a single decisive error is often accepted by downstream agents and cascades into trajectory-level failure. Existing work frames this as \emph{post-hoc…

计算与语言 · 计算机科学 2026-05-15 Boxuan Zhang , Jianing Zhu , Zeru Shi , Dongfang Liu , Ruixiang Tang

AI agents that interact with their environments through tools enable powerful applications, but in high-stakes business settings, unintended actions can cause unacceptable harm, such as privacy breaches and financial loss. Existing…

软件工程 · 计算机科学 2026-04-20 Yining Hong , Yining She , Eunsuk Kang , Christopher S. Timperley , Christian Kästner

Agentic AI systems - capable of goal interpretation, world modeling, planning, tool use, long-horizon operation, and autonomous coordination - introduce distinct control failures not addressed by existing safety frameworks. We identify six…

计算机与社会 · 计算机科学 2026-03-05 Subramanyam Sahoo

Unmanned Aerial Vehicles (UAVs) offer significant potential in dynamic, perception-intensive tasks such as search and rescue and environmental monitoring; however, their effectiveness is severely restricted by conventional pre-planned…

多智能体系统 · 计算机科学 2025-04-16 Yuhan Hu , Yirong Sun , Yanjun Chen , Xinghao Chen , Xiaoyu Shen , Wei Zhang

Foundation models have revolutionized artificial intelligence, yet their application in recommender systems remains limited by reasoning opacity and knowledge constraints. This paper introduces AgenticRAG, a novel framework that combines…

信息检索 · 计算机科学 2025-10-06 Bo Ma , Hang Li , ZeHua Hu , XiaoFan Gui , LuYao Liu , Simon Liu

In UAV dynamic decision, complex and variable hazardous factors pose severe challenges to the generalization capability of algorithms. Despite offering semantic understanding and scene generalization, Large Language Models (LLM) lack…

机器人学 · 计算机科学 2026-03-02 Wenzhe Zhao , Yang Zhao , Ganchao Liu , Zhiyu Jiang , Dandan Ma , Zihao Li , Xuelong Li

Modern Security Operations Centers struggle with alert fatigue, fragmented tooling, and limited cross-source event correlation. Challenges that current Security Information Event Management and Extended Detection and Response systems only…

密码学与安全 · 计算机科学 2026-04-08 Anes Abdennebi , Nadjia Kara , Laaziz Lahlou , Hakima Ould-Slimane

We present Multi-Agent gatekeeper, a framework that provides provable safety guarantees for leader-follower formation control in cluttered 3D environments. Existing methods face a trad-off: online planners and controllers lack formal safety…

机器人学 · 计算机科学 2025-11-26 Thomas Marshall Vielmetti , Devansh R Agrawal , Dimitra Panagou

As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line of defense against concrete deployment harms. In these settings, guardrail failures are no…