中文
相关论文

相关论文: Lessons From Red Teaming 100 Generative AI Product…

200 篇论文

Recent studies have discovered that large language models (LLM) may be ``fooled'' to output private information, including training data, system prompts, and personally identifiable information, under carefully crafted adversarial prompts.…

密码学与安全 · 计算机科学 2025-08-11 Yuzhou Nie , Zhun Wang , Ye Yu , Xian Wu , Xuandong Zhao , Wenbo Guo , Dawn Song

As artificial intelligence (AI) systems become increasingly adopted across sectors, the need for robust, proactive security strategies is paramount. Traditional defensive measures often fall short against the unique and evolving threats…

密码学与安全 · 计算机科学 2025-05-13 Josh Harguess , Chris M. Ward

As machine learning (ML) and artificial intelligence (AI) technologies become more widespread, concerns about their environmental impact are increasing due to the resource-intensive nature of training and inference processes. Green AI…

软件工程 · 计算机科学 2025-01-06 Vincenzo De Martino , Silverio Martínez-Fernández , Fabio Palomba

As AI becomes an integral part of our lives, the development of explainable AI, embodied in the decision-making process of an AI or robotic agent, becomes imperative. For a robotic teammate, the ability to generate explanations to justify…

人工智能 · 计算机科学 2020-09-01 Mehrdad Zakershahrak , Ze Gong , Nikhillesh Sadassivam , Yu Zhang

There are many unknowns regarding the characteristics and dynamics of human-AI teams, including a lack of understanding of how certain human-human teaming concepts may or may not apply to human-AI teams and how this composition affects team…

人机交互 · 计算机科学 2021-05-25 Nathan J. McNeese , Beau G. Schelble , Lorenzo Barberis Canonico , Mustafa Demir

This chapter formulates seven lessons for preventing harm in artificial intelligence (AI) systems based on insights from the field of system safety for software-based automation in safety-critical domains. New applications of AI across…

系统与控制 · 电气工程与系统科学 2022-02-21 Roel I. J. Dobbe

Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that…

Collaboration is key to STEM, where multidisciplinary team research can solve complex problems. However, inequality in STEM fields hinders their full potential, due to persistent psychological barriers in underrepresented students'…

计算机与社会 · 计算机科学 2024-02-02 Nia Nixon , Yiwen Lin , Lauren Snow

Generative AI is increasingly important in software engineering, including safety engineering, where its use ensures that software does not cause harm to people. This also leads to high quality requirements for generative AI. Therefore, the…

软件工程 · 计算机科学 2024-04-25 Florian Geissler , Karsten Roscher , Mario Trapp

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

密码学与安全 · 计算机科学 2025-07-30 Petr Spelda , Vit Stritecky

Federated Learning (FL) is increasingly being adopted in military collaborations to develop Large Language Models (LLMs) while preserving data sovereignty. However, prompt injection attacks-malicious manipulations of input prompts-pose new…

机器学习 · 计算机科学 2026-05-05 Youngjoon Lee , Taehyun Park , Yunho Lee , Jinu Gong , Joonhyuk Kang

The integration of Generative Artificial Intelligence (AI) into autonomous machines represents a major paradigm shift in how these systems operate and unlocks new solutions to problems once deemed intractable. Although generative AI agents…

机器人学 · 计算机科学 2024-10-22 Jason Jabbour , Vijay Janapa Reddi

In an era where digital threats are increasingly sophisticated, the intersection of Artificial Intelligence and cybersecurity presents both promising defenses and potent dangers. This paper delves into the escalating threat posed by the…

密码学与安全 · 计算机科学 2024-08-26 Yusuf Usman , Aadesh Upadhyay , Prashnna Gyawali , Robin Chataut

As AI agents are increasingly adopted to collaborate on complex objectives, ensuring the security of autonomous multi-agent systems becomes crucial. We develop simulations of agents collaborating on shared objectives to study these security…

The use of artificial intelligence (AI) in research across all disciplines is becoming ubiquitous. However, this ubiquity is largely driven by hyperspecific AI models developed during scientific studies for accomplishing a well-defined,…

计算机与社会 · 计算机科学 2023-12-19 Rishab Jain , Aditya Jain

With recent advancements in multi-agent generative AI (Gen AI), technology organizations like Microsoft are adopting these complex tools, redefining AI agents as active collaborators in complex workflows rather than as passive tools. In…

人机交互 · 计算机科学 2025-10-09 Suchismita Naik , Austin L. Toombs , Amanda Snellinger , Scott Saponas , Amanda K. Hall

Reducing the number of failures in a production system is one of the most challenging problems in technology driven industries, such as, the online retail industry. To address this challenge, change management has emerged as a promising…

机器学习 · 计算机科学 2021-08-19 Binay Gupta , Anirban Chatterjee , Harika Matha , Kunal Banerjee , Lalitdutt Parsai , Vijay Agneeswaran

We present a comprehensive AI risk taxonomy derived from eight government policies from the European Union, United States, and China and 16 company policies worldwide, making a significant step towards establishing a unified language for…

计算机与社会 · 计算机科学 2024-06-27 Yi Zeng , Kevin Klyman , Andy Zhou , Yu Yang , Minzhou Pan , Ruoxi Jia , Dawn Song , Percy Liang , Bo Li

In the contemporary digital age, Quantum Computing and Artificial Intelligence (AI) convergence is reshaping the cyber landscape, introducing unprecedented opportunities and potential vulnerabilities.This research, conducted over five…

计算机与社会 · 计算机科学 2023-10-10 Petar Radanliev , David De Roure , Omar Santos

Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-generation autonomous cars. In this context, safety is no longer about blocking harmful…

人工智能 · 计算机科学 2026-05-20 Ravi Pandya , Madison Bland , Duy P. Nguyen , Changliu Liu , Jaime Fernández Fisac , Andrea Bajcsy