中文
相关论文

相关论文: Red Teaming AI Red Teaming

200 篇论文

Collaboration is key to STEM, where multidisciplinary team research can solve complex problems. However, inequality in STEM fields hinders their full potential, due to persistent psychological barriers in underrepresented students'…

计算机与社会 · 计算机科学 2024-02-02 Nia Nixon , Yiwen Lin , Lauren Snow

Red teaming assesses how large language models (LLMs) can produce content that violates norms, policies, and rules set during their safety training. However, most existing automated methods in the literature are not representative of the…

In this paper, we argue that competitive pressures could incentivize AI companies to underinvest in ensuring their systems are safe, secure, and have a positive social impact. Ensuring that AI systems are developed responsibly may therefore…

计算机与社会 · 计算机科学 2019-07-11 Amanda Askell , Miles Brundage , Gillian Hadfield

In high-stakes disaster scenarios, timely and informed decision-making is critical yet often challenged by uncertainty, dynamic environments, and limited resources. This paper presents a systematic review of Human-AI collaboration patterns…

人工智能 · 计算机科学 2025-09-16 Emmanuel Adjei Domfeh , Christopher L. Dancy

Efforts to ensure the safety of large language models (LLMs) include safety fine-tuning, evaluation, and red teaming. However, despite the widespread use of the Retrieval-Augmented Generation (RAG) framework, AI safety work focuses on…

计算与语言 · 计算机科学 2025-04-28 Bang An , Shiyue Zhang , Mark Dredze

As large language models (LLMs) grow in power and influence, ensuring their safety and preventing harmful output becomes critical. Automated red teaming serves as a tool to detect security vulnerabilities in LLMs without manual labor.…

人工智能 · 计算机科学 2025-06-03 Weiyang Guo , Zesheng Shi , Zhuo Li , Yequan Wang , Xuebo Liu , Wenya Wang , Fangming Liu , Min Zhang , Jing Li

AI alignment research aims to develop techniques to ensure that AI systems do not cause harm. However, every alignment technique has failure modes, which are conditions in which there is a non-negligible chance that the technique fails to…

人工智能 · 计算机科学 2025-10-14 Leonard Dung , Florian Mai

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overlooked in traditional AI risk assessment frameworks. This…

密码学与安全 · 计算机科学 2024-10-21 Aviral Srivastava , Sourav Panda

AI is becoming increasingly integrated into everyday life, both in professional work environments and in leisure and entertainment contexts. This integration requires AI to move beyond acting as an assistant for informational or…

人机交互 · 计算机科学 2026-02-26 Christian Poelitz , Finale Doshi-Velez , Siân Lindley

The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures. However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more…

密码学与安全 · 计算机科学 2025-05-26 Weiyang Guo , Jing Li , Wenya Wang , YU LI , Daojing He , Jun Yu , Min Zhang

The main goal of design obfuscation schemes is to protect sensitive design details from untrusted parties in the VLSI supply chain, including but not limited to off-shore foundries and untrusted end users. In this work, we provide a…

密码学与安全 · 计算机科学 2025-08-20 Yuntao Liu , Abir Akib , Zelin Lu , Qian Xu , Ankur Srivastava , Gang Qu , David Kehlet , Nij Dorairaj

Artificial Intelligence's dual-use nature is revolutionizing the cybersecurity landscape, introducing new threats across four main categories: deepfakes and synthetic media, adversarial AI attacks, automated malware, and AI-powered social…

密码学与安全 · 计算机科学 2026-01-08 Sai Teja Erukude , Viswa Chaitanya Marella , Suhasnadh Reddy Veluru

This paper presents the vision, scientific contributions, and technical details of RedTWIZ: an adaptive and diverse multi-turn red teaming framework, to audit the robustness of Large Language Models (LLMs) in AI-assisted software…

As large language models grow in capability and agency, identifying vulnerabilities through red-teaming becomes vital for safe deployment. However, traditional prompt-engineering approaches may prove ineffective once red-teaming turns into…

人工智能 · 计算机科学 2026-02-10 Alexander Panfilov , Paul Kassianik , Maksym Andriushchenko , Jonas Geiping

This chapter explores the symbiotic relationship between Artificial Intelligence (AI) and trust in networked systems, focusing on how these two elements reinforce each other in strategic cybersecurity contexts. AI's capabilities in data…

人工智能 · 计算机科学 2024-11-21 Yunfei Ge , Quanyan Zhu

There is an increasing need to develop artificial intelligence systems that assist groups of humans working on coordinated tasks. These systems must recognize and understand the plans and relationships between actions for a team of humans…

人工智能 · 计算机科学 2023-08-01 Loren Rieffer-Champlin

We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-elicit attacks and overstate monitor performance. We identify three challenges with current…

密码学与安全 · 计算机科学 2026-05-12 Monika Jotautaitė , Maria Angelica Martinez , Ollie Matthews , Tyler Tracy

For AI technology to fulfill its full promises, we must have effective means to ensure Responsible AI behavior and curtail potential irresponsible use, e.g., in areas of privacy protection, human autonomy, robustness, and prevention of…

计算机与社会 · 计算机科学 2022-05-13 Wenjing Chu

With the increasing deployment of artificial intelligence (AI) technologies, the potential of humans working with AI agents has been growing at a great speed. Human-AI teaming is an important paradigm for studying various aspects when…

人机交互 · 计算机科学 2025-01-03 Lingyu Zhang , Zhengran Ji , Boyuan Chen

Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of…

计算机与社会 · 计算机科学 2025-02-17 Balint Gyevnar , Atoosa Kasirzadeh