中文
相关论文

相关论文: Red-Teaming for Generative AI: Silver Bullet or Se…

200 篇论文

The systemic risks posed by general-purpose AI models are a growing concern, yet the effectiveness of mitigations remains underexplored. Previous research has proposed frameworks for risk mitigation, but has left gaps in our understanding…

计算机与社会 · 计算机科学 2024-12-04 Risto Uuk , Annemieke Brouwer , Tim Schreier , Noemi Dreksler , Valeria Pulignano , Rishi Bommasani

Given rapid progress toward advanced AI and risks from frontier AI systems (advanced AI systems pushing the boundaries of the AI capabilities frontier), the creation and implementation of AI governance and regulatory schemes deserves…

Generative AI systems are increasingly used not only to produce content but also to retrieve data, invoke tools, and execute actions. This work examines the security and safety implications of that shift across content-level, model-level,…

密码学与安全 · 计算机科学 2026-05-19 Zelin Zhang , Qi Li , Jie Cao , Lingshuang Liu , Jianbing Ni

Generative AI systems may pose serious risks to individuals vulnerable to eating disorders. Existing safeguards tend to overlook subtle but clinically significant cues, leaving many risks unaddressed. To better understand the nature of…

人机交互 · 计算机科学 2025-12-05 Amy Winecoff , Kevin Klyman

Generative AI (genAI) is being used in education for different purposes. From the teachers' perspective, genAI can support activities such as learning design. However, there is a need to study the impact of genAI on the teachers' agency.…

人工智能 · 计算机科学 2024-07-10 Thomas B Frøsig , Margarida Romero

Explainable Artificial Intelligence (XAI) is a young but very promising field of research. Unfortunately, the progress in this field is currently slowed down by divergent and incompatible goals. We separate various threads tangled within…

人工智能 · 计算机科学 2024-07-30 Przemyslaw Biecek , Wojciech Samek

The rapid advancement of Vision-Language Models (VLMs) has brought their safety vulnerabilities into sharp focus. However, existing red teaming methods are fundamentally constrained by an inherent linear exploration paradigm, confining them…

机器学习 · 计算机科学 2026-03-25 Chunxiao Li , Lijun Li , Jing Shao

Generative AI models perturb the foundations of effective human communication. They present new challenges to contextual confidence, disrupting participants' ability to identify the authentic context of communication and their ability to…

人工智能 · 计算机科学 2024-01-26 Shrey Jain , Zoë Hitzig , Pamela Mishkin

Generative AI has made significant strides, yet concerns about the accuracy and reliability of its outputs continue to grow. Such inaccuracies can have serious consequences such as inaccurate decision-making, the spread of false…

数据库 · 计算机科学 2023-10-12 Nan Tang , Chenyu Yang , Ju Fan , Lei Cao , Yuyu Luo , Alon Halevy

Cyber threats continue to evolve in complexity, thereby traditional Cyber Threat Intelligence (CTI) methods struggle to keep pace. AI offers a potential solution, automating and enhancing various tasks, from data ingestion to resilience…

密码学与安全 · 计算机科学 2024-05-24 Lampis Alevizos , Martijn Dekker

Generative artificial intelligence (GenAI) is increasingly used to support a wide range of human tasks, yet empirical evidence on its effect on creativity remains scattered. Can GenAI generate ideas that are creative? To what extent can it…

人机交互 · 计算机科学 2025-05-26 Niklas Holzner , Sebastian Maier , Stefan Feuerriegel

Safety cases, structured arguments that a system is acceptably safe, are becoming central to the governance of AI systems. Yet, traditional safety-case practices from aviation or nuclear engineering rely on well-specified system boundaries,…

软件工程 · 计算机科学 2026-03-09 Sung Une Lee , Liming Zhu , Md Shamsujjoha , Liming Dong , Qinghua Lu , Jieshan Chen , Lionel Briand

The construction industry is a vital sector of the global economy, but it faces many productivity challenges in various processes, such as design, planning, procurement, inspection, and maintenance. Generative artificial intelligence (AI),…

The research investigates the crucial role of clear and intelligible terms of service in cultivating user trust and facilitating informed decision-making in the context of AI, in specific GenAI. It highlights the obstacles presented by…

计算机与社会 · 计算机科学 2024-06-21 Sundaraparipurnan Narayanan

Responsible AI (RAI) content work, such as annotation, moderation, or red teaming for AI safety, often exposes crowd workers to potentially harmful content. While prior work has underscored the importance of communicating well-being risk to…

人机交互 · 计算机科学 2026-04-02 Alice Qian , Ziqi Yang , Ryland Shaw , Jina Suh , Laura Dabbish , Hong Shen

Disinformation is among the top risks of generative artificial intelligence (AI) misuse. Global adoption of generative AI necessitates red-teaming evaluations (i.e., systematic adversarial probing) that are robust across diverse languages…

计算与语言 · 计算机科学 2025-09-24 Alejandro Cuevas , Saloni Dash , Bharat Kumar Nayak , Dan Vann , Madeleine I. G. Daepp

Generative AI is reshaping offensive cybersecurity by enabling autonomous red team agents that can plan, execute, and adapt during penetration tests. However, existing approaches face trade-offs between generality and specialization, and…

密码学与安全 · 计算机科学 2025-11-25 Strahinja Janjusevic , Anna Baron Garcia , Sohrob Kazerounian

Oversight and control, which we collectively call supervision, are often discussed as ways to ensure that AI systems are accountable, reliable, and able to fulfill governance and management requirements. However, the requirements for "human…

人工智能 · 计算机科学 2025-11-04 David Manheim , Aidan Homewood

The widespread use of generative AI systems is coupled with significant ethical and social challenges. As a result, policymakers, academic researchers, and social advocacy groups have all called for such systems to be audited. However,…

计算机与社会 · 计算机科学 2024-07-09 Jakob Mokander , Justin Curl , Mihir Kshirsagar

Ever since the launch of ChatGPT, Generative AI has been a focal point for concerns about AI's perceived existential risk. Once a niche topic in AI research and philosophy, AI safety and existential risk has now entered mainstream debate…

计算机与社会 · 计算机科学 2025-05-01 Lynette Webb
‹ 上一页 1 8 9 10 下一页 ›