English
Related papers

Related papers: A Safe Harbor for AI Evaluation and Red Teaming

200 papers

As intelligent systems are increasingly making decisions that directly affect society, perhaps the most important upcoming research direction in AI is to rethink the ethical implications of their actions. Means are needed to integrate…

Artificial Intelligence · Computer Science 2017-06-09 Virginia Dignum

Potential malicious misuse of civilian artificial intelligence (AI) poses serious threats to security on a national and international level. Besides defining autonomous systems from a technological viewpoint and explaining how AI…

Computers and Society · Computer Science 2024-03-25 Lukas Pöhler , Valentin Schrader , Alexander Ladwein , Florian von Keller

Generative AI systems are increasingly used not only to produce content but also to retrieve data, invoke tools, and execute actions. This work examines the security and safety implications of that shift across content-level, model-level,…

Cryptography and Security · Computer Science 2026-05-19 Zelin Zhang , Qi Li , Jie Cao , Lingshuang Liu , Jianbing Ni

Although AI has significant potential to transform society, there are serious concerns about its ability to behave and make decisions responsibly. Many ethical regulations, principles, and guidelines for responsible AI have been issued…

Artificial Intelligence · Computer Science 2022-09-21 Qinghua Lu , Liming Zhu , Xiwei Xu , Jon Whittle

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally…

Computers and Society · Computer Science 2025-06-05 Max Hellrigel-Holderbaum , Leonard Dung

Background: The rapid development and use of generative AI (GenAI) tools in academia presents complex and multifaceted ethical challenges for its users. Earlier research primarily focused on academic integrity concerns related to students'…

Computers and Society · Computer Science 2024-12-16 Sonja Bjelobaba , Lorna Waddington , Mike Perkins , Tomáš Foltýnek , Sabuj Bhattacharyya , Debora Weber-Wulff

Generative AI systems increasingly expose powerful reasoning and image refinement capabilities through user-facing chatbot interfaces. In this work, we show that the na\"ive exposure of such capabilities fundamentally undermines modern…

Cryptography and Security · Computer Science 2026-03-12 Sunpill Kim , Chanwoo Hwang , Minsu Kim , Jae Hong Seo

The rapid evolution of artificial intelligence (AI) systems, tools, and technologies has opened up novel, unprecedented opportunities for businesses to innovate, differentiate, and compete. However, growing concerns have emerged about the…

Human-Computer Interaction · Computer Science 2026-01-13 Nelly Elsayed

Enterprise AI systems, built on large language models, retrieval pipelines and autonomous agents, introduce a class of risks that traditional software quality assurance was never designed to address. These systems are probabilistic,…

Software Engineering · Computer Science 2026-05-25 Chitra Badagi , Divye Singh , Animesh Sen , Adinath Shirsath

Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant…

Over the past year, artificial intelligence (AI) companies have been increasingly adopting AI safety frameworks. These frameworks outline how companies intend to keep the potential risks associated with developing and deploying frontier AI…

Computers and Society · Computer Science 2024-09-16 Jide Alaga , Jonas Schuett , Markus Anderljung

AI systems for software development are rapidly gaining prominence, yet significant challenges remain in ensuring their safety. To address this, Amazon launched the Trusted AI track of the Amazon Nova AI Challenge, a global competition…

As Artificial Intelligence (AI) is increasingly promoted and used in qualitative research, it also raises profound methodological issues. This position paper critically interrogates the role of generative AI (genAI) in the context of…

Computers and Society · Computer Science 2025-11-12 Maria Couto Teixeira , Marisa Tschopp , Anna Jobin

Generative AI (e.g., Generative Adversarial Networks - GANs) has become increasingly popular in recent years. However, Generative AI introduces significant concerns regarding the protection of Intellectual Property Rights (IPR) (resp. model…

Although artificial intelligence (AI) shows growing promise for mental health care, current approaches to evaluating AI tools in this domain remain fragmented and poorly aligned with clinical practice, social context, and first-hand user…

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias,…

Artificial Intelligence · Computer Science 2025-06-23 Lina Berrayana , Sean Rooney , Luis Garcés-Erice , Ioana Giurgiu

Generative models are rapidly gaining popularity and being integrated into everyday applications, raising concerns over their safe use as various vulnerabilities are exposed. In light of this, the field of red teaming is undergoing…

Computation and Language · Computer Science 2024-11-27 Lizhi Lin , Honglin Mu , Zenan Zhai , Minghan Wang , Yuxia Wang , Renxi Wang , Junjie Gao , Yixuan Zhang , Wanxiang Che , Timothy Baldwin , Xudong Han , Haonan Li

The advent of advanced AI underscores the urgent need for comprehensive safety evaluations, necessitating collaboration across communities (i.e., AI, software engineering, and governance). However, divergent practices and terminologies…

Software Engineering · Computer Science 2024-05-17 Boming Xia , Qinghua Lu , Liming Zhu , Zhenchang Xing

With almost daily improvements in capabilities of artificial intelligence it is more important than ever to develop safety software for use by the AI research community. Building on our previous work on AI Containment Problem we propose a…

Artificial Intelligence · Computer Science 2017-07-27 James Babcock , Janos Kramar , Roman V. Yampolskiy

Generative AI models perturb the foundations of effective human communication. They present new challenges to contextual confidence, disrupting participants' ability to identify the authentic context of communication and their ability to…

Artificial Intelligence · Computer Science 2024-01-26 Shrey Jain , Zoë Hitzig , Pamela Mishkin
‹ Prev 1 8 9 10 Next ›