English
Related papers

Related papers: A Safe Harbor for AI Evaluation and Red Teaming

200 papers

In recent years, there has been an increased emphasis on understanding and mitigating adverse impacts of artificial intelligence (AI) technologies on society. Across academia, industry, and government bodies, a variety of endeavours are…

Computers and Society · Computer Science 2022-05-18 Ramya Srinivasan , Devi Parikh

Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet researchers, policymakers, and technology companies lack shared terminology for discussing AI risks.…

Various forms of implications of artificial intelligence that either exacerbate or decrease racial systemic injustice have been explored in this applied research endeavor. Taking each thematic area of identifying, analyzing, and debating an…

Computers and Society · Computer Science 2022-01-05 Alia Abbas

Generative AI is frequently portrayed as revolutionary or even apocalyptic, prompting calls for novel regulatory approaches. This essay argues that such views are misguided. Instead, generative AI should be understood as an evolutionary…

Computers and Society · Computer Science 2025-03-11 Gilad Abiri

In the next few years, applications of Generative AI are expected to revolutionize a number of different areas, ranging from science & medicine to education. The potential for these seismic changes has triggered a lively debate about…

In this paper, we study how different Reddit communities discuss generative AI in high school education, focusing on learning, academic integrity, AI detection, and emotional framing. Using 3,789 posts from five education-related…

Computers and Society · Computer Science 2026-03-27 Parth Gaba , Emiliano De Cristofaro

Many experts believe that AI systems will sooner or later pose uninsurable risks, including existential risks. This creates an extreme judgment-proof problem: few if any parties can be held accountable ex post in the event of such a…

Computers and Society · Computer Science 2025-07-15 Cristian Trout

Generative AI, with its tendency to "hallucinate" incorrect results, may pose a risk to knowledge work by introducing errors. On the other hand, it may also provide unprecedented opportunities for users, particularly non-experts, to learn…

Human-Computer Interaction · Computer Science 2024-12-20 Advait Sarkar , Xiaotong , Xu , Neil Toronto , Ian Drosos , Christian Poelitz

Generative AI is radically changing the creative arts, by fundamentally transforming the way we create and interact with cultural artefacts. While offering unprecedented opportunities for artistic expression and commercialisation, this…

Artificial Intelligence · Computer Science 2025-03-25 Jacopo de Berardinis , Lorenzo Porcaro , Albert Meroño-Peñuela , Angelo Cangelosi , Tess Buckley

Prominent AI experts have suggested that companies developing high-risk AI systems should be required to show that such systems are safe before they can be developed or deployed. The goal of this paper is to expand on this idea and explore…

Computers and Society · Computer Science 2024-06-25 Akash R. Wasil , Joshua Clymer , David Krueger , Emily Dardaman , Simeon Campos , Evan R. Murphy

Collaborative AI systems (CAISs) aim at working together with humans in a shared space to achieve a common goal. This critical setting yields hazardous circumstances that could harm human beings. Thus, building such systems with strong…

Human-Computer Interaction · Computer Science 2022-09-23 Jubril Gbolahan Adigun , Matteo Camilli , Michael Felderer , Andrea Giusti , Dominik T Matt , Anna Perini , Barbara Russo , Angelo Susi

Developing and implementing AI-based solutions help state and federal government agencies, research institutions, and commercial companies enhance decision-making processes, automate chain operations, and reduce the consumption of natural…

Artificial Intelligence · Computer Science 2021-12-02 Andrei Svetovidov , Abdul Rahman , Feras A. Batarseh

As AI systems grow more capable and autonomous, ensuring their safety and reliability requires not only model-level alignment but also strategic oversight of the humans and institutions involved in their development and deployment. Existing…

Artificial Intelligence · Computer Science 2026-02-10 Cheol Woo Kim , Davin Choo , Tzeh Yuan Neoh , Milind Tambe

Recent developments in AI safety research have called for red-teaming methods that effectively surface potential risks posed by generative AI models, with growing emphasis on how red-teamers' backgrounds and perspectives shape their…

Human-Computer Interaction · Computer Science 2026-05-12 Wesley Hanwen Deng , Mingxi Yan , Sunnie S. Y. Kim , Akshita Jha , Lauren Wilcox , Kenneth Holstein , Motahhare Eslami , Leon A. Gatys

Efforts to ensure the safety of large language models (LLMs) include safety fine-tuning, evaluation, and red teaming. However, despite the widespread use of the Retrieval-Augmented Generation (RAG) framework, AI safety work focuses on…

Computation and Language · Computer Science 2025-04-28 Bang An , Shiyue Zhang , Mark Dredze

The rapid growth of Large Language Models (LLMs) presents significant privacy, security, and ethical concerns. While much research has proposed methods for defending LLM systems against misuse by malicious actors, researchers have recently…

Computation and Language · Computer Science 2025-03-06 Alberto Purpura , Sahil Wadhwa , Jesse Zymet , Akshay Gupta , Andy Luo , Melissa Kazemi Rad , Swapnil Shinde , Mohammad Shahed Sorower

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed…

Artificial Intelligence · Computer Science 2025-05-26 Marie Davidsen Buhl , Jacob Pfau , Benjamin Hilton , Geoffrey Irving

As large language models (LLMs) achieve advanced persuasive capabilities, concerns about their potential risks have grown. The EU AI Act prohibits AI systems that use manipulative or deceptive techniques to undermine informed…

Computers and Society · Computer Science 2025-05-20 Haein Kong

Large generative AI models (GMs) like GPT and DALL-E are trained to generate content for general, wide-ranging purposes. GM content filters are generalized to filter out content which has a risk of harm in many cases, e.g., hate speech.…

Human-Computer Interaction · Computer Science 2023-06-07 Logan Stapleton , Jordan Taylor , Sarah Fox , Tongshuang Wu , Haiyi Zhu