English
Related papers

Related papers: A Safe Harbor for AI Evaluation and Red Teaming

200 papers

When AI systems are granted the agency to take impactful actions in the real world, there is an inherent risk that these systems behave in ways that are harmful. Typically, humans specify constraints on the AI system to prevent harmful…

Human-Computer Interaction · Computer Science 2022-11-09 Travis Mandel , Jahnu Best , Randall H. Tanaka , Hiram Temple , Chansen Haili , Kayla Schlectinger , Roy Szeto

Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of…

Computers and Society · Computer Science 2025-02-17 Balint Gyevnar , Atoosa Kasirzadeh

Generative Artificial Intelligence (GenAI) enables digital representatives to make decisions on behalf of team members in collaborative tasks, but faces challenges in accurately representing preferences. While supplying GenAI with detailed…

Computer Science and Game Theory · Computer Science 2025-12-16 Kexin Chen , Jianwei Huang , Yuan Luo

Deploying large language models (LMs) can pose hazards from harmful outputs such as toxic or false text. Prior work has introduced automated tools that elicit harmful outputs to identify these risks. While this is a valuable step toward…

Computation and Language · Computer Science 2023-10-12 Stephen Casper , Jason Lin , Joe Kwon , Gatlen Culp , Dylan Hadfield-Menell

Explanations for artificial intelligence (AI) systems are intended to support the people who are impacted by AI systems in high-stakes decision-making environments, such as doctors, patients, teachers, students, housing applicants, and many…

Human-Computer Interaction · Computer Science 2025-04-16 Gennie Mansi , Naveena Karusala , Mark Riedl

The widespread use of generative AI systems is coupled with significant ethical and social challenges. As a result, policymakers, academic researchers, and social advocacy groups have all called for such systems to be audited. However,…

Computers and Society · Computer Science 2024-07-09 Jakob Mokander , Justin Curl , Mihir Kshirsagar

As the deployment of artificial intelligence (AI) is changing many fields and industries, there are concerns about AI systems making decisions and recommendations without adequately considering various ethical aspects, such as…

Computers and Society · Computer Science 2023-10-02 Conrad Sanderson , Qinghua Lu , David Douglas , Xiwei Xu , Liming Zhu , Jon Whittle

Generative AI tools are increasingly entering academic peer review workflows, raising questions about fairness, accountability, and the legitimacy of evaluative judgment. While these systems promise efficiency gains amid growing reviewer…

Computers and Society · Computer Science 2026-03-24 Tatiana Chakravorti , Pranav Narayanan Venkit , Sourojit Ghosh , Sarah Rajtmajer

Artificial Intelligence (AI) has burrowed into our lives in various aspects; however, without appropriate testing, deployed AI systems are often being criticized to fail in critical and embarrassing cases. Existing testing approaches mainly…

Artificial Intelligence · Computer Science 2018-10-23 Siwei Fu , Anbang Xu , Xiaotong Liu , Huimin Zhou , Rama Akkiraju

Prominent AI companies are producing 'safety frameworks' as a type of voluntary self-governance. These statements purport to establish risk thresholds and safety procedures for the development and deployment of highly capable AI.…

Computers and Society · Computer Science 2025-10-14 Sam Coggins , Alexander K. Saeri , Katherine A. Daniell , Lorenn P. Ruster , Jessie Liu , Jenny L. Davis

The increased adoption of Artificial Intelligence (AI) presents an opportunity to solve many socio-economic and environmental challenges; however, this cannot happen without securing AI-enabled technologies. In recent years, most AI models…

Cryptography and Security · Computer Science 2021-02-10 Ayodeji Oseni , Nour Moustafa , Helge Janicke , Peng Liu , Zahir Tari , Athanasios Vasilakos

As Artificial Intelligence (AI) systems proliferate, the need for systematic, transparent, and actionable processes for evaluating them is growing. While many resources exist to support AI evaluation, they have several limitations. Few…

Computers and Society · Computer Science 2026-02-02 Rachel M. Kim , Blaine Kuehnert , Alice Lai , Kenneth Holstein , Hoda Heidari , Rayid Ghani

The tremendous rise of generative AI has reached every part of society - including the news environment. There are many concerns about the individual and societal impact of the increasing use of generative AI, including issues such as…

Computers and Society · Computer Science 2024-02-29 Kimon Kieslich , Nicholas Diakopoulos , Natali Helberger

Generative AI technologies have demonstrated significant potential across diverse applications. This study provides a comparative analysis of credit score modeling techniques, contrasting traditional approaches with those leveraging…

Machine Learning · Computer Science 2025-06-10 Nicola Lavecchia , Sid Fadanelli , Federico Ricciuti , Gennaro Aloe , Enrico Bagli , Pietro Giuffrida , Daniele Vergari

When users work with AI agents, they form conscious or subconscious expectations of them. Meeting user expectations is crucial for such agents to engage in successful interactions and teaming. However, users may form expectations of an…

Artificial Intelligence · Computer Science 2025-09-26 Akkamahadevi Hanni , Jonathan Montaño , Yu Zhang

Behind the scenes of maintaining the safety of technology products from harmful and illegal digital content lies unrecognized human labor. The recent rise in the use of generative AI technologies and the accelerating demands to meet…

Human-Computer Interaction · Computer Science 2024-11-05 Alice Qian Zhang , Judith Amores , Mary L. Gray , Mary Czerwinski , Jina Suh

AI projects often fail due to financial, technical, ethical, or user acceptance challenges -- failures frequently rooted in early-stage decisions. While HCI and Responsible AI (RAI) research emphasize this, practical approaches for…

Human-Computer Interaction · Computer Science 2025-06-24 Ji-Youn Jung , Devansh Saxena , Minjung Park , Jini Kim , Jodi Forlizzi , Kenneth Holstein , John Zimmerman

Red teaming is critical for identifying vulnerabilities and building trust in current LLMs. However, current automated methods for Large Language Models (LLMs) rely on brittle prompt templates or single-turn attacks, failing to capture the…

Machine Learning · Computer Science 2025-08-07 Roman Belaire , Arunesh Sinha , Pradeep Varakantham

The democratization of generative AI introduces new forms of human-AI interaction and raises urgent safety, ethical, and cybersecurity concerns. We develop a socio-technical explanation for how generative AI enables and scales cybercrime.…

Computers and Society · Computer Science 2025-12-04 Truong Jack Luu , Binny M. Samuel

Disinformation is among the top risks of generative artificial intelligence (AI) misuse. Global adoption of generative AI necessitates red-teaming evaluations (i.e., systematic adversarial probing) that are robust across diverse languages…

Computation and Language · Computer Science 2025-09-24 Alejandro Cuevas , Saloni Dash , Bharat Kumar Nayak , Dan Vann , Madeleine I. G. Daepp