English
Related papers

Related papers: A Safe Harbor for AI Evaluation and Red Teaming

200 papers

Robust access to trustworthy information is a critical need for society with implications for knowledge production, public health education, and promoting informed citizenry in democratic societies. Generative AI technologies may enable new…

Information Retrieval · Computer Science 2024-07-17 Bhaskar Mitra , Henriette Cramer , Olya Gurevich

There is general agreement that some form of regulation is necessary both for AI creators to be incentivised to develop trustworthy systems, and for users to actually trust those systems. But there is much debate about what form these…

As artificial intelligence (AI) becomes deeply integrated into critical infrastructures and everyday life, ensuring its safe deployment is one of humanity's most urgent challenges. Current AI models prioritize task optimization over safety,…

Artificial Intelligence · Computer Science 2024-11-08 Joshua T. S. Hewson

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overlooked in traditional AI risk assessment frameworks. This…

Cryptography and Security · Computer Science 2024-10-21 Aviral Srivastava , Sourav Panda

The integration of Generative Artificial Intelligence (AI) into autonomous machines represents a major paradigm shift in how these systems operate and unlocks new solutions to problems once deemed intractable. Although generative AI agents…

Robotics · Computer Science 2024-10-22 Jason Jabbour , Vijay Janapa Reddi

Agentic Artificial Intelligence (AI) can autonomously pursue long-term goals, make decisions, and execute complex, multi-turn workflows. Unlike traditional generative AI, which responds reactively to prompts, agentic AI proactively…

Computers and Society · Computer Science 2025-02-18 Anirban Mukherjee , Hannah Hanwen Chang

Governance institutions must respond to societal risks, including those posed by generative AI. This study empirically examines how public trust in institutions and AI technologies, along with perceived risks, shape preferences for AI…

Computers and Society · Computer Science 2025-05-01 Justin B. Bullock , Janet V. T. Pauketat , Hsini Huang , Yi-Fan Wang , Jacy Reese Anthis

Large Generative AI (GAI) models have the unparalleled ability to generate text, images, audio, and other forms of media that are increasingly indistinguishable from human-generated content. As these models often train on publicly available…

Computers and Society · Computer Science 2024-06-25 Tanja Šarčević , Alicja Karlowicz , Rudolf Mayer , Ricardo Baeza-Yates , Andreas Rauber

A number of leading AI companies, including OpenAI, Google DeepMind, and Anthropic, have the stated goal of building artificial general intelligence (AGI) - AI systems that achieve or exceed human performance across a wide range of…

Computers and Society · Computer Science 2023-05-15 Jonas Schuett , Noemi Dreksler , Markus Anderljung , David McCaffary , Lennart Heim , Emma Bluemke , Ben Garfinkel

Applications of Generative AI (Gen AI) are expected to revolutionize a number of different areas, ranging from science & medicine to education. The potential for these seismic changes has triggered a lively debate about the potential risks…

Collaboration in multi-agent autonomous systems is critical to increase performance while ensuring safety. However, due to heterogeneity of their features in, e.g., perception qualities, some autonomous systems have to be considered more…

Multiagent Systems · Computer Science 2023-05-22 Selma Saidi

In today's society, where Artificial Intelligence (AI) has gained a vital role, concerns regarding user's trust have garnered significant attention. The use of AI systems in high-risk domains have often led users to either under-trust it,…

Human-Computer Interaction · Computer Science 2025-04-16 Siddharth Mehrotra , Ujwal Gadiraju , Eva Bittner , Folkert van Delden , Catholijn M. Jonker , Myrthe L. Tielman

Automated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a standardized evaluation framework to rigorously assess new…

The expanding application of Artificial Intelligence (AI) in scientific fields presents unprecedented opportunities for discovery and innovation. However, this growth is not without risks. AI models in science, if misused, can amplify risks…

Artificial Intelligence · Computer Science 2023-12-12 Jiyan He , Weitao Feng , Yaosen Min , Jingwei Yi , Kunsheng Tang , Shuai Li , Jie Zhang , Kejiang Chen , Wenbo Zhou , Xing Xie , Weiming Zhang , Nenghai Yu , Shuxin Zheng

Generative AI is increasingly positioned as a peer in collaborative learning, yet its effects on ethical deliberation remain unclear. We report a between-subjects experiment with university students (N=217) who discussed an…

Human-Computer Interaction · Computer Science 2026-03-24 Yueqiao Jin , Roberto Martinez-Maldonado , Wanruo Shi , Songjie Huang , Mingmin Zheng , Xinbin Han , Dragan Gasevic , Lixiang Yan

The rapid and wide-scale adoption of AI to generate human speech poses a range of significant ethical and safety risks to society that need to be addressed. For example, a growing number of speech generation incidents are associated with…

Computation and Language · Computer Science 2024-05-16 Wiebke Hutiri , Oresiti Papakyriakopoulos , Alice Xiang

Legislation and public sentiment throughout the world have promoted fairness metrics, explainability, and interpretability as prescriptions for the responsible development of ethical artificial intelligence systems. Despite the importance…

Artificial Intelligence · Computer Science 2022-03-08 Erick Galinkin

The evidence on the effects of generative AI (GenAI) on critical thinking is mixed, with studies suggesting both potential harms and benefits depending on its implementation. Some argue that AI-driven provocations, such as questions asking…

Human-Computer Interaction · Computer Science 2026-03-23 Thomas Şerban von Davier , Hao-Ping Lee , Jodi Forlizzi , Sauvik Das

Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify…

As artificial intelligence systems grow more capable and autonomous, frontier AI development poses potential systemic risks that could affect society at a massive scale. Current practices at many AI labs developing these systems lack…

Computers and Society · Computer Science 2025-06-03 Aidan Kierans , Kaley Rittichier , Utku Sonsayar , Avijit Ghosh