中文
相关论文

相关论文: AI Deception: A Survey of Examples, Risks, and Pot…

200 篇论文

Conversational Artificial Intelligence (AI) used in industry settings can be trained to closely mimic human behaviors, including lying and deception. However, lying is often a necessary part of negotiation. To address this, we develop a…

计算机与社会 · 计算机科学 2021-03-16 Tae Wan Kim , Tong , Lu , Kyusong Lee , Zhaoqi Cheng , Yanhan Tang , John Hooker

Under the slogan of trustworthy AI, much of contemporary AI research is focused on designing AI systems and usage practices that inspire human trust and, thus, enhance adoption of AI systems. However, a person affected by an AI system may…

计算机与社会 · 计算机科学 2025-05-16 Benjamin Paaßen , Suzana Alpsancar , Tobias Matzner , Ingrid Scharlau

This research critically navigates the intricate landscape of AI deception, concentrating on deceptive behaviours of Large Language Models (LLMs). My objective is to elucidate this issue, examine the discourse surrounding it, and…

计算与语言 · 计算机科学 2024-03-18 Linge Guo

The promise of AI is huge. AI systems have already achieved good enough performance to be in our streets and in our homes. However, they can be brittle and unfair. For society to reap the benefits of AI systems, society needs to be able to…

人工智能 · 计算机科学 2020-02-18 Jeannette M. Wing

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today's AI models. This…

人工智能 · 计算机科学 2024-11-26 Robert West , Roland Aydin

Successful deployment of artificial intelligence (AI) in various settings has led to numerous positive outcomes for individuals and society. However, AI systems have also been shown to harm parts of the population due to biased predictions.…

计算机与社会 · 计算机科学 2023-07-21 Ondrej Bohdal , Timothy Hospedales , Philip H. S. Torr , Fazl Barez

The increasing use of artificial intelligence (AI) systems in our daily life through various applications, services, and products explains the significance of trust/distrust in AI from a user perspective. AI-driven systems (as opposed to…

计算机与社会 · 计算机科学 2024-04-05 Saleh Afroogh , Ali Akbari , Evan Malone , Mohammadali Kargar , Hananeh Alambeigi

In this work, we survey skepticism regarding AI risk and show parallels with other types of scientific skepticism. We start by classifying different types of AI Risk skepticism and analyze their root causes. We conclude by suggesting some…

人工智能 · 计算机科学 2021-07-20 Roman V. Yampolskiy

As AI systems proliferate in society, the AI community is increasingly preoccupied with the concept of AI Safety, namely the prevention of failures due to accidents that arise from an unanticipated departure of a system's behavior from…

计算机与社会 · 计算机科学 2024-01-23 Inioluwa Deborah Raji , Roel Dobbe

As large language models (LLMs) are increasingly deployed as interactive agents, open-ended human-AI interactions can involve deceptive behaviors with serious real-world consequences, yet existing evaluations remain largely…

人工智能 · 计算机科学 2026-02-09 Yichen Wu , Qianqian Gao , Xudong Pan , Geng Hong , Min Yang

Artificial intelligence (AI) technologies (re-)shape modern life, driving innovation in a wide range of sectors. However, some AI systems have yielded unexpected or undesirable outcomes or have been used in questionable manners. As a…

In recent years the use of Artificial Intelligence (AI) has become increasingly prevalent in a growing number of fields. As AI systems are being adopted in more high-stakes areas such as medicine and finance, ensuring that they are…

人机交互 · 计算机科学 2023-11-03 Tobias M. Peters , Roel W. Visser

The capabilities of artificial intelligence systems have been advancing to a great extent, but these systems still struggle with failure modes, vulnerabilities, and biases. In this paper, we study the current state of the field, and present…

密码学与安全 · 计算机科学 2025-06-12 Xingli Fang , Jianwei Li , Varun Mulchandani , Jung-Eun Kim

The rapid development of Artificial Intelligence (AI) technology has enabled the deployment of various systems based on it. However, many current AI systems are found vulnerable to imperceptible attacks, biased against underrepresented…

人工智能 · 计算机科学 2022-05-27 Bo Li , Peng Qi , Bo Liu , Shuai Di , Jingen Liu , Jiquan Pei , Jinfeng Yi , Bowen Zhou

Persuasion is a fundamental aspect of communication, influencing decision-making across diverse contexts, from everyday conversations to high-stakes scenarios such as politics, marketing, and law. The rise of conversational AI systems has…

The discourse on risks from advanced AI systems ("AIs") typically focuses on misuse, accidents and loss of control, but the question of AIs' moral status could have negative impacts which are of comparable significance and could be realised…

计算机与社会 · 计算机科学 2024-08-12 Ines Fernandez , Nicoleta Kyosovska , Jay Luong , Gabriel Mukobi

Artificial Intelligence (AI) has the potential to significantly benefit or harm humanity. At present, a few for-profit companies largely control the development and use of this technology, and therefore determine its outcomes. In an effort…

计算机与社会 · 计算机科学 2022-11-14 Casey Clifton , Richard Blythman , Kartika Tulusan

Artificial Intelligence (AI) systems have historically been used as tools that execute narrowly defined tasks. Yet recent advances in AI have unlocked possibilities for a new class of models that genuinely collaborate with humans in complex…

As artificial intelligence (AI) systems become increasingly integral to organizational processes, they introduce new forms of fraud that are often subtle, systemic, and concealed within technical complexity. This paper introduces the…

计算机与社会 · 计算机科学 2025-08-20 Benjamin Zweers , Diptish Dey , Debarati Bhaumik

Instances of Artificial Intelligence (AI) systems failing to deliver consistent, satisfactory performance are legion. We investigate why AI failures occur. We address only a narrow subset of the broader field of AI Safety. We focus on AI…

计算机与社会 · 计算机科学 2020-08-11 Debarag Narayan Banerjee , Sasanka Sekhar Chanda