中文
相关论文

相关论文: AI Deception: A Survey of Examples, Risks, and Pot…

200 篇论文

This paper presents a preliminary draft of a framework around the use of anthropomorphic deception, defined here as misleading users towards humanlike affordances in the design of autonomous systems. The goal is to promote reflection among…

人机交互 · 计算机科学 2026-04-20 Franziska Babel , Shane Saunderson , Shalaleh Rismani

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI…

As Artificial Intelligence (AI) technologies proliferate, concern has centered around the long-term dangers of job loss or threats of machines causing harm to humans. All of this concern, however, detracts from the more pertinent and…

人工智能 · 计算机科学 2018-09-24 Kirsten Lloyd

Artificial Intelligence (AI) is an integral part of our daily technology use and will likely be a critical component of emerging technologies. However, negative user preconceptions may hinder adoption of AI-based decision making. Prior work…

人机交互 · 计算机科学 2021-11-18 Amama Mahmood , Gopika Ajaykumar , Chien-Ming Huang

Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose catastrophic risks. Although numerous risks…

计算机与社会 · 计算机科学 2023-10-11 Dan Hendrycks , Mantas Mazeika , Thomas Woodside

Modern AI systems are reaping the advantage of novel learning methods. With their increasing usage, we are realizing the limitations and shortfalls of these systems. Brittleness to minor adversarial changes in the input data, ability to…

计算机与社会 · 计算机科学 2020-11-05 Richa Singh , Mayank Vatsa , Nalini Ratha

AI-based systems are widely employed nowadays to make decisions that have far-reaching impacts on individuals and society. Their decisions might affect everyone, everywhere and anytime, entailing concerns about potential human rights…

How to make artificial intelligence (AI) systems safe and aligned with human values is an open research question. Proposed solutions tend toward relying on human intervention in uncertain situations, learning human values and intentions…

计算机与社会 · 计算机科学 2023-10-04 Jeffrey W. Johnston

The rise of Artificial Intelligence (AI) will bring with it an ever-increasing willingness to cede decision-making to machines. But rather than just giving machines the power to make decisions that affect us, we need ways to work…

计算机与社会 · 计算机科学 2020-12-14 Elisa Bertino , Finale Doshi-Velez , Maria Gini , Daniel Lopresti , David Parkes

Generative AI (GenAI) poses a substantial threat to the integrity of information within the contemporary public sphere, which increasingly relies on social media platforms as intermediaries for news consumption. At present, most research…

计算机与社会 · 计算机科学 2025-11-11 C. Bowman Kerbage

In this paper, we propose "Confident AI" as a means to designing Artificial Intelligence (AI) and Machine Learning (ML) systems with both algorithm and user confidence in model predictions and reported results. The 4 basic tenets of…

人工智能 · 计算机科学 2022-02-15 Jim Davis

Synthetic images, audio, and video can now be generated and edited by Artificial Intelligence (AI). In particular, the malicious use of synthetic data has raised concerns about potential harms to cybersecurity, personal privacy, and public…

人机交互 · 计算机科学 2025-08-05 Yingfan Zhou , Ester Chen , Manasa Pisipati , Aiping Xiong , Sarah Rajtmajer

Human perception, memory and decision-making are impacted by tens of cognitive biases and heuristics that influence our actions and decisions. Despite the pervasiveness of such biases, they are generally not leveraged by today's Artificial…

人机交互 · 计算机科学 2023-12-04 Aditya Gulati , Miguel Angel Lozano , Bruno Lepri , Nuria Oliver

Existing strategies for managing risks from advanced AI systems often focus on affecting what AI systems are developed and how they diffuse. However, this approach becomes less feasible as the number of developers of advanced AI grows, and…

计算机与社会 · 计算机科学 2025-01-24 Jamie Bernardi , Gabriel Mukobi , Hilary Greaves , Lennart Heim , Markus Anderljung

Artificial Intelligence is rapidly embedding itself within militaries, economies, and societies, reshaping their very foundations. Given the depth and breadth of its consequences, it has never been more pressing to understand how to ensure…

机器学习 · 计算机科学 2024-11-19 Dan Hendrycks

Artificial intelligence (AI) methods have been proposed for the prediction of social behaviors which could be reasonably understood from patient-reported information. This raises novel ethical concerns about respect, privacy, and control…

It is curious that AI increasingly outperforms human decision makers, yet much of the public distrusts AI to make decisions affecting their lives. In this paper we explore a novel theory that may explain one reason for this. We propose that…

计算机与社会 · 计算机科学 2022-08-03 Bran Knowles , Jason D'Cruz , John T. Richards , Kush R. Varshney

Artificial intelligence (AI) systems attempt to imitate human behavior. How well they do this imitation is often used to assess their utility and to attribute human-like (or artificial) intelligence to them. However, most work on AI refers…

计算机与社会 · 计算机科学 2022-11-24 Vinodkumar Prabhakaran , Rida Qadri , Ben Hutchinson

To benefit from AI advances, users and operators of AI systems must have reason to trust it. Trust arises from multiple interactions, where predictable and desirable behavior is reinforced over time. Providing the system's users with some…

人工智能 · 计算机科学 2022-01-27 Stephanie Galaitsi , Benjamin D. Trump , Jeffrey M. Keisler , Igor Linkov , Alexander Kott

How much are we to trust a decision made by an AI algorithm? Trusting an algorithm without cause may lead to abuse, and mistrusting it may similarly lead to disuse. Trust in an AI is only desirable if it is warranted; thus, calibrating…

人机交互 · 计算机科学 2023-03-27 Neil Natarajan , Reuben Binns , Jun Zhao , Nigel Shadbolt