中文
相关论文

相关论文: AI Deception: A Survey of Examples, Risks, and Pot…

200 篇论文

Deception is a technique to mislead human or computer systems by manipulating beliefs and information. Successful deception is characterized by the information-asymmetric, dynamic, and strategic behaviors of the deceiver and the deceivee.…

密码学与安全 · 计算机科学 2018-10-02 Tao Zhang , Quanyan zhu

Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their performance leaps lie more insidious and destructive safety risks, namely deception. Unlike…

人工智能 · 计算机科学 2026-05-28 Sitong Fang , Shiyi Hou , Kaile Wang , Boyuan Chen , Donghai Hong , Jiayi Zhou , Josef Dai , Yaodong Yang , Jiaming Ji

Human-AI teams can be vulnerable to catastrophic failure when feedback from the AI is incorrect, especially under high cognitive workload. Traditional team aggregation methods, such as voting, are susceptible to these AI errors, which can…

A fascinating hypothesis is that human and animal intelligence could be explained by a few principles (rather than an encyclopedic list of heuristics). If that hypothesis was correct, we could more easily both understand our own…

机器学习 · 计算机科学 2022-08-02 Anirudh Goyal , Yoshua Bengio

Recent advancements in artificial intelligence (AI) systems, including large language models like ChatGPT, offer promise and peril for scholarly peer review. On the one hand, AI can enhance efficiency by addressing issues like long…

计算机与社会 · 计算机科学 2023-09-25 Laurie A. Schintler , Connie L. McNeely , James Witte

AI safety practitioners invest considerable resources in AI system evaluations, but these investments may be wasted if evaluations fail to realize their impact. This paper questions the core value proposition of evaluations: that they…

计算机与社会 · 计算机科学 2024-08-06 Gabriel Mukobi

Enterprise AI systems, built on large language models, retrieval pipelines and autonomous agents, introduce a class of risks that traditional software quality assurance was never designed to address. These systems are probabilistic,…

软件工程 · 计算机科学 2026-05-25 Chitra Badagi , Divye Singh , Animesh Sen , Adinath Shirsath

Artificial Intelligence (AI) has made impressive progress in recent years and represents a key technology that has a crucial impact on the economy and society. However, it is clear that AI and business models based on it can only reach…

Due to the cultural and governance differences of countries around the world, there currently exists a wide spectrum of AI regulation policy proposals that have created a chaos in the global AI regulatory space. Properly regulating AI…

计算机与社会 · 计算机科学 2023-07-25 Weiyue Wu , Shaoshan Liu

In AI-assisted decision-making, it is critical for human decision-makers to know when to trust AI and when to trust themselves. However, prior studies calibrated human trust only based on AI confidence indicating AI's correctness likelihood…

人机交互 · 计算机科学 2023-01-18 Shuai Ma , Ying Lei , Xinru Wang , Chengbo Zheng , Chuhan Shi , Ming Yin , Xiaojuan Ma

AI agents have been boosted by large language models. AI agents can function as intelligent assistants and complete tasks on behalf of their users with access to tools and the ability to execute commands in their environments. Through…

密码学与安全 · 计算机科学 2024-12-19 Yifeng He , Ethan Wang , Yuyang Rong , Zifei Cheng , Hao Chen

Artificial intelligence and machine learning are increasingly used to offload decision making from people. In the past, one of the rationales for this replacement was that machines, unlike people, can be fair and unbiased. Evidence suggests…

计算机与社会 · 计算机科学 2024-09-27 Will Bridewell , Paul F. Bello , Selmer Bringsjord

The use of Artificial Intelligence (AI) in high-risk, decision-making scenarios presents technical, safety, and normative challenges; problems that may only be ameliorated by human oversight. However, notions of human oversight lack a…

Risk thresholds provide a measure of the level of risk exposure that a society or individual is willing to withstand, ultimately shaping how we determine the safety of technological systems. Against the backdrop of the Cold War, the first…

计算机与社会 · 计算机科学 2025-04-22 Heidy Khlaaf , Sarah Myers West

Background: Deception detection is a prevalent problem for security practitioners. With a need for more large-scale approaches, automated methods using machine learning have gained traction. However, detection performance still implies…

计算与语言 · 计算机科学 2020-03-31 Bennett Kleinberg , Bruno Verschuere

Organizations of all sizes, across all industries and domains are leveraging artificial intelligence (AI) technologies to solve some of their biggest challenges around operations, customer experience, and much more. However, due to the…

计算机与社会 · 计算机科学 2022-11-24 Navdeep Gill , Abhishek Mathur , Marcos V. Conde

Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify…

The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans can perform. Despite how useful these systems might be,…

Artificial intelligence (AI) is increasingly of tremendous interest in the medical field. However, failures of medical AI could have serious consequences for both clinical outcomes and the patient experience. These consequences could erode…

人工智能 · 计算机科学 2020-08-19 Thomas P. Quinn , Manisha Senadeera , Stephan Jacobs , Simon Coghlan , Vuong Le

Since the release of ChatGPT, there has been a lot of debate about whether AI systems pose an existential risk to humanity. This paper develops a general framework for thinking about the existential risk of AI systems. We analyze a two…

人工智能 · 计算机科学 2026-01-16 Herman Cappelen , Simon Goldstein , John Hawthorne