中文
相关论文

相关论文: Information-theoretic Distinctions Between Decepti…

200 篇论文

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today's AI models. This…

人工智能 · 计算机科学 2024-11-26 Robert West , Roland Aydin

Alignment faking is a form of strategic deception in AI in which models selectively comply with training objectives when they infer that they are in training, while preserving different behavior outside training. The phenomenon was first…

Readers can have different goals with respect to the text that they are reading. Can these goals be decoded from their eye movements over the text? In this work, we examine for the first time whether it is possible to distinguish between…

计算与语言 · 计算机科学 2025-02-28 Omer Shubi , Cfir Avraham Hadar , Yevgeni Berzak

Automated verbal deception detection using methods from Artificial Intelligence (AI) has been shown to outperform humans in disentangling lies from truths. Research suggests that transparency and interpretability of computational methods…

人机交互 · 计算机科学 2026-04-10 Riccardo Loconte , Merylin Monaro , Pietro Pietrini , Bruno Verschuere , Bennett Kleinberg

The notion of concept drift refers to the phenomenon that the distribution, which is underlying the observed data, changes over time; as a consequence machine learning models may become inaccurate and need adjustment. While there do exist…

机器学习 · 计算机科学 2020-06-24 Fabian Hinder , Barbara Hammer

Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something "really is" and how something "appears to be", and this gap…

神经元与认知 · 定量生物学 2024-12-30 Tomer Ullman

Deception is a technique to mislead human or computer systems by manipulating beliefs and information. Successful deception is characterized by the information-asymmetric, dynamic, and strategic behaviors of the deceiver and the deceivee.…

密码学与安全 · 计算机科学 2018-10-02 Tao Zhang , Quanyan zhu

Large language models (LLMs) are currently at the forefront of intertwining artificial intelligence (AI) systems with human communication and everyday life. Thus, aligning them with human values is of great importance. However, given the…

计算与语言 · 计算机科学 2024-06-06 Thilo Hagendorff

The widespread adoption of large language models (LLMs) across diverse AI applications is proof of the outstanding achievements obtained in several tasks, such as text mining, text generation, and question answering. However, LLMs are not…

计算与语言 · 计算机科学 2023-11-15 Alessandro Bruno , Pier Luigi Mazzeo , Aladine Chetouani , Marouane Tliba , Mohamed Amine Kerkouri

While much research in artificial intelligence (AI) has focused on scaling capabilities, the accelerating pace of development makes countervailing work on producing harmless, "aligned" systems increasingly urgent. Yet research on alignment…

人工智能 · 计算机科学 2025-12-12 Dani Roytburg , Beck Miller

Confusion is a mental state triggered by cognitive disequilibrium that can occur in many types of task-oriented interaction, including Human-Robot Interaction (HRI). People may become confused while interacting with robots due to…

计算与语言 · 计算机科学 2022-08-22 Na Li , Robert Ross

When AI systems allow human-like communication, they elicit increasingly complex relational responses. Knowledge workers face a particular challenge: They approach these systems as tools while interacting with them in ways that resemble…

人机交互 · 计算机科学 2026-03-02 Emrecan Gulay , Eleonora Picco , Enrico Glerean , Corinna Coupette

Despite significant strides in factual reliability, errors -- often termed hallucinations -- remain a major concern for generative AI, especially as LLMs are increasingly expected to be helpful in more complex or nuanced setups. Yet even in…

计算与语言 · 计算机科学 2026-05-05 Gal Yona , Mor Geva , Yossi Matias

Artificial intelligence (AI) has demonstrated strong potential in clinical diagnostics, often achieving accuracy comparable to or exceeding that of human experts. A key challenge, however, is that AI reasoning frequently diverges from…

人工智能 · 计算机科学 2026-05-25 Belona Sonna , Alban Grastien

In the face of rapidly advancing AI technology, individuals will increasingly rely on AI agents to navigate life's growing complexities, raising critical concerns about maintaining both human agency and autonomy. This paper addresses a…

计算机与社会 · 计算机科学 2025-04-29 Philipp Koralus

Generative AI increasingly supports scientific inference, from protein structure prediction to weather forecasting. Yet its distinctive failure mode, hallucination, raises epistemic alarm bells. I argue that this failure mode can be…

计算机与社会 · 计算机科学 2026-01-14 Charles Rathkopf

Understanding the actions of both humans and artificial intelligence (AI) agents is important before modern AI systems can be fully integrated into our daily life. In this paper, we show that, despite their current huge success, deep…

人工智能 · 计算机科学 2021-01-19 Nodens Koren , Qiuhong Ke , Yisen Wang , James Bailey , Xingjun Ma

Machine learning algorithms operating on mobile networks can be characterized into three different categories. First is the classical situation in which the end-user devices send their data to a central server where this data is used to…

机器学习 · 计算机科学 2020-05-07 Semih Yagli , Alex Dytso , H. Vincent Poor

Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep epistemic divergence. Using two major reasoning benchmarks -…

计算与语言 · 计算机科学 2026-02-13 Eddie Yang , Dashun Wang

Recent empirical research by Sharma et al. (2026) demonstrated that AI assistant interactions carry meaningful potential for situational human disempowerment, including reality distortion, value judgment distortion, and action distortion.…

人机交互 · 计算机科学 2026-02-18 Aleksey Komissarov