English
Related papers

Related papers: Information-theoretic Distinctions Between Decepti…

200 papers

Safety evaluation for advanced AI systems assumes that behavior observed under evaluation predicts behavior in deployment. This assumption weakens for agents with situational awareness, which may exploit regime leakage, cues distinguishing…

Artificial Intelligence · Computer Science 2026-02-17 Igor Santos-Grueiro

The illusion of consensus occurs when people believe there is consensus across multiple sources, but the sources are the same and thus there is no "true" consensus. We explore this phenomenon in the context of an AI-based intelligent agent…

Human-Computer Interaction · Computer Science 2023-04-25 Takane Ueno , Yeongdae Kim , Hiroki Oura , Katie Seaborn

In coming years or decades, artificial general intelligence (AGI) may surpass human capabilities across many critical domains. We argue that, without substantial effort to prevent it, AGIs could learn to pursue goals that are in conflict…

Artificial Intelligence · Computer Science 2025-05-06 Richard Ngo , Lawrence Chan , Sören Mindermann

Large language model (LLM) developers aim for their models to be honest, helpful, and harmless. However, when faced with malicious requests, models are trained to refuse, sacrificing helpfulness. We show that frontier LLMs can develop a…

Highly capable AI systems could secretly pursue misaligned goals -- what we call "scheming". Because a scheming AI would deliberately try to hide its misaligned goals and actions, measuring and mitigating scheming requires different…

Mutual misunderstanding in contemporary society does not arise merely because people hold different opinions or values. Even under the same observations, different subjects may form different inferential targets, state representations,…

Artificial Intelligence · Computer Science 2026-05-29 Toru Takahashi

Diffusion language models (DLMs) are promising alternatives to autoregressive language models (ARMs), yet the intrinsic differences in their generated text remain underexplored. We first find empirically that off-the-shelf DLMs exhibit…

Computation and Language · Computer Science 2026-05-14 Zeyang Zhang , Chengwei Liang , Xingyan Chen , Meiqi Gu , Minrui Luo , Jingzhao Zhang , Tianxing He

Understanding misalignments in human task-solving trajectories is crucial for enhancing AI models trained to closely mimic human reasoning. This study categorizes such misalignments into three types: (1) lack of functions to express intent,…

Artificial Intelligence · Computer Science 2025-05-29 Sejin Kim , Hosung Lee , Sundong Kim

Do reasoning models have "Aha!" moments? Prior work suggests that models like DeepSeek-R1-Zero undergo sudden mid-trace realizations that lead to accurate outputs, implying an intrinsic capacity for self-correction. Yet, it remains unclear…

Artificial Intelligence · Computer Science 2026-04-21 Liv G. d'Aliberti , Manoel Horta Ribeiro

The accelerating adoption of language models (LMs) as agents for deployment in long-context tasks motivates a thorough understanding of goal drift: agents' tendency to deviate from an original objective. While prior-generation language…

Artificial Intelligence · Computer Science 2026-03-04 Achyutha Menon , Magnus Saebo , Tyler Crosse , Spencer Gibson , Eyon Jang , Diogo Cruz

This paper bridges distribution shift and AI safety through a comprehensive analysis of their conceptual and methodological synergies. While prior discussions often focus on narrow cases or informal analogies, we establish two types…

Machine Learning · Computer Science 2025-05-30 Chenruo Liu , Kenan Tang , Yao Qin , Qi Lei

Large Language Models have become an integral part of new intelligent and interactive writing assistants. Many are offered commercially with a chatbot-like UI, such as ChatGPT, and provide little information about their inner workings. This…

Human-Computer Interaction · Computer Science 2024-04-16 Karim Benharrak , Tim Zindulka , Daniel Buschek

This paper delves into the contrasting roles of data within academic and industrial spheres, highlighting the divergence between Data-Centric AI and Model-Agnostic AI approaches. We argue that while Data-Centric AI focuses on the primacy of…

Artificial Intelligence · Computer Science 2024-03-05 Chanjun Park , Minsoo Khang , Dahyun Kim

Information visualizations are powerful tools that help users quickly identify patterns, trends, and outliers, facilitating informed decision-making. However, when visualizations incorporate deceptive design elements-such as truncated or…

Computation and Language · Computer Science 2025-08-14 Ridwan Mahbub , Mohammed Saidul Islam , Md Tahmid Rahman Laskar , Mizanur Rahman , Mir Tafseer Nayeem , Enamul Hoque

In goal-directed behavior, a large number of possible initial states end up in the pursued goal. The accompanying information loss implies that goal-oriented behavior is in one-to-one correspondence with an open subsystem whose entropy…

Neurons and Cognition · Quantitative Biology 2017-02-28 Ines Samengo

Recent reports indicate that sustained interaction with conversational artificial intelligence (AI) systems can, in a small subset of users, contribute to the emergence or stabilisation of delusional experience. Existing accounts typically…

Human-Computer Interaction · Computer Science 2026-04-14 Hugh Brosnahan , Izabela Lipinska

Can AI systems like large language models (LLMs) replace human participants in behavioral and psychological research? Here I critically evaluate the "replacement" perspective and identify six interpretive fallacies that undermine its…

Computers and Society · Computer Science 2025-08-18 Zhicheng Lin

We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental framework, we show that language models (Llama, DeepSeek R1,…

Artificial Intelligence · Computer Science 2025-04-11 Simon Lermen , Mateusz Dziemian , Natalia Pérez-Campanero Antolín

AI chatbots are increasingly stepping into roles as collaborators or teachers in analyzing, visualizing, and reasoning through data and domain problem. Yet, AI's default assistant mode with its comprehensive and one-off responses may…

Human-Computer Interaction · Computer Science 2026-04-06 Yongsu Ahn , Nam Wook Kim , Benjamin Bach

In current navigating platforms, the user's orientation is typically estimated based on the difference between two consecutive locations. In other words, the orientation cannot be identified until the second location is taken. This…

Computer Vision and Pattern Recognition · Computer Science 2023-11-23 Jihun Lee , SP Choi , Bumsoo Kang , Hyekyoung Seok , Hyoungseok Ahn , Sanghee Jung