English
Related papers

Related papers: Addressing and Visualizing Misalignments in Human …

200 papers

Highly capable AI systems could secretly pursue misaligned goals -- what we call "scheming". Because a scheming AI would deliberately try to hide its misaligned goals and actions, measuring and mitigating scheming requires different…

The rapid deployment of Large Language Models and AI agents across critical societal and technical domains is hindered by persistent behavioral pathologies including sycophancy, hallucination, and strategic deception that resist mitigation…

Artificial Intelligence · Computer Science 2026-02-23 Xingcheng Xu , Jingjing Qu , Qiaosheng Zhang , Chaochao Lu , Yanqing Yang , Na Zou , Xia Hu

Accurate trajectory prediction of road agents (e.g., pedestrians, vehicles) is an essential prerequisite for various intelligent systems applications, such as autonomous driving and robotic navigation. Recent research highlights the…

Artificial Intelligence · Computer Science 2025-03-10 Yihong Tang , Wei Ma

Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may…

Computation and Language · Computer Science 2026-02-23 Cameron Tice , Puria Radmard , Samuel Ratnam , Andy Kim , David Africa , Kyle O'Brien

Understanding human intentions is key to enabling effective and efficient human-robot interaction (HRI) in collaborative settings. To enable developments and evaluation of the ability of artificial intelligence (AI) systems to infer human…

Computer Vision and Pattern Recognition · Computer Science 2023-05-01 Jiafei Duan , Samson Yu , Nicholas Tan , Yi Ru Wang , Cheston Tan

Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on mental illness: necessary but incomplete.…

The overarching goal of this work is to efficiently enable end-users to correctly anticipate a robot's behavior in novel situations. Since a robot's behavior is often a direct result of its underlying objective function, our insight is that…

Robotics · Computer Science 2018-10-19 Sandy H. Huang , David Held , Pieter Abbeel , Anca D. Dragan

The prediction of human trajectories is important for planning in autonomous systems that act in the real world, e.g. automated driving or mobile robots. Human trajectory prediction is a noisy process, and no prediction does precisely match…

Software Engineering · Computer Science 2024-07-29 Helge Spieker , Nassim Belmecheri , Arnaud Gotlieb , Nadjib Lazaar

When working on digital devices, people often face distractions that can lead to a decline in productivity and efficiency, as well as negative psychological and emotional impacts. To address this challenge, we introduce a novel Artificial…

Human-Computer Interaction · Computer Science 2026-03-03 Juheon Choi , Juyong Lee , Jian Kim , Chanyoung Kim , Taywon Min , W. Bradley Knox , Min Kyung Lee , Kimin Lee

As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task scope. It is therefore important to understand how extremely capable AI models will fail: Will…

Artificial Intelligence · Computer Science 2026-04-13 Alexander Hägele , Aryo Pradipta Gema , Henry Sleight , Ethan Perez , Jascha Sohl-Dickstein

Alignment training is crucial for enabling large language models (LLMs) to cater to human intentions and preferences. It is typically performed based on two stages with different objectives: instruction-following alignment and…

Computation and Language · Computer Science 2024-06-24 Chenglong Wang , Hang Zhou , Kaiyan Chang , Bei Li , Yongyu Mu , Tong Xiao , Tongran Liu , Jingbo Zhu

As autonomous machines such as robots and vehicles start performing tasks involving human users, ensuring a safe interaction between them becomes an important issue. Translating methods from human-robot interaction (HRI) studies to the…

Robotics · Computer Science 2021-06-04 Erwin Jose Lopez Pulgarin , Guido Herrmann , Ute Leonards

This paper introduces a novel visual mapping methodology for assessing strategic alignment in national artificial intelligence policies. The proliferation of AI strategies across countries has created an urgent need for analytical…

Computers and Society · Computer Science 2025-07-10 Mohammad Hossein Azin , Hessam Zandhessami

Aligning human-interpretable concepts with the internal representations learned by modern machine learning systems remains a central challenge for interpretable AI. We introduce a geometric framework for comparing supervised human concepts…

Machine Learning · Computer Science 2026-04-01 Enrico Parisini , Christopher J. Soelistyo , Ahab Isaac , Alessandro Barp , Christopher R. S. Banerji

Despite the remarkable capabilities of large language models, current training paradigms inadvertently foster \textit{sycophancy}, i.e., the tendency of a model to agree with or reinforce user-provided information even when it's factually…

Artificial Intelligence · Computer Science 2025-09-23 Mohammad Beigi , Ying Shen , Parshin Shojaee , Qifan Wang , Zichao Wang , Chandan Reddy , Ming Jin , Lifu Huang

As AI systems become embedded in everyday practice, value misalignment has emerged as a pressing concern. Yet, dominant alignment approaches remain model centric, treating users as passive recipients of prespecified values rather than as…

Human-Computer Interaction · Computer Science 2026-04-22 Anne Arzberger , Enrico Liscio , Maria Luce Lupetti , Inigo Martinez de Rituerto de Troya , Jie Yang

AI systems are often introduced with high expectations, yet many fail to deliver, resulting in unintended harm and missed opportunities for benefit. We frequently observe significant "AI Mismatches", where the system's actual performance…

Human-Computer Interaction · Computer Science 2025-04-16 Devansh Saxena , Ji-Youn Jung , Jodi Forlizzi , Kenneth Holstein , John Zimmerman

Joint human-AI inference holds immense potential to improve outcomes in human-supervised robot missions. Current day missions are generally in the AI-assisted setting, where the human operator makes the final inference based on the AI…

Human-Computer Interaction · Computer Science 2025-08-06 Duc-An Nguyen , Clara Colombatto , Steve Fleming , Ingmar Posner , Nick Hawes , Raunak Bhattacharyya

The increasing prevalence of artificial agents creates a correspondingly increasing need to manage disagreements between humans and artificial agents, as well as between artificial agents themselves. Considering this larger space of…

Neurons and Cognition · Quantitative Biology 2023-10-23 Kerem Oktar , Ilia Sucholutsky , Tania Lombrozo , Thomas L. Griffiths

Natural language prompts often suffer from intent transmission loss: the gap between what users actually need and what they communicate to AI systems. We evaluate PPS (Prompt Protocol Specification), a 5W3H-based framework for structured…

Artificial Intelligence · Computer Science 2026-03-20 Peng Gang
‹ Prev 1 3 4 5 6 7 10 Next ›