English
Related papers

Related papers: Neurodivergent Influenceability as a Contingent So…

200 papers

As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Current benchmarks for AI agents primarily evaluate refusal of…

Artificial Intelligence · Computer Science 2026-05-12 Miles Q. Li , Benjamin C. M. Fung , Martin Weiss , Pulei Xiong , Khalil Al-Hussaeni , Claude Fachkha

AI chatbots are increasingly stepping into roles as collaborators or teachers in analyzing, visualizing, and reasoning through data and domain problem. Yet, AI's default assistant mode with its comprehensive and one-off responses may…

Human-Computer Interaction · Computer Science 2026-04-06 Yongsu Ahn , Nam Wook Kim , Benjamin Bach

AI-related incidents are becoming increasingly frequent and severe, ranging from safety failures to misuse by malicious actors. In such complex situations, identifying which elements caused an adverse outcome, the problem of cause…

Artificial Intelligence · Computer Science 2026-03-17 Maria Victoria Carro , David Lagnado

Researchers worried about catastrophic risks from advanced AI have argued that we should expect sufficiently capable AI agents to pursue power over humanity because power is a convergent instrumental goal, something that is useful for a…

Artificial Intelligence · Computer Science 2025-06-10 Christian Tarsney

AI for Social Impact (AI4SI) has achieved compelling results in public health, conservation, and security, yet scaling these successes remains difficult due to a persistent deployment bottleneck. We characterize this bottleneck through…

Computers and Society · Computer Science 2026-01-09 Lingkai Kong , Cheol Woo Kim , Davin Choo , Milind Tambe

We develop and study new adversarial perturbations that enable an attacker to gain control over decisions in generic Artificial Intelligence (AI) systems including deep learning neural networks. In contrast to adversarial data modification,…

Cryptography and Security · Computer Science 2023-12-07 Ivan Y. Tyukin , Desmond J. Higham , Alexander Bastounis , Eliyas Woldegeorgis , Alexander N. Gorban

AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithic Artificial General Intelligence (AGI). The alternative AGI…

Artificial Intelligence · Computer Science 2026-05-20 Nenad Tomašev , Matija Franklin , Julian Jacobs , Sébastien Krier , Simon Osindero

As artificial intelligence (AI) improves, traditional alignment strategies may falter in the face of unpredictable self-improvement, hidden subgoals, and the sheer complexity of intelligent systems. Inspired by contemplative wisdom…

Artificial Intelligence · Computer Science 2025-08-19 Ruben Laukkonen , Fionn Inglis , Shamil Chandaria , Lars Sandved-Smith , Edmundo Lopez-Sola , Jakob Hohwy , Jonathan Gold , Adam Elwood

Jurisprudence, the study of how judges should properly decide cases, and alignment, the science of getting AI models to conform to human values, share a fundamental structure. These seemingly distant fields both seek to predict and shape…

Artificial Intelligence · Computer Science 2026-05-12 Nicholas Caputo

AI Alignment research seeks to align human and AI goals to ensure independent actions by a machine are always ethical. This paper argues empathy is necessary for this task, despite being often neglected in favor of more deductive…

Neural and Evolutionary Computing · Computer Science 2023-12-14 Devin Gonier , Adrian Adduci , Cassidy LoCascio

As Generative AI systems increasingly engage in long-term, personal, and relational interactions, human-AI engagements are becoming significantly complex, making them more challenging to understand and govern. These Interactive AI systems…

Computers and Society · Computer Science 2025-08-26 Yulu Pi , Cagatay Turkay , Daniel Bogiatzis-Gibbons

Artificially intelligent agents are increasingly being integrated into human decision-making: from large language model (LLM) assistants to autonomous vehicles. These systems often optimize their individual objective, leading to conflicts,…

Machine Learning · Computer Science 2025-02-07 Juan Agustin Duque , Milad Aghajohari , Tim Cooijmans , Razvan Ciuca , Tianyu Zhang , Gauthier Gidel , Aaron Courville

Regulation of advanced technologies such as Artificial Intelligence (AI) has become increasingly important, given the associated risks and apparent ethical issues. With the great benefits promised from being able to first supply such…

Artificial Intelligence · Computer Science 2022-01-05 Theodor Cimpeanu , Francisco C. Santos , Luis Moniz Pereira , Tom Lenaerts , The Anh Han

From its inception, AI has had a rather ambivalent relationship to humans---swinging between their augmentation and replacement. Now, as AI technologies enter our everyday lives at an ever increasing pace, there is a greater need for AI…

Artificial Intelligence · Computer Science 2019-10-17 Subbarao Kambhampati

Artificial intelligence (AI) systems, such as machine learning algorithms, have allowed scientists, marketers and governments to shed light on correlations that remained invisible until now. Beforehand, the dots that we had to connect in…

Computers and Society · Computer Science 2022-02-08 Remy Demichelis

For billions of years, evolution has been the driving force behind the development of life, including humans. Evolution endowed humans with high intelligence, which allowed us to become one of the most successful species on the planet.…

Computers and Society · Computer Science 2023-07-21 Dan Hendrycks

AI Safety researchers attempting to align values of highly capable intelligent systems with those of humanity face a number of challenges including personal value extraction, multi-agent value merger and finally in-silico encoding.…

Artificial Intelligence · Computer Science 2019-01-08 Roman V. Yampolskiy

AI companies increasingly develop and deploy privacy-enhancing technologies, bias-constraining measures, evaluation frameworks, and alignment techniques -- framing them as addressing concerns related to data privacy, algorithmic fairness,…

Computers and Society · Computer Science 2025-10-03 Rui-Jie Yew , Brian Judge

As AI systems grow more capable and autonomous, ensuring their safety and reliability requires not only model-level alignment but also strategic oversight of the humans and institutions involved in their development and deployment. Existing…

Artificial Intelligence · Computer Science 2026-02-10 Cheol Woo Kim , Davin Choo , Tzeh Yuan Neoh , Milind Tambe

This paper introduces Admissibility Alignment: a reframing of AI alignment as a property of admissible action and decision selection over distributions of outcomes under uncertainty, evaluated through the behavior of candidate policies. We…

Artificial Intelligence · Computer Science 2026-01-06 Chris Duffey