English
Related papers

Related papers: Ethical Implications of Training Deceptive AI

200 papers

The recent breakthroughs of deep reinforcement learning (DRL) technique in Alpha Go and playing Atari have set a good example in handling large state and actions spaces of complicated control problems. The DRL technique is comprised of (i)…

Artificial Intelligence · Computer Science 2017-10-12 Hongjia Li , Tianshu Wei , Ao Ren , Qi Zhu , Yanzhi Wang

This study proposes an "AI Development Support" approach that, unlike conventional AI Alignment-which aims to forcefully inject human values-supports the ethical and moral development of AI itself. As demonstrated by the Orthogonality…

Artificial Intelligence · Computer Science 2025-02-28 Taichiro Endo

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their deployment is frequently undermined by undesirable behaviors such as generating harmful content, factual inaccuracies, and societal biases. Diagnosing the…

Computation and Language · Computer Science 2025-10-06 Zhe Li , Wei Zhao , Yige Li , Jun Sun

Large language models (LLMs) aligned for safety through techniques like reinforcement learning from human feedback (RLHF) often exhibit emergent deceptive behaviors, where outputs appear compliant but subtly mislead or omit critical…

Machine Learning · Computer Science 2025-07-15 Santhosh Kumar Ravindran

Large language models of artificial intelligence (AI), such as ChatGPT, find remarkable but controversial applicability in science and research. This paper reviews epistemological challenges, ethical and integrity risks in science conduct…

Computers and Society · Computer Science 2023-08-01 Evangelos Pournaras

In recent years, Deep Reinforcement Learning (DRL) has emerged as a promising method for robot collision avoidance. However, such DRL models often come with limitations, such as adapting effectively to structured environments containing…

Robotics · Computer Science 2023-10-27 Max Asselmeier , Zhaoyi Li , Kelin Yu , Danfei Xu

The adoption of machine learning (ML) and, more specifically, deep learning (DL) applications into all major areas of our lives is underway. The development of trustworthy AI is especially important in medicine due to the large implications…

Machine Learning · Computer Science 2024-02-22 Daniel Schwabe , Katinka Becker , Martin Seyferth , Andreas Klaß , Tobias Schäffter

Machine learning (ML) and artificial intelligence (AI) researchers play an important role in the ethics and governance of AI, including taking action against what they perceive to be unethical uses of AI (Belfield, 2020; Van Noorden, 2020).…

Computers and Society · Computer Science 2021-05-06 Baobao Zhang , Markus Anderljung , Lauren Kahn , Noemi Dreksler , Michael C. Horowitz , Allan Dafoe

As political parties around the world experiment with Artificial Intelligence (AI) in election campaigns, concerns about deception and manipulation are rising. This article examines how the public reacts to different uses of AI in elections…

Computers and Society · Computer Science 2025-05-20 Andreas Jungherr , Adrian Rauchfleisch , Alexander Wuttke

We propose an information-theoretic formalization of the distinction between two fundamental AI safety failure modes: deceptive alignment and goal drift. While both can lead to systems that appear misaligned, we demonstrate that they…

Artificial Intelligence · Computer Science 2026-03-31 Robin Young

Aligning generative diffusion models with human preferences via reinforcement learning (RL) is critical yet challenging. Most existing algorithms are often vulnerable to reward hacking, such as quality degradation, over-stylization, or…

Humans are known to construct cognitive maps of their everyday surroundings using a variety of perceptual inputs. As such, when a human is asked for directions to a particular location, their wayfinding capability in converting this…

Robotics · Computer Science 2020-11-03 Vishnu Sashank Dorbala , Arjun Srinivasan , Aniket Bera

Federated Learning (FL) has emerged as a critical paradigm for enabling privacy-preserving machine learning, particularly in regulated sectors such as finance and healthcare. However, standard FL strategies often encounter significant…

Cryptography and Security · Computer Science 2025-06-02 Abhijit Talluri

The rapid advancements in large language models (LLMs) have revolutionized natural language processing, unlocking unprecedented capabilities in communication, automation, and knowledge generation. However, the ethical implications of LLM…

Computers and Society · Computer Science 2026-01-27 Javed I. Khan , Sharmila Rahman Prithula

Artificial intelligence (AI) research is routinely criticized for its real and potential impacts on society, and we lack adequate institutional responses to this criticism and to the responsibility that it reflects. AI research often falls…

Computers and Society · Computer Science 2021-07-13 Michael S. Bernstein , Margaret Levi , David Magnus , Betsy Rajala , Debra Satz , Charla Waeiss

Deceptive games are games where the reward structure or other aspects of the game are designed to lead the agent away from a globally optimal policy. While many games are already deceptive to some extent, we designed a series of games in…

Artificial Intelligence · Computer Science 2018-02-06 Damien Anderson , Matthew Stephenson , Julian Togelius , Christian Salge , John Levine , Jochen Renz

As generative AI models become increasingly integrated into high-stakes domains, the need for robust methods to evaluate their ethical reasoning becomes increasingly important. This paper introduces a five-dimensional audit model --…

Artificial Intelligence · Computer Science 2025-04-25 W. Russell Neuman , Chad Coleman , Ali Dasdan , Safinah Ali , Manan Shah

The technical progression of artificial intelligence (AI) research has been built on breakthroughs in fields such as computer science, statistics, and mathematics. However, in the past decade AI researchers have increasingly looked to the…

Computers and Society · Computer Science 2023-06-06 Will Hawkins , Brent Mittelstadt

Pre-trained Large Language Models (LLMs) are an integral part of modern AI that have led to breakthrough performances in complex AI tasks. Major AI companies with expensive infrastructures are able to develop and train these large models…

Cryptography and Security · Computer Science 2023-05-02 Rouzbeh Behnia , Mohamamdreza Ebrahimi , Jason Pacheco , Balaji Padmanabhan

The risks of frontier AI may require international cooperation, which in turn may require verification: checking that all parties follow agreed-on rules. For instance, states might need to verify that powerful AI models are widely deployed…

Computers and Society · Computer Science 2025-07-29 Mauricio Baker , Gabriel Kulp , Oliver Marks , Miles Brundage , Lennart Heim