English
Related papers

Related papers: On the possibility of deep alignment

200 papers

One of the remarkable feats of intelligent life is that it restructures the world it lives in for its own benefit. This extended abstract outlines how the information-theoretic principle of empowerment, as an intrinsic motivation, can be…

Adaptation and Self-Organizing Systems · Physics 2013-10-15 Christoph Salge , Daniel Polani

In reinforcement learning, an agent learns to reach a set of goals by means of an external reward signal. In the natural world, intelligent organisms learn from internal drives, bypassing the need for external signals, which is beneficial…

Machine Learning · Computer Science 2020-06-16 Rui Zhao , Yang Gao , Pieter Abbeel , Volker Tresp , Wei Xu

Ensuring fairness in decentralized multi-agent systems presents significant challenges due to emergent biases, systemic inefficiencies, and conflicting agent incentives. This paper provides a comprehensive survey of fairness in multi-agent…

Multiagent Systems · Computer Science 2025-03-04 Rajesh Ranjan , Shailja Gupta , Surya Narayan Singh

We propose a multi-agent distributed reinforcement learning algorithm that balances between potentially conflicting short-term reward and sparse, delayed long-term reward, and learns with partial information in a dynamic environment. We…

Machine Learning · Computer Science 2022-04-06 Jing Tan , Ramin Khalili , Holger Karl

This paper develops a control-theoretic framework for analyzing agentic systems embedded within feedback control loops, where an AI agent may adapt controller parameters, select among control strategies, invoke external tools, reconfigure…

Systems and Control · Electrical Eng. & Systems 2026-03-26 Ali Eslami , Jiangbo Yu

Artificial Intelligence (AI) agents have rapidly evolved from specialized, rule-based programs to versatile, learning-driven autonomous systems capable of perception, reasoning, and action in complex environments. The explosion of data,…

The artificial intelligence industry is not an isolated economic phenomenon; it is the current physical substrate for a broader, multi-billion-year process: the evolution of an abstract intelligence on Earth. As the scale of computation…

Physics and Society · Physics 2026-05-29 William Yicheng Zhu , Lei Zhu

The reasoning capabilities of embodied agents introduce a critical, under-explored inferential privacy challenge, where the risk of an agent generate sensitive conclusions from ambient data. This capability creates a fundamental tension…

Human-Computer Interaction · Computer Science 2025-09-24 Shuning Zhang , Hong Jia , Simin Li , Ting Dang , Yongquan `Owen' Hu , Xin Yi , Hewu Li

Cooperation is vital to our survival and progress. Evolutionary game theory offers a lens to understand the structures and incentives that enable cooperation to be a successful strategy. As artificial intelligence agents become integral to…

Multiagent Systems · Computer Science 2025-04-29 Tomer Jordi Chaffer , Justin Goldston , Gemach D. A. T. A.

This paper proposes an intent-aware multi-agent planning framework as well as a learning algorithm. Under this framework, an agent plans in the goal space to maximize the expected utility. The planning process takes the belief of other…

Artificial Intelligence · Computer Science 2018-03-07 Siyuan Qi , Song-Chun Zhu

We model endogenous perception of private information in single-agent screening problems, with potential evaluation errors. The agent's evaluation of their type depends on their cognitive state: either attentive (i.e., they correctly…

Theoretical Economics · Economics 2025-03-12 Benjamin Balzer , Benjamin Young

We propose the creation of a systematic effort to identify and replicate key findings in neuropsychology and allied fields related to understanding human values. Our aim is to ensure that research underpinning the value alignment problem of…

Artificial Intelligence · Computer Science 2018-09-11 Gopal P. Sarma , Nick J. Hay , Adam Safron

Active inference is an ambitious theory that treats perception, inference and action selection of autonomous agents under the heading of a single principle. It suggests biologically plausible explanations for many cognitive phenomena,…

Artificial Intelligence · Computer Science 2018-06-22 Martin Biehl , Christian Guckelsberger , Christoph Salge , Simón C. Smith , Daniel Polani

An important step in the development of value alignment (VA) systems in AI is understanding how values can interrelate with facts. Designers of future VA systems will need to utilize a hybrid approach in which ethical reasoning and…

Artificial Intelligence · Computer Science 2019-07-15 Tae Wan Kim , Thomas Donaldson , John Hooker

The objective of a reinforcement learning agent is to behave so as to maximise the sum of a suitable scalar function of state: the reward. These rewards are typically given and immutable. In this paper, we instead consider the proposition…

Artificial Intelligence · Computer Science 2020-08-25 Zeyu Zheng , Junhyuk Oh , Matteo Hessel , Zhongwen Xu , Manuel Kroiss , Hado van Hasselt , David Silver , Satinder Singh

Value-alignment in normative multi-agent systems is used to promote a certain value and to ensure the consistent behavior of agents in autonomous intelligent systems with human values. However, the current literature is limited to…

Multiagent Systems · Computer Science 2023-05-15 Maha Riad , Vinicius Renan de Carvalho , Fatemeh Golpayegani

This paper proposes a definition of system health in the context of multiple agents optimizing a joint reward function. We use this definition as a credit assignment term in a policy gradient algorithm to distinguish the contributions of…

Machine Learning · Computer Science 2021-01-06 Ross E. Allen , Jayesh K. Gupta , Jaime Pena , Yutai Zhou , Javona White Bear , Mykel J. Kochenderfer

Many real-world human behaviors can be characterized as a sequential decision making processes, such as urban travelers choices of transport modes and routes (Wu et al. 2017). Differing from choices controlled by machines, which in general…

Artificial Intelligence · Computer Science 2019-07-12 Guojun Wu , Yanhua Li , Zhenming Liu , Jie Bao , Yu Zheng , Jieping Ye , Jun Luo

Ensuring artificial intelligence behaves in such a way that is aligned with human values is commonly referred to as the alignment challenge. Prior work has shown that rational agents, behaving in such a way that maximizes a utility…

Artificial Intelligence · Computer Science 2024-02-16 Paulo Garcia

AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is intended to provide value for a principal: the user on whose…

Artificial Intelligence · Computer Science 2021-02-09 Simon Zhuang , Dylan Hadfield-Menell