English
Related papers

Related papers: A testable framework for AI alignment: Simulation …

200 papers

Scaling test-time compute has emerged as a powerful mechanism for enhancing Large Language Model (LLM) performance. However, standard post-training paradigms, Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), optimize the…

Machine Learning · Computer Science 2026-05-21 Adam Ousherovitch , Ambuj Tewari

Artificial intelligence (AI) is increasingly being considered to assist human decision-making in high-stake domains (e.g. health). However, researchers have discussed an issue that humans can over-rely on wrong suggestions of the AI model…

Human-Computer Interaction · Computer Science 2023-08-09 Min Hun Lee , Chong Jun Chew

The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizing the satisfaction of preferences, and (3) that AI systems…

Artificial Intelligence · Computer Science 2024-11-12 Tan Zhi-Xuan , Micah Carroll , Matija Franklin , Hal Ashton

Safety validation is a crucial component in the development and deployment of autonomous systems, such as self-driving vehicles and robotic systems. Ensuring safe operation necessitates extensive testing and verification of control…

Systems and Control · Electrical Eng. & Systems 2023-05-11 Ali Baheri , Mykel J. Kochenderfer

The field of fair AI aims to counter biased algorithms through computational modelling. However, it faces increasing criticism for perpetuating the use of overly technical and reductionist methods. As a result, novel approaches appear in…

Computers and Society · Computer Science 2024-09-26 Miriam Fahimi , Mayra Russo , Kristen M. Scott , Maria-Esther Vidal , Bettina Berendt , Katharina Kinder-Kurlanda

The rise of artificial intelligence (A.I.) based systems is already offering substantial benefits to the society as a whole. However, these systems may also enclose potential conflicts and unintended consequences. Notably, people will tend…

Computers and Society · Computer Science 2020-12-23 Pedro Fernandes , Francisco C. Santos , Manuel Lopes

Artificial intelligence is humanity's most promising technology because of the remarkable capabilities offered by foundation models. Yet, the same technology brings confusion and consternation: foundation models are poorly understood and…

Artificial Intelligence · Computer Science 2025-07-01 Rishi Bommasani

Artificial intelligence systems are increasingly deployed in domains that shape human behaviour, institutional decision-making, and societal outcomes. Existing responsible AI and governance efforts provide important normative principles but…

Artificial Intelligence · Computer Science 2025-12-19 Otman A. Basir

In this paper, we propose a test-time adaptive agent that performs exploratory inference through posterior-guided belief refinement without relying on gradient-based updates or additional training for LLM agent operating under partial…

Artificial Intelligence · Computer Science 2026-01-01 Seohui Bae , Jeonghye Kim , Youngchul Sung , Woohyung Lim

Populating our world with hyperintelligent machines obliges us to examine cognitive behaviors observed across domains that suggest autonomy may be a fundamental property of cognitive systems, and while not inherently adversarial, it…

Neurons and Cognition · Quantitative Biology 2025-06-09 Andrea Morris

AI agents -- systems that combine foundation models with reasoning, planning, memory, and tool use -- are rapidly becoming a practical interface between natural-language intent and real-world computation. This survey synthesizes the…

Artificial Intelligence · Computer Science 2026-01-06 Bin Xu

The emergence of large language models (LLMs) has sparked the possibility of about Artificial Superintelligence (ASI), a hypothetical AI system surpassing human intelligence. However, existing alignment paradigms struggle to guide such…

Machine Learning · Computer Science 2024-12-30 HyunJin Kim , Xiaoyuan Yi , Jing Yao , Jianxun Lian , Muhua Huang , Shitong Duan , JinYeong Bak , Xing Xie

Robotic manipulation with deformable objects represents a data-intensive regime in embodied learning, where shape, contact, and topology co-evolve in ways that far exceed the variability of rigids. Although simulation promises relief from…

Recent work proposes using world models to generate controlled virtual environments in which AI agents can be tested before deployment to ensure their reliability and safety. However, accurate world models often have high computational…

Artificial Intelligence · Computer Science 2025-04-08 Fernando Rosas , Alexander Boyd , Manuel Baltieri

Much of the recent work developing formal methods techniques to specify or learn the behavior of autonomous systems is predicated on a belief that formal specifications are interpretable and useful for humans when checking systems. Though…

Artificial Intelligence · Computer Science 2023-05-30 Ho Chit Siu , Kevin Leahy , Makai Mann

Foundation models (FMs) face a critical safety challenge: as capabilities scale, instrumental convergence drives default trajectories toward loss of human control, potentially culminating in existential catastrophe. Current alignment…

Artificial Intelligence · Computer Science 2025-06-04 Ram Potham , Max Harms

In AI-assisted decision-making, it is critical for human decision-makers to know when to trust AI and when to trust themselves. However, prior studies calibrated human trust only based on AI confidence indicating AI's correctness likelihood…

Human-Computer Interaction · Computer Science 2023-01-18 Shuai Ma , Ying Lei , Xinru Wang , Chengbo Zheng , Chuhan Shi , Ming Yin , Xiaojuan Ma

The development of LLMs has elevated AI agents from task-specific tools to long-lived, decision-making entities. Yet, most architectures remain static and reactive, tethered to manually defined, narrow scenarios. These systems excel at…

Artificial Intelligence · Computer Science 2025-12-23 Mingyang Sun , Feng Hong , Weinan Zhang

This paper contributes to the nascent debate around safety cases for frontier AI systems. Safety cases are structured, defensible arguments that a system is acceptably safe to deploy in a given context. Historically, they have been used in…

Computers and Society · Computer Science 2026-03-11 Shaun Feakins , Ibrahim Habli , Phillip Morgan