English
Related papers

Related papers: AI Alignment at Your Discretion

200 papers

Big models have achieved revolutionary breakthroughs in the field of AI, but they might also pose potential concerns. Addressing such concerns, alignment technologies were introduced to make these models conform to human preferences and…

Artificial Intelligence · Computer Science 2024-03-08 Xinpeng Wang , Shitong Duan , Xiaoyuan Yi , Jing Yao , Shanlin Zhou , Zhihua Wei , Peng Zhang , Dongkuan Xu , Maosong Sun , Xing Xie

Understanding the decisions made and actions taken by increasingly complex AI system remains a key challenge. This has led to an expanding field of research in explainable artificial intelligence (XAI), highlighting the potential of…

Numerous fairness metrics have been proposed and employed by artificial intelligence (AI) experts to quantitatively measure bias and define fairness in AI models. Recognizing the need to accommodate stakeholders' diverse fairness…

Artificial Intelligence · Computer Science 2025-02-11 Lin Luo , Yuri Nakao , Mathieu Chollet , Hiroya Inakoshi , Simone Stumpf

It has recently been argued that AI models' representations are becoming aligned as their scale and performance increase. Empirical analyses have been designed to support this idea and conjecture the possible alignment of different…

Machine Learning · Computer Science 2025-02-21 Francesco Insulla , Shuo Huang , Lorenzo Rosasco

AI and ML models have already found many applications in critical domains, such as healthcare and criminal justice. However, fully automating such high-stakes applications can raise ethical or fairness concerns. Instead, in such cases,…

Artificial Intelligence · Computer Science 2023-04-28 Ioannis Papantonis , Vaishak Belle

Oversight and control, which we collectively call supervision, are often discussed as ways to ensure that AI systems are accountable, reliable, and able to fulfill governance and management requirements. However, the requirements for "human…

Artificial Intelligence · Computer Science 2025-11-04 David Manheim , Aidan Homewood

AI Alignment is often presented as an interaction between a single designer and an artificial agent in which the designer attempts to ensure the agent's behavior is consistent with its purpose, and risks arise solely because of conflicts…

Artificial Intelligence · Computer Science 2023-09-14 Steve Phelps , Rebecca Ranson

Recent proposals aiming at regulating artificial intelligence (AI) and automated decision-making (ADM) suggest a particular form of risk regulation, i.e. a risk-based approach. The most salient example is the Artificial Intelligence Act…

Computers and Society · Computer Science 2022-11-14 Carsten Orwat , Jascha Bareis , Anja Folberth , Jutta Jahnel , Christian Wadephul

In AI-assisted decision-making, a central promise of having a human-in-the-loop is that they should be able to complement the AI system by overriding its wrong recommendations. In practice, however, we often see that humans cannot assess…

Human-Computer Interaction · Computer Science 2025-02-05 Jakob Schoeffer , Johannes Jakubik , Michael Voessing , Niklas Kuehl , Gerhard Satzger

Apologies are a powerful tool used in human-human interactions to provide affective support, regulate social processes, and exchange information following a trust violation. The emerging field of AI apology investigates the use of apologies…

Computers and Society · Computer Science 2024-12-23 Hadassah Harland , Richard Dazeley , Hashini Senaratne , Peter Vamplew , Francisco Cruz , Bahareh Nakisa

This paper explores the tension between openness and prudence in AI research, evident in two core principles of the Montr\'eal Declaration for Responsible AI. While the AI community has strong norms around open sharing of research, concerns…

Computers and Society · Computer Science 2020-01-14 Jess Whittlestone , Aviv Ovadya

As algorithms increasingly inform and influence decisions made about individuals, it becomes increasingly important to address concerns that these algorithms might be discriminatory. The output of an algorithm can be discriminatory for many…

Machine Learning · Computer Science 2018-03-19 Úrsula Hébert-Johnson , Michael P. Kim , Omer Reingold , Guy N. Rothblum

AI-related incidents are becoming increasingly frequent and severe, ranging from safety failures to misuse by malicious actors. In such complex situations, identifying which elements caused an adverse outcome, the problem of cause…

Artificial Intelligence · Computer Science 2026-03-17 Maria Victoria Carro , David Lagnado

AI-based systems have been used widely across various industries for different decisions ranging from operational decisions to tactical and strategic ones in low- and high-stakes contexts. Gradually the weaknesses and issues of these…

Human-Computer Interaction · Computer Science 2022-01-13 Morteza Saberi

Frontier AI systems are rapidly advancing in their capabilities to persuade, deceive, and influence human behaviour, with current models already demonstrating human-level persuasion and strategic deception in specific contexts. Humans are…

Artificial Intelligence · Computer Science 2025-07-18 Rishane Dassanayake , Mario Demetroudi , James Walpole , Lindley Lentati , Jason R. Brown , Edward James Young

Aligning AI systems with human privacy preferences requires understanding individuals' nuanced disclosure behaviors beyond general norms. Yet eliciting such boundaries remains challenging due to the context-dependent nature of privacy…

Cryptography and Security · Computer Science 2025-09-29 Bingcan Guo , Eryue Xu , Zhiping Zhang , Tianshi Li

Machine learning (ML) and artificial intelligence (AI) approaches are often criticized for their inherent bias and for their lack of control, accountability, and transparency. Consequently, regulatory bodies struggle with containing this…

Artificial Intelligence · Computer Science 2025-01-06 Benjamin Roth , Pedro Henrique Luz de Araujo , Yuxi Xia , Saskia Kaltenbrunner , Christoph Korab

This manuscript establishes information-theoretic limitations for robustness of AI security and alignment by extending G\"odel's incompleteness theorem to AI. Knowing these limitations and preparing for the challenges they bring is…

Artificial Intelligence · Computer Science 2026-05-18 Apostol Vassilev

Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may…

Computation and Language · Computer Science 2026-02-23 Cameron Tice , Puria Radmard , Samuel Ratnam , Andy Kim , David Africa , Kyle O'Brien

Motivated by mitigating potentially harmful impacts of technologies, the AI community has formulated and accepted mathematical definitions for certain pillars of accountability: e.g. privacy, fairness, and model transparency. Yet, we argue…

Machine Learning · Computer Science 2022-12-16 Teresa Datta , Daniel Nissani , Max Cembalest , Akash Khanna , Haley Massa , John P. Dickerson