English
Related papers

Related papers: Researching Alignment Research: Unsupervised Analy…

200 papers

The success of artificial intelligence (AI), and deep learning models in particular, has led to their widespread adoption across various industries due to their ability to process huge amounts of data and learn complex patterns. However,…

Artificial Intelligence · Computer Science 2023-09-22 Wei Jie Yeo , Wihan van der Heever , Rui Mao , Erik Cambria , Ranjan Satapathy , Gianmarco Mengaldo

We describe cases where real recommender systems were modified in the service of various human values such as diversity, fairness, well-being, time well spent, and factual accuracy. From this we identify the current practice of values…

Information Retrieval · Computer Science 2021-07-26 Jonathan Stray , Ivan Vendrov , Jeremy Nixon , Steven Adler , Dylan Hadfield-Menell

It is widely agreed that when AI models assist decision-makers in high-stakes domains by predicting an outcome of interest, they should communicate the confidence of their predictions. However, empirical evidence suggests that…

Machine Learning · Computer Science 2026-05-14 Nina Corvelo Benz , Eleni Straitouri , Manuel Gomez-Rodriguez

Artificial intelligence (AI) is rapidly advancing in healthcare, enhancing the efficiency and effectiveness of services across various specialties, including cardiology, ophthalmology, dermatology, emergency medicine, etc. AI applications…

The psychological science of artificial intelligence (AI) can be broadly defined as an emerging field of psychology that examines all AI-related mental and behavioral processes from the perspective of psychology. This field has been growing…

Human-Computer Interaction · Computer Science 2026-02-03 Zheng Yan , Ru-Yuan Zhang

We formalize AI alignment as a multi-objective optimization problem called $\langle M,N,\varepsilon,\delta\rangle$-agreement, in which a set of $N$ agents (including humans) must reach approximate ($\varepsilon$) agreement across $M$…

Artificial Intelligence · Computer Science 2025-11-20 Aran Nayebi

This essay examines how what is considered to be artificial intelligence (AI) has changed over time and come to intersect with the expertise of the author. Initially, AI developed on a separate trajectory, both topically and…

Artificial Intelligence · Computer Science 2018-04-02 Kush R. Varshney

This position paper argues that formal optimal control theory should be central to AI alignment research, offering a distinct perspective from prevailing AI safety and security approaches. While recent work in AI safety and mechanistic…

Artificial Intelligence · Computer Science 2025-06-24 Elija Perrier

Over the past decade, AI has made a remarkable progress. It is agreed that this is due to the recently revived Deep Learning technology. Deep Learning enables to process large amounts of data using simplified neuron networks that simulate…

Artificial Intelligence · Computer Science 2015-02-19 Emanuel Diamant

Artificial Intelligence is a field that lives many lives, and the term has come to encompass a motley collection of scientific and commercial endeavours. In this paper, I articulate the contours of a rather neglected but central scientific…

Artificial Intelligence · Computer Science 2024-08-19 Dimitri Coelho Mollo

Mathematical reasoning is a fundamental aspect of human intelligence and is applicable in various fields, including science, engineering, finance, and everyday life. The development of artificial intelligence (AI) systems capable of solving…

Artificial Intelligence · Computer Science 2023-06-23 Pan Lu , Liang Qiu , Wenhao Yu , Sean Welleck , Kai-Wei Chang

Inner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debated. Moreover, recent critiques raise issues that question…

Artificial Intelligence · Computer Science 2024-08-01 Martina G. Vilas , Federico Adolfi , David Poeppel , Gemma Roig

The growing importance of understanding and addressing algorithmic bias in artificial intelligence (AI) has led to a surge in research on AI fairness, which often assumes that the underlying data is independent and identically distributed…

Machine Learning · Computer Science 2025-01-28 Wenbin Zhang , Shuigeng Zhou , Toby Walsh , Jeremy C. Weiss

Today, AI is increasingly being used in many high-stakes decision-making applications in which fairness is an important concern. Already, there are many examples of AI being biased and making questionable and unfair decisions. The AI…

Artificial Intelligence · Computer Science 2020-02-06 Yunfeng Zhang , Rachel K. E. Bellamy , Kush R. Varshney

Neuroscience and AI have an intertwined history, largely relayed in the literature of both fields. In recent years, due to the engineering orientations of AI research and the monopoly of industry for its large-scale applications, the mutual…

Digital Libraries · Computer Science 2025-11-13 Sylvain Fontaine

Aligning AI systems with human values fundamentally relies on effective human feedback. While significant research has addressed training algorithms, the role of user interface is often overlooked and only treated as an implementation…

Human-Computer Interaction · Computer Science 2026-02-13 Danqing Shi

"AI as a Service" (AIaaS) is a rapidly growing market, offering various plug-and-play AI services and tools. AIaaS enables its customers (users) - who may lack the expertise, data, and/or resources to develop their own systems - to easily…

Machine Learning · Computer Science 2023-02-06 Kornel Lewicki , Michelle Seng Ah Lee , Jennifer Cobbe , Jatinder Singh

After the tremendous advances of deep learning and other AI methods, more attention is flowing into other properties of modern approaches, such as interpretability, fairness, etc. combined in frameworks like Responsible AI. Two research…

Artificial Intelligence · Computer Science 2021-05-26 Dominik Seuß

Modern AI enables a high-level, declarative form of interaction: Users describe the intended outcome they wish an AI to produce, but do not actually create the outcome themselves. In contrast, in traditional user interfaces, users invoke…

Human-Computer Interaction · Computer Science 2024-09-18 Michael Terry , Chinmay Kulkarni , Martin Wattenberg , Lucas Dixon , Meredith Ringel Morris

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally…

Computers and Society · Computer Science 2025-06-05 Max Hellrigel-Holderbaum , Leonard Dung