English
Related papers

Related papers: Specific versus General Principles for Constitutio…

200 papers

Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's Model Spec (OpenAI, 2025a), integrated into post-training via methods like character…

Artificial Intelligence · Computer Science 2026-05-26 Arya Jakkli , Senthooran Rajamanoharan , Neel Nanda

Companies have considered adoption of various high-level artificial intelligence (AI) principles for responsible AI, but there is less clarity on how to implement these principles as organizational practices. This paper reviews the…

Computers and Society · Computer Science 2020-06-09 Daniel Schiff , Bogdana Rakova , Aladdin Ayesh , Anat Fanti , Michael Lennon

The rapid adoption of generative artificial intelligence (AI) in scientific research, particularly large language models (LLMs), has outpaced the development of ethical guidelines, leading to a "Triple-Too" problem: too many high-level…

Computers and Society · Computer Science 2026-04-07 Zhicheng Lin

Anthropomorphic language describing artificial intelligence (AI) is widespread in media, policy, and everyday discourse; so too are discussions of AI bad behavior, from hallucinations to inappropriate comments. How does humanizing language…

Human-Computer Interaction · Computer Science 2026-04-29 Jaime Banks , Nicholas David Bowman , Roman Saladino

Using the example of the film 2001: A Space Odyssey, this chapter illustrates the challenges posed by an AI capable of making decisions that go against human interests. But are human decisions always rational and ethical? In reality, the…

Computers and Society · Computer Science 2025-12-05 Charlotte Jacquemot

Agentic AI systems, possessing capabilities for autonomous planning and action, show great potential across diverse domains. However, their practical deployment is hindered by challenges in aligning their behavior with varied human values,…

Artificial Intelligence · Computer Science 2025-08-12 Nell Watson , Ahmed Amer , Evan Harris , Preeti Ravindra , Shujun Zhang

We propose to build directly upon our longstanding, prior r&d in AI/machine ethics in order to attempt to make real the blue-sky idea of AI that can thwart mass shootings, by bringing to bear its ethical reasoning. The r&d in question is…

Computers and Society · Computer Science 2021-02-19 Selmer Bringsjord , Naveen Sundar Govindarajulu , Michael Giancola

For an artificial intelligence (AI) to be aligned with human values (or human preferences), it must first learn those values. AI systems that are trained on human behavior, risk miscategorising human irrationalities as human values -- and…

Artificial Intelligence · Computer Science 2022-03-02 Rebecca Gorman , Stuart Armstrong

Amidst the race to create more intelligent machines there is a risk that we will rely on AI in ways that reduce our own agency as humans. To reduce this risk, we could aim to create tools that prioritize and enhance the human role in…

Human-Computer Interaction · Computer Science 2026-01-15 Sean Koon

Constitutional AI has focused on single-model alignment using fixed principles. However, multi-agent systems create novel alignment challenges through emergent social dynamics. We present Constitutional Evolution, a framework for…

Multiagent Systems · Computer Science 2026-02-04 Ujwal Kumar , Alice Saito , Hershraj Niranjani , Rayan Yessou , Phan Xuan Tan

Advanced AI systems capable of generating humanlike text and multimodal content are now widely available. In this paper, we discuss the impacts that generative artificial intelligence may have on democratic processes. We consider the…

The introduction of artificial intelligence into activities traditionally carried out by human beings produces brutal changes. This is not without consequences for human values. This paper is about designing and implementing models of…

Artificial Intelligence · Computer Science 2020-10-16 Fabrice Muhlenbach

The prevailing discourse around AI ethics lacks the language and formalism necessary to capture the diverse ethical concerns that emerge when AI systems interact with individuals. Drawing on Sen and Nussbaum's capability approach, we…

Artificial Intelligence · Computer Science 2023-09-08 Alex John London , Hoda Heidari

This article presents a critique of ethics in the context of artificial intelligence (AI). It argues for the need to question established patterns of thought and traditional authorities, including core concepts such as autonomy, morality,…

Computers and Society · Computer Science 2024-08-09 Irina Spiegel

We fine-tune large language models to write natural language critiques (natural language critical comments) using behavioral cloning. On a topic-based summarization task, critiques written by our models help humans find flaws in summaries…

Computation and Language · Computer Science 2022-06-15 William Saunders , Catherine Yeh , Jeff Wu , Steven Bills , Long Ouyang , Jonathan Ward , Jan Leike

Given that Artificial Intelligence (AI) increasingly permeates our lives, it is critical that we systematically align AI objectives with the goals and values of humans. The human-AI alignment problem stems from the impracticality of…

Computers and Society · Computer Science 2022-07-05 John Nay , James Daily

High-quality feedback is essential for effective human-AI interaction. It bridges knowledge gaps, corrects digressions, and shapes system behavior; both during interaction and throughout model development. Yet despite its importance, human…

Human-Computer Interaction · Computer Science 2026-03-31 Nikhil Sharma , Zheng Zhang , Daniel Lee , Namita Krishnan , Guang-Jie Ren , Ziang Xiao , Yunyao Li

We introduce a black-box interpretability framework that learns a verifiable constitution: a natural language summary of how changes to a prompt affect a model's specific behavior, such as its alignment, correctness, or adherence to…

Machine Learning · Computer Science 2026-02-03 Neha Kalibhat , Zi Wang , Prasoon Bajpai , Drew Proud , Wenjun Zeng , Been Kim , Mani Malek

We are currently unable to specify human goals and societal values in a way that reliably directs AI behavior. Law-making and legal interpretation form a computational engine that converts opaque human values into legible directives. "Law…

Computers and Society · Computer Science 2023-05-17 John J. Nay

Computational social choice and algorithmic decision theory offer rich aggregation theory but no comprehensive process for egalitarian self-governance: aggregation, deliberation, amendment, and consensus are each considered in isolation,…

Multiagent Systems · Computer Science 2026-05-15 Ehud Shapiro , Nimrod Talmon