English
Related papers

Related papers: Does Claude's Constitution Have a Culture?

200 papers

There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To…

Artificial Intelligence · Computer Science 2024-06-13 Saffron Huang , Divya Siddarth , Liane Lovitt , Thomas I. Liao , Esin Durmus , Alex Tamkin , Deep Ganguli

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of principles (the…

Machine Learning · Computer Science 2026-01-27 Henry Bell , Lara Neubauer da Costa Schertel , Bochu Ding , Brandon Fain

AI assistants can impart value judgments that shape people's decisions and worldviews, yet little is known empirically about what values these systems rely on in practice. To address this, we develop a bottom-up, privacy-preserving method…

Computation and Language · Computer Science 2025-04-22 Saffron Huang , Esin Durmus , Miles McCain , Kunal Handa , Alex Tamkin , Jerry Hong , Michael Stern , Arushi Somani , Xiuruo Zhang , Deep Ganguli

Traditional methods for aligning Large Language Models (LLMs), such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on implicit principles, limiting interpretability. Constitutional AI…

Machine Learning · Computer Science 2025-04-01 Carl-Leander Henneking , Claas Beger

Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (\textit{Crafting Constitutions for CAI…

Artificial Intelligence · Computer Science 2025-02-25 Yara Kyrychenko , Ke Zhou , Edyta Bogucka , Daniele Quercia

In January 2026, Anthropic published a 79-page "constitution" for its AI model Claude, the most comprehensive corporate AI governance document ever released. This Article offers the first legal and democratic-theoretic analysis of that…

Computers and Society · Computer Science 2026-04-06 Gilad Abiri

The growing capabilities of large language models (LLMs) have led to their use as substitutes for human feedback for training and assessing other LLMs. These methods often rely on `constitutions', written guidelines which a critic model…

Artificial Intelligence · Computer Science 2024-11-18 Saskia Redgate , Andrew M. Bean , Adam Mahdi

We are increasingly subjected to the power of AI authorities. As AI decisions become inescapable, entering domains such as healthcare, education, and law, we must confront a vital question: how can we ensure AI systems have the legitimacy…

Computers and Society · Computer Science 2025-05-15 Gilad Abiri

Culture fundamentally shapes people's reasoning, behavior, and communication. As people increasingly use generative artificial intelligence (AI) to expedite and automate personal and professional tasks, cultural values embedded in AI models…

Computation and Language · Computer Science 2024-09-20 Yan Tao , Olga Viberg , Ryan S. Baker , Rene F. Kizilcec

As AI systems become increasingly prevalent and impactful, the need for effective AI governance and accountability measures is paramount. This paper examines the AI governance landscape, focusing on Anthropic's Claude, a foundational AI…

Computers and Society · Computer Science 2024-07-03 Aman Priyanshu , Yash Maurya , Zuofei Hong

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preferences are used to train…

Computation and Language · Computer Science 2025-04-22 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Samuel Albanie , Robert Mullins

As language models continue to grow larger, the cost of acquiring high-quality training data has increased significantly. Collecting human feedback is both expensive and time-consuming, and manual labels can be noisy, leading to an…

Artificial Intelligence · Computer Science 2025-04-08 Xue Zhang

Large language models have become the latest trend in natural language processing, heavily featuring in the digital tools we use every day. However, their replies often reflect a narrow cultural viewpoint that overlooks the diversity of…

Computation and Language · Computer Science 2025-10-22 Alistair Plum , Anne-Marie Lutgen , Christoph Purschke , Achim Rettinger

Large language models increasingly function as artificial reasoners: they evaluate arguments, assign credibility, and express confidence. Yet their belief-forming behavior is governed by implicit, uninspected epistemic policies. This paper…

Artificial Intelligence · Computer Science 2026-04-23 Michele Loi

When you ask an AI assistant for advice about your career, your marriage, or a conflict with your family, does it give you the same answer regardless of where you are from? We tested this systematically by presenting three leading AI…

Computation and Language · Computer Science 2026-04-27 Pruthvinath Jeripity Venkata

Are AI systems truly representing human values, or merely averaging across them? Our study suggests a concerning reality: Large Language Models (LLMs) fail to represent diverse cultural moral frameworks despite their linguistic…

Computation and Language · Computer Science 2025-08-01 Simon Münker

Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative,…

There is an urgent need to incorporate the perspectives of culturally diverse groups into AI developments. We present a novel conceptual framework for research that aims to expand, reimagine, and reground mainstream visions of AI using…

Human-Computer Interaction · Computer Science 2024-03-11 Xiao Ge , Chunchen Xu , Daigo Misaki , Hazel Rose Markus , Jeanne L Tsai

Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's Model Spec (OpenAI, 2025a), integrated into post-training via methods like character…

Artificial Intelligence · Computer Science 2026-05-26 Arya Jakkli , Senthooran Rajamanoharan , Neel Nanda

Objective: This study examines how well leading Chinese and Western large language models understand and apply Chinese social work principles, focusing on their foundational knowledge within a non-Western professional setting. We test…

Computers and Society · Computer Science 2025-03-10 Zia Qi , Brian E. Perron , Miao Wang , Cao Fang , Sitao Chen , Bryan G. Victor
‹ Prev 1 2 3 10 Next ›