中文
相关论文

相关论文: C3AI: Crafting and Evaluating Constitutions for Co…

200 篇论文

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of principles (the…

机器学习 · 计算机科学 2026-01-27 Henry Bell , Lara Neubauer da Costa Schertel , Bochu Ding , Brandon Fain

Traditional methods for aligning Large Language Models (LLMs), such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on implicit principles, limiting interpretability. Constitutional AI…

机器学习 · 计算机科学 2025-04-01 Carl-Leander Henneking , Claas Beger

There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To…

人工智能 · 计算机科学 2024-06-13 Saffron Huang , Divya Siddarth , Liane Lovitt , Thomas I. Liao , Esin Durmus , Alex Tamkin , Deep Ganguli

Constitutional AI is a method to oversee and control LLMs based on a set of rules written in natural language. These rules are typically written by human experts, but could in principle be learned automatically given sufficient training…

人工智能 · 计算机科学 2026-03-18 Rushil Thareja , Gautam Gupta , Francesco Pinto , Nils Lukas

The growing capabilities of large language models (LLMs) have led to their use as substitutes for human feedback for training and assessing other LLMs. These methods often rely on `constitutions', written guidelines which a critic model…

人工智能 · 计算机科学 2024-11-18 Saskia Redgate , Andrew M. Bean , Adam Mahdi

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preferences are used to train…

计算与语言 · 计算机科学 2025-04-22 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Samuel Albanie , Robert Mullins

As language models continue to grow larger, the cost of acquiring high-quality training data has increased significantly. Collecting human feedback is both expensive and time-consuming, and manual labels can be noisy, leading to an…

人工智能 · 计算机科学 2025-04-08 Xue Zhang

Recent incidents highlight safety risks in Large Language Models (LLMs), motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-critique mechanism on small, uncensored 7-9B parameter models:…

机器学习 · 计算机科学 2025-04-14 Antonio-Gabriel Chacón Menke , Phan Xuan Tan

Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative,…

We are increasingly subjected to the power of AI authorities. As AI decisions become inescapable, entering domains such as healthcare, education, and law, we must confront a vital question: how can we ensure AI systems have the legitimacy…

计算机与社会 · 计算机科学 2025-05-15 Gilad Abiri

With the rapid development of large language models (LLMs), aligning LLMs with human values and societal norms to ensure their reliability and safety has become crucial. Reinforcement learning with human feedback (RLHF) and Constitutional…

计算与语言 · 计算机科学 2024-03-28 Xiusi Chen , Hongzhi Wen , Sreyashi Nag , Chen Luo , Qingyu Yin , Ruirui Li , Zheng Li , Wei Wang

From moderating content within an online community to producing socially-appropriate generative outputs, decision-making tasks -- conducted by either humans or AI -- often depend on subjective or socially-established criteria. To ensure…

人机交互 · 计算机科学 2025-09-26 Quan Ze Chen , Amy X. Zhang

Predicting agents impacted by legal policies, physical limitations, and operational preferences is inherently difficult. In recent years, neuro-symbolic methods have emerged, integrating machine learning and symbolic reasoning models into…

机器人学 · 计算机科学 2025-07-22 Simon Kohaut , Felix Divo , Benedict Flade , Devendra Singh Dhami , Julian Eggert , Kristian Kersting

Constitutional AI (CAI) aligns language models with explicitly stated normative principles, offering a transparent alternative to implicit alignment through human feedback alone. However, because constitutions are authored by specific…

计算机与社会 · 计算机科学 2026-03-31 Parham Pourdavood

Alignment of artificial intelligence (AI) encompasses the normative problem of specifying how AI systems should act and the technical problem of ensuring AI systems comply with those specifications. To date, AI alignment has generally…

Transformative AI systems may pose unprecedented catastrophic risks, but the U.S. Constitution places significant constraints on the government's ability to govern this technology. This paper examines how the First Amendment, administrative…

计算机与社会 · 计算机科学 2025-12-19 Alex Mark , Aaron Scher

Large language models increasingly function as artificial reasoners: they evaluate arguments, assign credibility, and express confidence. Yet their belief-forming behavior is governed by implicit, uninspected epistemic policies. This paper…

人工智能 · 计算机科学 2026-04-23 Michele Loi

Case studies commonly form the pedagogical backbone in law, ethics, and many other domains that face complex and ambiguous societal questions informed by human values. Similar complexities and ambiguities arise when we consider how AI…

人工智能 · 计算机科学 2023-11-28 K. J. Kevin Feng , Quan Ze Chen , Inyoung Cheong , King Xia , Amy X. Zhang

Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-explored. We propose Reverse Constitutional AI (R-CAI), a framework for automated and…

计算与语言 · 计算机科学 2026-04-21 Yuan Fang , Yiming Luo , Aimin Zhou , Fei Tan
‹ 上一页 1 2 3 10 下一页 ›