中文
相关论文

相关论文: How Well Do Models Follow Their Constitutions?

200 篇论文

As language models continue to grow larger, the cost of acquiring high-quality training data has increased significantly. Collecting human feedback is both expensive and time-consuming, and manual labels can be noisy, leading to an…

人工智能 · 计算机科学 2025-04-08 Xue Zhang

In January 2026, Anthropic published a 79-page "constitution" for its AI model Claude, the most comprehensive corporate AI governance document ever released. This Article offers the first legal and democratic-theoretic analysis of that…

计算机与社会 · 计算机科学 2026-04-06 Gilad Abiri

As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability (knowing what they do not know, detecting errors, seeking clarification) under adversarial pressure is a critical safety…

人工智能 · 计算机科学 2026-05-15 Rahul Kumar

Companies that develop foundation models publish behavioral guidelines they pledge their models will follow, but it remains unclear if models actually do so. While providers such as OpenAI, Anthropic, and Google have published detailed…

计算与语言 · 计算机科学 2025-10-24 Ahmed Ahmed , Kevin Klyman , Yi Zeng , Sanmi Koyejo , Percy Liang

The character of the "AI assistant" persona generated by modern chatbot large language models influences both surface-level behavior and apparent values, beliefs, and ethics. These all affect interaction quality, perceived intelligence, and…

计算与语言 · 计算机科学 2025-11-04 Sharan Maiya , Henning Bartsch , Nathan Lambert , Evan Hubinger

Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (\textit{Crafting Constitutions for CAI…

人工智能 · 计算机科学 2025-02-25 Yara Kyrychenko , Ke Zhou , Edyta Bogucka , Daniele Quercia

We are increasingly subjected to the power of AI authorities. As AI decisions become inescapable, entering domains such as healthcare, education, and law, we must confront a vital question: how can we ensure AI systems have the legitimacy…

计算机与社会 · 计算机科学 2025-05-15 Gilad Abiri

Constitutional AI is a method to oversee and control LLMs based on a set of rules written in natural language. These rules are typically written by human experts, but could in principle be learned automatically given sufficient training…

人工智能 · 计算机科学 2026-03-18 Rushil Thareja , Gautam Gupta , Francesco Pinto , Nils Lukas

Agentic AI systems, possessing capabilities for autonomous planning and action, show great potential across diverse domains. However, their practical deployment is hindered by challenges in aligning their behavior with varied human values,…

人工智能 · 计算机科学 2025-08-12 Nell Watson , Ahmed Amer , Evan Harris , Preeti Ravindra , Shujun Zhang

Constitutional AI (CAI) aligns language models with explicitly stated normative principles, offering a transparent alternative to implicit alignment through human feedback alone. However, because constitutions are authored by specific…

计算机与社会 · 计算机科学 2026-03-31 Parham Pourdavood

Fine-tuning APIs offered by major AI providers create new attack surfaces where adversaries can bypass safety measures through targeted fine-tuning. We introduce Trojan-Speak, an adversarial fine-tuning method that bypasses Anthropic's…

密码学与安全 · 计算机科学 2026-04-01 Bilgehan Sel , Xuanli He , Alwin Peng , Ming Jin , Jerry Wei

Technical and legal debates frequently suggest that "accuracy" is an objective, measurable, and purely technical property. We challenge this view, showing that evaluating AI performance fundamentally depends on context-dependent normative…

AI agents powered by reasoning models require access to sensitive user data. However, their reasoning traces are difficult to control, which can result in the unintended leakage of private information to external parties. We propose…

计算与语言 · 计算机科学 2026-03-02 Haritz Puerto , Haonan Li , Xudong Han , Timothy Baldwin , Iryna Gurevych

Privacy concerns have led to the development of privacy-preserving approaches for learning models from sensitive data. Yet, in practice, even models learned with privacy guarantees can inadvertently memorize unique training examples or leak…

机器学习 · 统计学 2019-11-11 Mario Diaz , Peter Kairouz , Jiachun Liao , Lalitha Sankar

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preferences are used to train…

计算与语言 · 计算机科学 2025-04-22 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Samuel Albanie , Robert Mullins

Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative,…

Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tuning, called reinforcement learning from human feedback,…

Leading language model (LM) providers like OpenAI and Anthropic allow customers to fine-tune frontier LMs for specific use cases. To prevent abuse, these providers apply filters to block fine-tuning on overtly harmful data. In this setting,…

密码学与安全 · 计算机科学 2025-07-15 Joshua Kazdan , Abhay Puri , Rylan Schaeffer , Lisa Yu , Chris Cundy , Jason Stanley , Sanmi Koyejo , Krishnamurthy Dvijotham

The rapid emergence of large language models (LLMs) has raised urgent questions across the modern workforce about this new technology's strengths, weaknesses, and capabilities. For privacy professionals, the question is whether these AI…

计算机与社会 · 计算机科学 2025-08-13 Zane Witherspoon , Thet Mon Aye , YingYing Hao

Existing legal frameworks on AI rely on training compute thresholds as a proxy to identify potentially-dangerous AI models and trigger increased regulatory attention. In the United States, Section 4.2(a) of Executive Order 14110 instructs…

计算机与社会 · 计算机科学 2025-02-04 Matteo Pistillo , Pablo Villalobos
‹ 上一页 1 2 3 10 下一页 ›