中文
相关论文

相关论文: How Well Do Models Follow Their Constitutions?

200 篇论文

This work focuses on the problem of hyper-parameter tuning (HPT) for robust (i.e., adversarially trained) models, shedding light on the new challenges and opportunities arising during the HPT process for robust models. To this end, we…

机器学习 · 计算机科学 2024-06-14 Pedro Mendes , Paolo Romano , David Garlan

As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Current benchmarks for AI agents primarily evaluate refusal of…

人工智能 · 计算机科学 2026-05-12 Miles Q. Li , Benjamin C. M. Fung , Martin Weiss , Pulei Xiong , Khalil Al-Hussaeni , Claude Fachkha

Large language models increasingly function as artificial reasoners: they evaluate arguments, assign credibility, and express confidence. Yet their belief-forming behavior is governed by implicit, uninspected epistemic policies. This paper…

人工智能 · 计算机科学 2026-04-23 Michele Loi

Constitutional AI has focused on single-model alignment using fixed principles. However, multi-agent systems create novel alignment challenges through emergent social dynamics. We present Constitutional Evolution, a framework for…

多智能体系统 · 计算机科学 2026-02-04 Ujwal Kumar , Alice Saito , Hershraj Niranjani , Rayan Yessou , Phan Xuan Tan

We introduce a black-box interpretability framework that learns a verifiable constitution: a natural language summary of how changes to a prompt affect a model's specific behavior, such as its alignment, correctness, or adherence to…

机器学习 · 计算机科学 2026-02-03 Neha Kalibhat , Zi Wang , Prasoon Bajpai , Drew Proud , Wenjun Zeng , Been Kim , Mani Malek

The same prompt -- "best CRM software" -- reaches AI assistants from buyers in widely different contexts: a solo founder, an enterprise VP, a UK SMB owner. We audit how strongly that contextual variation reshapes which brands the model…

人工智能 · 计算机科学 2026-05-29 Will Jack , Noah Lehman , Keller Maloney , Sarah Xu

The growing capabilities of large language models (LLMs) have led to their use as substitutes for human feedback for training and assessing other LLMs. These methods often rely on `constitutions', written guidelines which a critic model…

人工智能 · 计算机科学 2024-11-18 Saskia Redgate , Andrew M. Bean , Adam Mahdi

The enterprise governance of Generative AI (GenAI) in regulated sectors, such as Human Resources (HR), demands scalable yet reproducible auditing mechanisms. While Large Language Model (LLM)-as-a-Judge approaches offer scalability, their…

软件工程 · 计算机科学 2026-01-21 Murtuza N. Shergadwala

Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. Simple category definitions are not detailed enough for labelers to produce the accurate,…

计算与语言 · 计算机科学 2026-05-26 Konstantin Berlin , Adam Swanda

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of principles (the…

机器学习 · 计算机科学 2026-01-27 Henry Bell , Lara Neubauer da Costa Schertel , Bochu Ding , Brandon Fain

The dominant industry response to AI-generated code quality problems is to deploy AI reviewers. This paper argues that this response is structurally circular when executable specifications are absent: without an external reference, both the…

软件工程 · 计算机科学 2026-03-30 Christo Zietsman

Predicting agents impacted by legal policies, physical limitations, and operational preferences is inherently difficult. In recent years, neuro-symbolic methods have emerged, integrating machine learning and symbolic reasoning models into…

机器人学 · 计算机科学 2025-07-22 Simon Kohaut , Felix Divo , Benedict Flade , Devendra Singh Dhami , Julian Eggert , Kristian Kersting

Autonomous AI agents are being deployed with filesystem access, email control, and multi-step planning. This thesis contributes to four open problems in AI safety: understanding dangerous internal computations, removing dangerous behaviors…

机器学习 · 计算机科学 2026-04-02 Aengus Lynch

Generative AI is rapidly moving from research to deployment, elevating the need for responsible development, evaluation, and governance. We conduct a PRISMA guided review of 232 studies (November 2022 - December 2025), spanning large…

Artificial Intelligence (AI) is taking on increasingly autonomous roles, e.g., browsing the web as a research assistant and managing money. But specifying goals and restrictions for AI behavior is difficult. Similar to how parties to a…

计算与语言 · 计算机科学 2023-01-31 John J. Nay

Real-world natural language processing systems need to be robust to human adversaries. Collecting examples of human adversaries for training is an effective but expensive solution. On the other hand, training on synthetic attacks with small…

机器学习 · 计算机科学 2024-02-16 Aradhana Sinha , Ananth Balashankar , Ahmad Beirami , Thi Avrahami , Jilin Chen , Alex Beutel

As AI systems become increasingly prevalent and impactful, the need for effective AI governance and accountability measures is paramount. This paper examines the AI governance landscape, focusing on Anthropic's Claude, a foundational AI…

计算机与社会 · 计算机科学 2024-07-03 Aman Priyanshu , Yash Maurya , Zuofei Hong

Alignment research focuses on making individual AI systems reliable. Human institutions achieve reliable collective behaviour differently: they mitigate the risk posed by misaligned individuals through organisational structure. Multi-agent…

人工智能 · 计算机科学 2026-02-17 William Waites

There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To…

人工智能 · 计算机科学 2024-06-13 Saffron Huang , Divya Siddarth , Liane Lovitt , Thomas I. Liao , Esin Durmus , Alex Tamkin , Deep Ganguli