中文
相关论文

相关论文: Evaluating the role of `Constitutions' for learnin…

200 篇论文

Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative,…

As language models continue to grow larger, the cost of acquiring high-quality training data has increased significantly. Collecting human feedback is both expensive and time-consuming, and manual labels can be noisy, leading to an…

人工智能 · 计算机科学 2025-04-08 Xue Zhang

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of principles (the…

机器学习 · 计算机科学 2026-01-27 Henry Bell , Lara Neubauer da Costa Schertel , Bochu Ding , Brandon Fain

Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (\textit{Crafting Constitutions for CAI…

人工智能 · 计算机科学 2025-02-25 Yara Kyrychenko , Ke Zhou , Edyta Bogucka , Daniele Quercia

Traditional methods for aligning Large Language Models (LLMs), such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on implicit principles, limiting interpretability. Constitutional AI…

机器学习 · 计算机科学 2025-04-01 Carl-Leander Henneking , Claas Beger

There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To…

人工智能 · 计算机科学 2024-06-13 Saffron Huang , Divya Siddarth , Liane Lovitt , Thomas I. Liao , Esin Durmus , Alex Tamkin , Deep Ganguli

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preferences are used to train…

计算与语言 · 计算机科学 2025-04-22 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Samuel Albanie , Robert Mullins

With the rapid development of large language models (LLMs), aligning LLMs with human values and societal norms to ensure their reliability and safety has become crucial. Reinforcement learning with human feedback (RLHF) and Constitutional…

计算与语言 · 计算机科学 2024-03-28 Xiusi Chen , Hongzhi Wen , Sreyashi Nag , Chen Luo , Qingyu Yin , Ruirui Li , Zheng Li , Wei Wang

Constitutional AI is a method to oversee and control LLMs based on a set of rules written in natural language. These rules are typically written by human experts, but could in principle be learned automatically given sufficient training…

人工智能 · 计算机科学 2026-03-18 Rushil Thareja , Gautam Gupta , Francesco Pinto , Nils Lukas

Large language model (LLM) prompting is a promising new approach for users to create and customize their own chatbots. However, current methods for steering a chatbot's outputs, such as prompt engineering and fine-tuning, do not support…

The importance of managing feedback practices in higher education has been widely recognised, as they play a crucial role in enhancing teaching, learning, and assessment processes. In today's educational landscape, feedback practices are…

计算机与社会 · 计算机科学 2026-02-04 Daniele Agostini , Federica Picasso

Recent works have shown that Large Language Models (LLMs) have a tendency to memorize patterns and biases present in their training data, raising important questions about how such memorized content influences model behavior. One such…

计算与语言 · 计算机科学 2025-07-01 Shanshan Xu , T. Y. S. S Santosh , Yanai Elazar , Quirin Vogel , Barbara Plank , Matthias Grabmair

Large language models (LLMs) are increasingly used as conversational partners for learning, yet the interactional dynamics supporting users' learning and engagement are understudied. We analyze the linguistic and interactional features from…

计算与语言 · 计算机科学 2026-03-13 Shaz Furniturewala , Gerard Christopher Yeo , Kokil Jaidka

In aligning large language models (LLMs), utilizing feedback from existing advanced AI rather than humans is an important method to scale supervisory signals. However, it is highly challenging for AI to understand human intentions and…

计算与语言 · 计算机科学 2024-06-18 Rong Bao , Rui Zheng , Shihan Dou , Xiao Wang , Enyu Zhou , Bo Wang , Qi Zhang , Liang Ding , Dacheng Tao

Large language models (LLMs) are used by over a billion people globally, most often to assist with writing. In this work, we demonstrate that LLMs not only alter the voice and tone of human writing, but also consistently alter the intended…

计算与语言 · 计算机科学 2026-03-20 Marwa Abdulhai , Isadora White , Yanming Wan , Ibrahim Qureshi , Joel Leibo , Max Kleiman-Weiner , Natasha Jaques

In this paper, we conduct an empirical analysis of how large language models (LLMs), specifically GPT-4, interpret constitutional principles in complex decision-making scenarios. We examine rulings from the Italian Constitutional Court on…

计算与语言 · 计算机科学 2024-08-12 Camilla Bignotti , Carolina Camassa

Large language models (LLMs) are highly capable at a variety of tasks given the right prompt, but writing one is still a difficult and tedious process. In this work, we introduce ConstitutionalExperts, a method for learning a prompt…

计算与语言 · 计算机科学 2024-03-11 Savvas Petridis , Ben Wedin , Ann Yuan , James Wexler , Nithum Thain

Large language models increasingly function as artificial reasoners: they evaluate arguments, assign credibility, and express confidence. Yet their belief-forming behavior is governed by implicit, uninspected epistemic policies. This paper…

人工智能 · 计算机科学 2026-04-23 Michele Loi

Traditional methods for eliciting people's opinions face a trade-off between depth and scale: structured surveys enable large-scale data collection but limit respondents' ability to voice their opinions in their own words, while…

‹ 上一页 1 2 3 10 下一页 ›