中文
相关论文

相关论文: Value alignment: a formal approach

200 篇论文

The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH). At the same time, there are mounting concerns that LLMs exhibit a left-wing political bias. Yet, the commitment to AI…

计算与语言 · 计算机科学 2025-07-22 Thilo Hagendorff

The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussions on Artificial Superintelligence (ASI)-a system surpassing…

人工智能 · 计算机科学 2026-02-10 HyunJin Kim , Xiaoyuan Yi , Jing Yao , Muhua Huang , JinYeong Bak , James Evans , Xing Xie

We develop an optimization framework centered around a core idea: once a (parametric) policy is specified, control authority is transferred to the policy, resulting in an autonomous dynamical system. Thus we should be able to optimize…

机器学习 · 计算机科学 2025-06-11 Emo Todorov

In an effort to regulate Machine Learning-driven (ML) systems, current auditing processes mostly focus on detecting harmful algorithmic biases. While these strategies have proven to be impactful, some values outlined in documents dealing…

机器学习 · 计算机科学 2022-06-20 Mireia Yurrita , Dave Murray-Rust , Agathe Balayn , Alessandro Bozzon

Large language models implicitly encode preferences over human values, yet steering them often requires large training data. In this work, we investigate a simple approach: Can we reliably modify a model's value system in downstream…

计算与语言 · 计算机科学 2025-08-18 Shangrui Nie , Florian Mai , David Kaczér , Charles Welch , Zhixue Zhao , Lucie Flek

Technical and legal debates frequently suggest that "accuracy" is an objective, measurable, and purely technical property. We challenge this view, showing that evaluating AI performance fundamentally depends on context-dependent normative…

Value learning is a crucial aspect of safe and ethical AI. This is primarily pursued by methods inferring human values from behaviour. However, humans care about much more than we are able to demonstrate through our actions. Consequently,…

人工智能 · 计算机科学 2025-05-28 Paul de Font-Reaulx

Large language models (LLMs) are increasingly being used as decision aids. However, users have diverse values and preferences that can affect their decision-making, which requires novel methods for LLM alignment and personalization.…

Big models have greatly advanced AI's ability to understand, generate, and manipulate information and content, enabling numerous applications. However, as these models become increasingly integrated into everyday life, their inherent…

计算机与社会 · 计算机科学 2023-10-27 Xiaoyuan Yi , Jing Yao , Xiting Wang , Xing Xie

Agents based on Large Language Models (LLMs) are increasingly permeating various domains of human production and life, highlighting the importance of aligning them with human values. The current alignment of AI systems primarily focuses on…

计算与语言 · 计算机科学 2024-02-21 Shimin Li , Tianxiang Sun , Qinyuan Cheng , Xipeng Qiu

As Artificial Intelligence (AI) technologies continue to advance, protecting human autonomy and promoting ethical decision-making are essential to fostering trust and accountability. Human agency (the capacity of individuals to make…

计算机与社会 · 计算机科学 2025-10-13 Laxmiraju Kandikatla , Branislav Radeljic

In AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are `better' or `safer.' We refer to this latitude as alignment discretion. Such discretion remains largely…

Autoformalization aims to convert informal mathematical proofs into machine-verifiable formats, bridging the gap between natural and formal languages. However, ensuring semantic alignment between the informal and formalized statements…

计算与语言 · 计算机科学 2024-10-15 Jianqiao Lu , Yingjia Wan , Yinya Huang , Jing Xiong , Zhengying Liu , Zhijiang Guo

AI ethics is an emerging field with multiple, competing narratives about how to best solve the problem of building human values into machines. Two major approaches are focused on bias and compliance, respectively. But neither of these ideas…

人工智能 · 计算机科学 2023-02-24 Thomas Krendl Gilbert , Megan Welle Brozek , Andrew Brozek

Agentic AIs $-$ AIs that are capable and permitted to undertake complex actions with little supervision $-$ mark a new frontier in AI capabilities and raise new questions about how to safely create and align such systems with users,…

计算机与社会 · 计算机科学 2024-10-04 Hayley Clatterbuck , Clinton Castro , Arvo Muñoz Morán

In this paper, we aim to align large language models with the ever-changing, complex, and diverse human values (e.g., social norms) across time and locations. This presents a challenge to existing alignment techniques, such as supervised…

计算与语言 · 计算机科学 2023-12-27 Chunpu Xu , Steffi Chern , Ethan Chern , Ge Zhang , Zekun Wang , Ruibo Liu , Jing Li , Jie Fu , Pengfei Liu

Big models, exemplified by Large Language Models (LLMs), are models typically pre-trained on massive data and comprised of enormous parameters, which not only obtain significantly improved performance across diverse tasks but also present…

人工智能 · 计算机科学 2023-09-06 Jing Yao , Xiaoyuan Yi , Xiting Wang , Jindong Wang , Xing Xie

Deep neural networks excel in medical imaging but remain prone to biases, leading to fairness gaps across demographic groups. We provide the first systematic exploration of Human-AI alignment and fairness in this domain. Our results show…

计算机视觉与模式识别 · 计算机科学 2025-05-16 Haozhe Luo , Ziyu Zhou , Zixin Shu , Aurélie Pahud de Mortanges , Robert Berke , Mauricio Reyes

Algorithms and other formal models purportedly incorporating human values like fairness have grown increasingly popular in computer science. In response to sociotechnical challenges in the use of these models, designers and researchers have…

计算机与社会 · 计算机科学 2021-06-18 Benjamin Fish , Luke Stark

Generative AI models ought to be useful and safe across cross-cultural contexts. One critical step toward this goal is understanding how AI models adhere to sociocultural norms. While this challenge has gained attention in NLP, existing…

‹ 上一页 1 8 9 10 下一页 ›