中文
相关论文

相关论文: Beyond Preferences: Learning Alignment Principles …

200 篇论文

We are introducing Aligned, a platform for global governance and alignment of frontier models, and eventually superintelligence. While previous efforts at the major AI labs have attempted to gather inputs for alignment, these are often…

计算机与社会 · 计算机科学 2023-11-16 Ethan Shaotran , Ido Pesok , Sam Jones , Emi Liu

Large language models (LLMs), initially developed for generative AI, are now evolving into agentic AI systems, which make decisions in complex, real-world contexts. Unfortunately, while their generative capabilities are well-documented,…

人工智能 · 计算机科学 2026-04-02 Matthew DosSantos DiSorbo , Harang Ju , Sinan Aral

Artificial Intelligence principles define social and ethical considerations to develop future AI. They come from research institutes, government organizations and industries. All versions of AI principles are with different considerations…

人工智能 · 计算机科学 2018-12-13 Yi Zeng , Enmeng Lu , Cunqing Huangfu

Constitutional AI has focused on single-model alignment using fixed principles. However, multi-agent systems create novel alignment challenges through emergent social dynamics. We present Constitutional Evolution, a framework for…

多智能体系统 · 计算机科学 2026-02-04 Ujwal Kumar , Alice Saito , Hershraj Niranjani , Rayan Yessou , Phan Xuan Tan

Large language models (LLMs) have shown promising accuracy in predicting survey responses and policy preferences, which has increased interest in their potential to represent human interests in various domains. Most existing research has…

计算机与社会 · 计算机科学 2025-11-18 Suyash Fulay , Jocelyn Zhu , Michiel Bakker

Due to the remarkable capabilities and growing impact of large language models (LLMs), they have been deeply integrated into many aspects of society. Thus, ensuring their alignment with human values and intentions has emerged as a critical…

Recent work in large language modeling (LLMs) has used fine-tuning to align outputs with the preferences of a prototypical user. This work assumes that human preferences are static and homogeneous across individuals, so that aligning to a a…

Learning from preference feedback is essential for aligning large language models (LLMs) with human values and improving the quality of generated responses. However, existing preference learning methods rely heavily on curated data from…

Deployed artificial intelligence (AI) often impacts humans, and there is no one-size-fits-all metric to evaluate these tools. Human-centered evaluation of AI-based systems combines quantitative and qualitative analysis and human input. It…

人机交互 · 计算机科学 2023-03-14 Teresa Datta , John P. Dickerson

Political biases in Large Language Model (LLM)-based artificial intelligence (AI) systems, such as OpenAI's ChatGPT or Google's Gemini, have been previously reported. While several prior studies have attempted to quantify these biases using…

计算机与社会 · 计算机科学 2025-03-17 David Rozado

Aligning AI agents to human intentions and values is a key bottleneck in building safe and deployable AI applications. But whose values should AI agents be aligned with? Reinforcement learning with human feedback (RLHF) has emerged as the…

人工智能 · 计算机科学 2023-10-25 Abhilash Mishra

The challenge of finding compromises between agent proposals is fundamental to AI sub-fields such as argumentation, mediation, and negotiation. Building on this tradition, Elkind et al. (2021) introduced a process for coalition formation…

多智能体系统 · 计算机科学 2025-12-09 Eyal Briman , Ehud Shapiro , Nimrod Talmon

Legal argumentation is a vital cornerstone of justice, underpinning an adversarial form of law, and extensive research has attempted to augment or undertake legal argumentation via the use of computer-based automation including Artificial…

计算机与社会 · 计算机科学 2020-09-24 Lance Eliot

Modern Artificial Intelligence (AI) systems excel at diverse tasks, from image classification to strategy games, even outperforming humans in many of these domains. After making astounding progress in language learning in the recent decade,…

计算与语言 · 计算机科学 2022-01-11 Marina Dubova

Agentic AI systems, possessing capabilities for autonomous planning and action, show great potential across diverse domains. However, their practical deployment is hindered by challenges in aligning their behavior with varied human values,…

人工智能 · 计算机科学 2025-08-12 Nell Watson , Ahmed Amer , Evan Harris , Preeti Ravindra , Shujun Zhang

We conduct an incentivized laboratory experiment to study people's perception of generative artificial intelligence (GenAI) alignment in the context of economic decision-making. Using a panel of economic problems spanning the domains of…

理论经济学 · 经济学 2026-04-03 Kevin He , Ran Shorrer , Mengjia Xia

Recent advances enable Large Language Models (LLMs) to generate AI personas, yet their lack of deep contextual, cultural, and emotional understanding poses a significant limitation. This study quantitatively compared human responses with…

计算机与社会 · 计算机科学 2025-12-03 Tabia Tanzin Prama , Christopher M. Danforth , Peter Sheridan Dodds

The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH). At the same time, there are mounting concerns that LLMs exhibit a left-wing political bias. Yet, the commitment to AI…

计算与语言 · 计算机科学 2025-07-22 Thilo Hagendorff

As general-purpose artificial intelligence systems become increasingly integrated into society and are used for information seeking, content generation, problem solving, textual analysis, coding, and running processes, it is crucial to…

计算机与社会 · 计算机科学 2025-08-28 Ljubisa Bojic , Dylan Seychell , Milan Cabarkapa

We identify several important and unsettled legal questions with profound ethical and societal implications arising from generative artificial intelligence (GenAI), focusing on its distinguishable characteristics from traditional software…

计算机与社会 · 计算机科学 2024-07-03 David Atkinson , Jacob Morrison