中文
相关论文

相关论文: Full-Stack Alignment: Co-Aligning AI and Instituti…

200 篇论文

Artificial Intelligence (AI) started out with an ambition to reproduce the human mind, but, as the sheer scale of that ambition became manifest, it quickly retreated into either studying specialized intelligent behaviours, or proposing…

人工智能 · 计算机科学 2021-06-17 Alexander Boer , Giovanni Sileno

The rapid integration of generative AI into everyday life underscores the need to move beyond unidirectional alignment models that only adapt AI to human values. This workshop focuses on bidirectional human-AI alignment, a dynamic,…

Value-aware AI should recognise human values and adapt to the value systems (value-based preferences) of different users. This requires operationalization of values, which can be prone to misspecification. The social nature of values…

人工智能 · 计算机科学 2026-02-12 Andrés Holgado-Sánchez , Peter Vamplew , Richard Dazeley , Sascha Ossowski , Holger Billhardt

Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. We argue that there exist many different coherent identity boundaries (e.g.\ instance, model, persona), and…

As large language models (LLMs) increasingly shape content generation, interaction, and decision-making across the Web, aligning them with human values has become a central objective in trustworthy AI. This challenge becomes even more…

机器学习 · 计算机科学 2026-05-12 Hefei Xu , Le Wu , Yu Wang , Min Hou , Han Wu , Zhen Zhang , Meng Wang

Artificial intelligence is humanity's most promising technology because of the remarkable capabilities offered by foundation models. Yet, the same technology brings confusion and consternation: foundation models are poorly understood and…

人工智能 · 计算机科学 2025-07-01 Rishi Bommasani

This paper grounds ethics in evolutionary biology, viewing moral norms as adaptive mechanisms that render cooperation fitness-viable under selection pressure. Current alignment approaches add ethics post hoc, treating it as an external…

计算机与社会 · 计算机科学 2025-10-17 Dylan Waldner

An important step in the development of value alignment (VA) systems in AI is understanding how VA can reflect valid ethical principles. We propose that designers of VA systems incorporate ethics by utilizing a hybrid approach in which both…

人工智能 · 计算机科学 2020-12-23 Tae Wan Kim , John Hooker , Thomas Donaldson

Present practice of deciding on regulation faces numerous problems that make adopted regulations static, unexplained, unduly influenced by powerful interest groups, and stained with a perception of illegitimacy. These well-known problems…

计算机与社会 · 计算机科学 2026-04-02 Thomas Hofweber , Andreas Sudmann , Evangelos Pournaras

As data-driven modeling of physical dynamical systems becomes more prevalent, a new challenge is emerging: making these models more compatible and aligned with existing human knowledge. AI-driven scientific modeling processes typically…

机器学习 · 计算机科学 2024-10-11 Kevin Zhang , Hod Lipson

With the growing attention and investment in recent AI approaches such as large language models, the narrative that the larger the AI system the more valuable, powerful and interesting it is is increasingly seen as common sense. But what is…

计算机与社会 · 计算机科学 2025-03-04 Gaël Varoquaux , Alexandra Sasha Luccioni , Meredith Whittaker

This paper develops a comprehensive framework for artificial intelligence systems that operate under strict epistemic constraints, moving beyond stochastic language prediction to support structured reasoning, propositional commitment, and…

计算机科学中的逻辑 · 计算机科学 2025-06-24 Craig Steven Wright

Determining the similarities and differences between humans and artificial intelligence (AI) is an important goal both in computational cognitive neuroscience and machine learning, promising a deeper understanding of human cognition and…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Florian P. Mahner , Lukas Muttenthaler , Umut Güçlü , Martin N. Hebart

Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). We sketch a roadmap…

人工智能 · 计算机科学 2025-03-19 Florian Mai , David Kaczér , Nicholas Kluge Corrêa , Lucie Flek

Increasingly, laws are being proposed and passed by governments around the world to regulate Artificial Intelligence (AI) systems implemented into the public and private sectors. Many of these regulations address the transparency of AI…

计算机与社会 · 计算机科学 2022-07-05 Andrew Bell , Oded Nov , Julia Stoyanovich

Artificial Intelligence (AI) systems are not intrinsically neutral and biases trickle in any type of technological tool. In particular when dealing with people, the impact of AI algorithms' technical errors originating with mislabeled data…

人工智能 · 计算机科学 2025-04-03 Camilla Quaresmini , Giuseppe Primiero

Innovations in AI have focused primarily on the questions of "what" and "how"-algorithms for finding patterns in web searches, for instance-without adequate attention to the possible harms (such as privacy, bias, or manipulation) and…

计算机与社会 · 计算机科学 2020-12-14 Suresh Venkatasubramanian , Nadya Bliss , Helen Nissenbaum , Melanie Moses

Fairness in AI-driven decision-making systems has become a critical concern, especially when these systems directly affect human lives. This paper explores the public's comprehension of fairness in healthcare recommendations. We conducted a…

机器学习 · 计算机科学 2024-09-10 Veronica Kecki , Alan Said

Instruction-tuned Large Language Models (LLMs) are increasingly deployed as AI Assistants in firms for support in cognitive tasks. These AI assistants carry embedded perspectives which influence factors across the firm including…

计算机与社会 · 计算机科学 2025-05-27 Noah Broestl , Benjamin Lange , Cristina Voinea , Geoff Keeling , Rachael Lam

A rising vision for AI in the open world centers on the development of systems that can complement humans for perceptual, diagnostic, and reasoning tasks. To date, systems aimed at complementing the skills of people have employed models…

人工智能 · 计算机科学 2020-05-05 Bryan Wilder , Eric Horvitz , Ece Kamar