中文
相关论文

相关论文: Moral Anchor System: A Predictive Framework for AI…

200 篇论文

AI alignment considers how we can encode AI systems in a way that is compatible with human values. The normative side of this problem asks what moral values or principles, if any, we should encode in AI. To this end, we present a framework…

计算机与社会 · 计算机科学 2023-01-11 Betty Li Hou , Brian Patrick Green

Artificial Intelligence (AI) is poised to transform healthcare delivery through revolutionary advances in clinical decision support and diagnostic capabilities. While human expertise remains foundational to medical practice, AI-powered…

As Artificial Intelligence (AI) technologies continue to advance, protecting human autonomy and promoting ethical decision-making are essential to fostering trust and accountability. Human agency (the capacity of individuals to make…

计算机与社会 · 计算机科学 2025-10-13 Laxmiraju Kandikatla , Branislav Radeljic

Organisations that design and deploy artificial intelligence (AI) systems increasingly commit themselves to high-level, ethical principles. However, there still exists a gap between principles and practices in AI ethics. One major obstacle…

计算机与社会 · 计算机科学 2024-07-09 Jakob Mokander , Margi Sheth , David Watson , Luciano Floridi

AI agents are beginning to complete valuable, long-horizon business operations tasks, but training and evaluation environments for enterprise work still struggle to balance realism, verifiability, and scale. Environment and task creation…

人工智能 · 计算机科学 2026-05-27 Maksim Ivanov , Abhijay Rana

Agentic AI systems capable of reasoning, planning, and executing actions present fundamentally distinct governance challenges compared to traditional AI models. Unlike conventional AI, these systems exhibit emergent and unexpected behaviors…

人工智能 · 计算机科学 2025-11-19 Charles L. Wang , Trisha Singhal , Ameya Kelkar , Jason Tuo

Artificial Intelligence (AI) technology epitomizes the complex challenges posed by human-made artifacts, particularly those widely integrated into society and exerting significant influence, highlighting potential benefits and their…

人工智能 · 计算机科学 2025-10-06 Michael Papademas , Xenia Ziouvelou , Antonis Troumpoukis , Vangelis Karkaletsis

This paper aims to provide an overview of the ethical concerns in artificial intelligence (AI) and the framework that is needed to mitigate those risks, and to suggest a practical path to ensure the development and use of AI at the United…

计算机与社会 · 计算机科学 2021-04-27 Lambert Hogenhout

Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued to be one of the central safety problems in AI.…

人工智能 · 计算机科学 2023-02-10 Malek Mechergui , Sarath Sreedharan

Humans and machines interact more frequently than ever and our societies are becoming increasingly hybrid. A consequence of this hybridisation is the degradation of societal trust due to the prevalence of AI-enabled deception. Yet, despite…

多智能体系统 · 计算机科学 2024-06-12 Stefan Sarkadi

AI Alignment research seeks to align human and AI goals to ensure independent actions by a machine are always ethical. This paper argues empathy is necessary for this task, despite being often neglected in favor of more deductive…

神经与进化计算 · 计算机科学 2023-12-14 Devin Gonier , Adrian Adduci , Cassidy LoCascio

AI-based systems can increasingly perform work tasks autonomously. In safety-critical tasks, human oversight of these systems is required to mitigate risks and to ensure responsibility in case something goes wrong. Since people often…

人机交互 · 计算机科学 2026-02-12 Cedric Faas , Richard Uth , Sarah Sterz , Markus Langer , Anna Maria Feit

Risk thresholds provide a measure of the level of risk exposure that a society or individual is willing to withstand, ultimately shaping how we determine the safety of technological systems. Against the backdrop of the Cold War, the first…

计算机与社会 · 计算机科学 2025-04-22 Heidy Khlaaf , Sarah Myers West

Reducing the number of failures in a production system is one of the most challenging problems in technology driven industries, such as, the online retail industry. To address this challenge, change management has emerged as a promising…

机器学习 · 计算机科学 2021-08-19 Binay Gupta , Anirban Chatterjee , Harika Matha , Kunal Banerjee , Lalitdutt Parsai , Vijay Agneeswaran

Enhancing the moral alignment of Large Language Models (LLMs) is a critical challenge in AI safety. Current alignment techniques often act as superficial guardrails, leaving the intrinsic moral representations of LLMs largely untouched. In…

计算与语言 · 计算机科学 2026-01-16 Luoming Hu , Jingjie Zeng , Liang Yang , Hongfei Lin

Artificial intelligence (AI) ethics has gained significant momentum, evidenced by the growing body of published literature, policy guidelines, and public discourse. However, the practical implementation and adoption of AI ethics principles…

计算机与社会 · 计算机科学 2025-03-03 Sarah Hladikova , Yuling Wang , Andreia Martinho

Detecting AI risks becomes more challenging as stronger models emerge and find novel methods such as Alignment Faking to circumvent these detection attempts. Inspired by how risky behaviors in humans (i.e., illegal activities that may hurt…

计算与语言 · 计算机科学 2025-05-21 Yu Ying Chiu , Zhilin Wang , Sharan Maiya , Yejin Choi , Kyle Fish , Sydney Levine , Evan Hubinger

Agentic artificial intelligence (AI) -- multi-agent systems that combine large language models with external tools and autonomous planning -- are rapidly transitioning from research laboratories into high-stakes domains. Our earlier "Basic"…

人工智能 · 计算机科学 2025-09-16 Manish Shukla

LLM-powered Multi-Agent Systems (LLM-MAS) unlock new potentials in distributed reasoning, collaboration, and task generalization but also introduce additional risks due to unguaranteed agreement, cascading uncertainty, and adversarial…

多智能体系统 · 计算机科学 2025-10-22 Jinwei Hu , Yi Dong , Shuang Ao , Zhuoyun Li , Boxuan Wang , Lokesh Singh , Guangliang Cheng , Sarvapali D. Ramchurn , Xiaowei Huang

Given rapid progress toward advanced AI and risks from frontier AI systems (advanced AI systems pushing the boundaries of the AI capabilities frontier), the creation and implementation of AI governance and regulatory schemes deserves…