中文
相关论文

相关论文: Aligning Artificial Superintelligence via a Multi-…

200 篇论文

Given that AI systems are set to play a pivotal role in future decision-making processes, their trustworthiness and reliability are of critical concern. Due to their scale and complexity, modern AI systems resist direct interpretation, and…

人工智能 · 计算机科学 2025-01-03 Binxia Xu , Antonis Bikakis , Daniel Onah , Andreas Vlachidis , Luke Dickens

In coming years or decades, artificial general intelligence (AGI) may surpass human capabilities across many critical domains. We argue that, without substantial effort to prevent it, AGIs could learn to pursue goals that are in conflict…

人工智能 · 计算机科学 2025-05-06 Richard Ngo , Lawrence Chan , Sören Mindermann

In the age of information overload, professionals across various fields face the challenge of navigating vast amounts of documentation and ever-evolving standards. Ensuring compliance with standards, regulations, and contractual obligations…

密码学与安全 · 计算机科学 2024-07-22 Shohreh Deldari , Mohammad Goudarzi , Aditya Joshi , Arash Shaghaghi , Simon Finn , Flora D. Salim , Sanjay Jha

The proliferation of AI agents, with their complex and context-dependent actions, renders conventional privacy paradigms obsolete. This position paper argues that the current model of privacy management, rooted in a user's unilateral…

人机交互 · 计算机科学 2025-08-12 Shuning Zhang , Ying Ma , Jingruo Chen , Simin Li , Xin Yi , Hewu Li

This paper reviews the current development of artificial intelligence (AI) techniques for the application area of robot communication. The study of the control and operation of multiple robots collaboratively toward a common goal is fast…

信号处理 · 电气工程与系统科学 2018-05-01 S. H. Alsamhi , Ou Ma , M. S. Ansari

The external evaluation of AI systems is increasingly recognised as a crucial approach for understanding their potential risks. However, facilitating external evaluation in practice faces significant challenges in balancing evaluators' need…

计算机与社会 · 计算机科学 2025-03-04 Ben Bucknall , Robert F. Trager , Michael A. Osborne

We study AI alignment through the lens of law-and-economics models of deterrence and enforcement. In these models, misconduct is not treated as an external failure, but as a strategic response to incentives: an actor weighs the gain from…

机器学习 · 计算机科学 2026-05-12 Rohit Agarwal , Joshua Lin , Mark Braverman , Elad Hazan

We propose that future AI transparency and accountability regulations are based on an open global standard for exchanging information about AI systems, which allows co-existence of potentially conflicting local regulations. Then, we discuss…

计算机与社会 · 计算机科学 2026-01-22 Warren Buckley , Adrian Byrne , Nicholas Perello , Cyrus Cousins , Taha Yasseri , Yair Zick , Przemyslaw Grabowicz

The Artificial intelligence in critical sectors-healthcare, finance, and public safety-has made system integrity paramount for maintaining societal trust. Current verification methods for AI systems lack comprehensive lifecycle assurance,…

密码学与安全 · 计算机科学 2024-11-04 Mahesh Vaijainthymala Krishnamoorthy

Decentralized, agentic AI marketplaces are rapidly emerging to support software engineering tasks such as debugging, patch generation, and security auditing, often operating without centralized oversight. However, existing reputation…

人工智能 · 计算机科学 2026-05-04 Mohd Sameen Chishti , Damilare Peter Oyinloye , Jingyue Li

Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes increasingly challenging. This paper explores how we can…

人工智能 · 计算机科学 2025-10-31 Rishub Jain , Sophie Bridgers , Lili Janzer , Rory Greig , Tian Huey Teh , Vladimir Mikulik

Productive human-AI collaboration requires appropriate reliance, yet contemporary AI systems are often miscalibrated, exhibiting systematic overconfidence or underconfidence. We investigate whether humans can learn to mentally recalibrate…

人机交互 · 计算机科学 2026-03-25 ZhaoBin Li , Mark Steyvers

In human-AI decision making, designing AI that complements human expertise has been a natural strategy to enhance human-AI collaboration, yet it often comes at the cost of decreased AI performance in areas of human strengths. This can…

人工智能 · 计算机科学 2026-02-24 Hasan Amin , Ming Yin , Rajiv Khanna

Recent advances in artificial intelligence (AI) have lead to an explosion of multimedia applications (e.g., computer vision (CV) and natural language processing (NLP)) for different domains such as commercial, industrial, and intelligence.…

计算机与社会 · 计算机科学 2019-11-14 Erik Blasch , James Sung , Tao Nguyen , Chandra P. Daniel , Alisa P. Mason

Atomic multicast is a communication primitive that delivers messages to multiple groups of processes according to some total order, with each group receiving the projection of the total order onto messages addressed to it. To be scalable,…

分布式、并行与集群计算 · 计算机科学 2019-04-16 Alexey Gotsman , Anatole Lefort , Gregory Chockler

Collaborative problem solving and learning are shaped by who or what is on the team. As large language models (LLMs) increasingly function as collaborators rather than tools, a key question is whether AI teammates can be aligned to express…

人机交互 · 计算机科学 2026-03-03 Mohammad Amin Samadi , Nia Nixon

Ensuring artificial intelligence behaves in such a way that is aligned with human values is commonly referred to as the alignment challenge. Prior work has shown that rational agents, behaving in such a way that maximizes a utility…

人工智能 · 计算机科学 2024-02-16 Paulo Garcia

Artificial intelligence (AI) systems, such as machine learning algorithms, have allowed scientists, marketers and governments to shed light on correlations that remained invisible until now. Beforehand, the dots that we had to connect in…

计算机与社会 · 计算机科学 2022-02-08 Remy Demichelis

Classical robot ethics is often framed around obedience, including Asimov's laws. This framing is insufficient for contemporary AI systems, which are increasingly adaptive, generative, embodied, and embedded in physical, psychological, and…

计算机与社会 · 计算机科学 2026-04-30 Somyajit Chakraborty

Explainable Artificial Intelligence (XAI) has become critical in enhancing the transparency and trustworthiness of AI systems, especially as these systems are increasingly deployed in high-stakes domains such as healthcare and finance.…

符号计算 · 计算机科学 2024-08-13 Shengxin Hong , Xiuyi Fan