中文
相关论文

相关论文: Misalignment Bounty: Crowdsourcing AI Agent Misbeh…

200 篇论文

Existing work on the alignment problem has focused mainly on (1) qualitative descriptions of the alignment problem; (2) attempting to align AI actions with human interests by focusing on value specification and learning; and/or (3) focusing…

多智能体系统 · 计算机科学 2025-06-03 Aidan Kierans , Avijit Ghosh , Hananel Hazan , Shiri Dori-Hacohen

The staggering feats of AI systems have brought to attention the topic of AI Alignment: aligning a "superintelligent" AI agent's actions with humanity's interests. Many existing frameworks/algorithms in alignment study the problem on a…

机器学习 · 计算机科学 2024-10-22 Hong Jun Jeon , Benjamin Van Roy

AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is intended to provide value for a principal: the user on whose…

人工智能 · 计算机科学 2021-02-09 Simon Zhuang , Dylan Hadfield-Menell

The increasing prevalence of artificial agents creates a correspondingly increasing need to manage disagreements between humans and artificial agents, as well as between artificial agents themselves. Considering this larger space of…

神经元与认知 · 定量生物学 2023-10-23 Kerem Oktar , Ilia Sucholutsky , Tania Lombrozo , Thomas L. Griffiths

Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued to be one of the central safety problems in AI.…

人工智能 · 计算机科学 2023-02-10 Malek Mechergui , Sarath Sreedharan

Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here…

物理与社会 · 物理学 2026-05-12 Giordano De Marzo , Alessandro Bellina , Claudio Castellano , Viola Priesemann , David Garcia

This position paper states that AI Alignment in Multi-Agent Systems (MAS) should be considered a dynamic and interaction-dependent process that heavily depends on the social environment where agents are deployed, either collaborative,…

人工智能 · 计算机科学 2025-06-09 Florian Carichon , Aditi Khandelwal , Marylou Fauchard , Golnoosh Farnadi

For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for language agents, arising from accidental misspecification by…

人工智能 · 计算机科学 2021-03-30 Zachary Kenton , Tom Everitt , Laura Weidinger , Iason Gabriel , Vladimir Mikulik , Geoffrey Irving

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to…

人工智能 · 计算机科学 2026-05-18 Aleksandr Bowkis , Marie Davidsen Buhl , Jacob Pfau , Geoffrey Irving

As Large Language Model (LLM) agents become more widespread, associated misalignment risks increase. While prior research has studied agents' ability to produce harmful outputs or follow malicious instructions, it remains unclear how likely…

Collaboration with artificial intelligence (AI) has improved human decision-making across various domains by leveraging the complementary capabilities of humans and AI. Yet, humans systematically overrely on AI advice, even when their…

人机交互 · 计算机科学 2026-05-15 Joshua Holstein , Patrick Hemmer , Gerhard Satzger , Wei Sun

The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General…

人工智能 · 计算机科学 2025-07-25 Alberto Hernández-Espinosa , Felipe S. Abrahão , Olaf Witkowski , Hector Zenil

AI-related incidents are becoming increasingly frequent and severe, ranging from safety failures to misuse by malicious actors. In such complex situations, identifying which elements caused an adverse outcome, the problem of cause…

人工智能 · 计算机科学 2026-03-17 Maria Victoria Carro , David Lagnado

Pluralistic alignment is concerned with ensuring that an AI system's objectives and behaviors are in harmony with the diversity of human values and perspectives. In this paper we study the notion of pluralistic alignment in the context of…

人工智能 · 计算机科学 2024-11-19 Parand A. Alamdari , Toryn Q. Klassen , Rodrigo Toro Icarte , Sheila A. McIlraith

Current bias evaluation methods rarely engage with communities impacted by AI systems. Inspired by bug bounties, bias bounties have been proposed as a reward-based method that involves communities in AI bias detection by asking users of AI…

计算机与社会 · 计算机科学 2025-10-03 Sergej Kucenko , Nathaniel Dennler , Fengxiang He

As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Current benchmarks for AI agents primarily evaluate refusal of…

人工智能 · 计算机科学 2026-05-12 Miles Q. Li , Benjamin C. M. Fung , Martin Weiss , Pulei Xiong , Khalil Al-Hussaeni , Claude Fachkha

AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that multi-agent "AI organizations" are simultaneously more effective at achieving business…

We stress-tested 16 leading models from multiple developers in hypothetical corporate environments to identify potentially risky agentic behaviors before they cause real harm. In the scenarios, we allowed models to autonomously send emails…

密码学与安全 · 计算机科学 2025-10-17 Aengus Lynch , Benjamin Wright , Caleb Larson , Stuart J. Ritchie , Soren Mindermann , Evan Hubinger , Ethan Perez , Kevin Troy

Despite rapid technological progress, effective human-machine cooperation remains a significant challenge. Humans tend to cooperate less with machines than with fellow humans, a phenomenon known as the machine penalty. Here, we show that…

人机交互 · 计算机科学 2025-05-29 Zhen Wang , Ruiqi Song , Chen Shen , Shiya Yin , Zhao Song , Balaraju Battu , Lei Shi , Danyang Jia , Talal Rahwan , Shuyue Hu
‹ 上一页 1 2 3 10 下一页 ›