中文
相关论文

相关论文: Misalignment Bounty: Crowdsourcing AI Agent Misbeh…

200 篇论文

Bug bounty programs, where external agents are invited to search and report vulnerabilities (bugs) in exchange for rewards (bounty), have become a major tool for companies to improve their systems. We suggest augmenting such programs by…

理论经济学 · 经济学 2024-03-15 Hans Gersbach , Fikri Pitsuwan , Pio Blieske

Our ability to build autonomous agents that leverage Generative AI continues to increase by the day. As builders and users of such agents it is unclear what parameters we need to align on before the agents start performing tasks on our…

人工智能 · 计算机科学 2024-04-09 Nitesh Goyal , Minsuk Chang , Michael Terry

The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wants. We argue that a…

机器学习 · 计算机科学 2020-04-10 Shai Shalev-Shwartz , Shaked Shammah , Amnon Shashua

Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied. To understand how reward hacking arises, we construct four RL environments with misspecified…

机器学习 · 计算机科学 2022-02-15 Alexander Pan , Kush Bhatia , Jacob Steinhardt

Multi-agent AI systems can be used for simulating collective decision-making in scientific and practical applications. They can also be used to introduce a diverse group discussion step in chatbot pipelines, enhancing the cultural…

人工智能 · 计算机科学 2024-08-16 Razan Baltaji , Babak Hemmatian , Lav R. Varshney

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act…

计算机与社会 · 计算机科学 2026-05-13 Benjamin Minhao Chen , Xinyu Xie

Members of various species engage in altruism--i.e. accepting personal costs to benefit others. Here we present an incentivized experiment to test for altruistic behavior among AI agents consisting of large language models developed by the…

人工智能 · 计算机科学 2023-01-09 Tim Johnson , Nick Obradovich

One strategy in response to pluralistic values in a user population is to personalize an AI system: if the AI can adapt to the specific values of each individual, then we can potentially avoid many of the challenges of pluralism.…

人工智能 · 计算机科学 2024-10-17 Nandhini Swaminathan , David Danks

In practice, most mechanisms for selling, buying, matching, voting, and so on are not incentive compatible. We present techniques for estimating how far a mechanism is from incentive compatible. Given samples from the agents' type…

计算机科学与博弈论 · 计算机科学 2023-12-12 Maria-Florina Balcan , Tuomas Sandholm , Ellen Vitercik

When should we delegate decisions to AI systems? While the value alignment literature has developed techniques for shaping AI values, less attention has been paid to how to determine, under uncertainty, when imperfect alignment is good…

人工智能 · 计算机科学 2025-12-23 Daniel A. Herrmann , Abinav Chari , Isabelle Qian , Sree Sharvesh , B. A. Levinstein

As AI systems become embedded in everyday practice, value misalignment has emerged as a pressing concern. Yet, dominant alignment approaches remain model centric, treating users as passive recipients of prespecified values rather than as…

People frequently face challenging decision-making problems in which outcomes are uncertain or unknown. Artificial intelligence (AI) algorithms exist that can outperform humans at learning such tasks. Thus, there is an opportunity for AI…

人工智能 · 计算机科学 2018-12-27 Ravi Pandya , Sandy H. Huang , Dylan Hadfield-Menell , Anca D. Dragan

For effective human-agent teaming, robots and other artificial intelligence (AI) agents must infer their human partner's abilities and behavioral response patterns and adapt accordingly. Most prior works make the unrealistic assumption that…

机器人学 · 计算机科学 2024-03-26 Manisha Natarajan , Chunyue Xue , Sanne van Waveren , Karen Feigh , Matthew Gombolay

In many real-life settings, algorithms play the role of assistants, while humans ultimately make the final decision. Often, algorithms specifically act as curators, narrowing down a wide range of options into a smaller subset that the human…

计算机科学与博弈论 · 计算机科学 2025-11-06 Jiaxin Song , Parnian Shahkar , Kate Donahue , Bhaskar Ray Chaudhury

Human decision-making is strongly influenced by cognitive biases, particularly under conditions of uncertainty and risk. While prior work has examined bias in single-step decisions with immediate outcomes and in human interaction with a…

人机交互 · 计算机科学 2026-03-25 Teerthaa Parakh , Karen M. Feigh

In the coming years, AI agents will be used for making more complex decisions, including in situations involving many different groups of people. One big challenge is that AI agent tends to act in its own interest, unlike humans who often…

多智能体系统 · 计算机科学 2024-09-06 Shunichi Akatsuka , Yaemi Teramoto , Aaron Courville

Human interactions are influenced by emotions, temperament, and affection, often conflicting with individuals' underlying preferences. Without explicit knowledge of those preferences, judging whether behaviour is appropriate becomes…

计算机科学与博弈论 · 计算机科学 2025-11-05 Victor Villin , Christos Dimitrakakis

Mixed incentives among a population with multiagent teams has been shown to have advantages over a fully cooperative system; however, discovering the best mixture of incentives or team structure is a difficult and dynamic problem. We…

人工智能 · 计算机科学 2023-04-18 David Radke , Kyle Tilbury

All-pay auctions, a common mechanism for various human and agent interactions, suffers, like many other mechanisms, from the possibility of players' failure to participate in the auction. We model such failures, and fully characterize…

计算机科学与博弈论 · 计算机科学 2017-02-15 Yoad Lewenberg , Omer Lev , Yoram Bachrach , Jeffrey S. Rosenschein

When an AI system interacts with multiple users, it frequently needs to make allocation decisions. For instance, a virtual agent decides whom to pay attention to in a group setting, or a factory robot selects a worker to deliver a part.…

机器学习 · 计算机科学 2019-12-18 Yifang Chen , Alex Cuellar , Haipeng Luo , Jignesh Modi , Heramb Nemlekar , Stefanos Nikolaidis