中文
相关论文

相关论文: Optimal Anonymous Independent Reward Scheme Design

200 篇论文

Autonomous agents optimize the reward function we give them. What they don't know is how hard it is for us to design a reward function that actually captures what we want. When designing the reward, we might think of some specific training…

人工智能 · 计算机科学 2020-10-08 Dylan Hadfield-Menell , Smitha Milli , Pieter Abbeel , Stuart Russell , Anca Dragan

Designers of AI agents often iterate on the reward function in a trial-and-error process until they get the desired behavior, but this only guarantees good behavior in the training environment. We propose structuring this process as a…

机器学习 · 计算机科学 2023-10-17 Sören Mindermann , Rohin Shah , Adam Gleave , Dylan Hadfield-Menell

In the contest design problem, there are $n$ strategic contestants, each of whom decides an effort level. A contest designer with a fixed budget must then design a mechanism that allocates a prize $p_i$ to the $i$-th rank based on the…

计算机科学与博弈论 · 计算机科学 2026-04-07 Negin Golrezaei , MohammadTaghi Hajiaghayi , Suho Shin

We analyze a two-period principal-agent model in which the principal faces a budget constraint, and the agent's private costs of performing tasks across the two periods may be correlated. We examine the optimal design of the reward scheme…

最优化与控制 · 数学 2025-12-30 Eilon Solan , Avraham Tabbach , Chang Zhao

Reward design is a critical part of the application of reinforcement learning, the performance of which strongly depends on how well the reward signal frames the goal of the designer and how well the signal assesses progress in reaching…

机器学习 · 计算机科学 2022-08-01 Yixiang Wang , Yujing Hu , Feng Wu , Yingfeng Chen

To convey desired behavior to a Reinforcement Learning (RL) agent, a designer must choose a reward function for the environment, arguably the most important knob designers have in interacting with RL agents. Although many reward functions…

机器学习 · 计算机科学 2022-06-01 Henry Sowerby , Zhiyuan Zhou , Michael L. Littman

We consider schemes for obtaining truthful reports on a common but hidden signal from large groups of rational, self-interested agents. One example are online feedback mechanisms, where users provide observations about the quality of a…

计算机科学与博弈论 · 计算机科学 2014-01-16 Radu Jurca , Boi Faltings

Online platforms in the Internet Economy commonly incorporate recommender systems that recommend products (or "arms") to users (or "agents"). A key challenge in this domain arises from myopic agents who are naturally incentivized to exploit…

信息检索 · 计算机科学 2024-06-19 Xiaowu Dai , Wenlu Xu , Yuan Qi , Michael I. Jordan

We consider an outsourcing problem where a software agent procures multiple services from providers with uncertain reliabilities to complete a computational task before a strict deadline. The service consumer requires a procurement strategy…

计算机科学与博弈论 · 计算机科学 2021-10-26 Farzaneh Farhadi , Maria Chli , Nicholas R. Jennings

We study reward design strategies for incentivizing a reinforcement learning agent to adopt a policy from a set of admissible policies. The goal of the reward designer is to modify the underlying reward function cost-efficiently while…

机器学习 · 计算机科学 2022-01-07 Kiarash Banihashem , Adish Singla , Jiarui Gan , Goran Radanovic

We study a decision-maker's problem of finding optimal monetary incentive schemes for retention when faced with agents whose participation decisions (stochastically) depend on the incentive they receive. Our focus is on policies constrained…

计算机科学与博弈论 · 计算机科学 2024-07-31 Daniel Freund , Chamsi Hssaine

Reward design is a fundamental problem in reinforcement learning (RL). A misspecified or poorly designed reward can result in low sample efficiency and undesired behaviors. In this paper, we propose the idea of programmatic reward design,…

机器学习 · 计算机科学 2022-01-10 Weichao Zhou , Wenchao Li

We study allocation problems without monetary transfers where agents have correlated types, i.e., hold private information about one another. Such peer information is relevant in various settings, including science funding, allocation of…

理论经济学 · 经济学 2025-03-21 Axel Niemeyer , Justus Preusser

We consider a generalization of the densest subhypergraph problem where nonnegative rewards are given for including partial hyperedges in a dense subhypergraph. Prior work addressed this problem only in cases where reward functions are…

数据结构与算法 · 计算机科学 2025-06-17 Vedangi Bengali , Nikolaj Tatti , Iiro Kumpulainen , Florian Adriaens , Nate Veldt

Sparse rewards are a major bottleneck in multi-agent reinforcement learning (MARL), where simultaneous learning induces non-stationarity and makes reward design especially delicate. Reward shaping can accelerate learning, but in the…

多智能体系统 · 计算机科学 2026-05-25 Elie Abboud , Oren Gal

For selling a single item to agents with independent but non-identically distributed values, the revenue optimal auction is complex. With respect to it, Hartline and Roughgarden (2009) showed that the approximation factor of the…

计算机科学与博弈论 · 计算机科学 2016-11-17 Saeed Alaei , Jason Hartline , Rad Niazadeh , Emmanouil Pountourakis , Yang Yuan

Points-based rewards programs are a prevalent way to incentivize customer loyalty; in these programs, customers who make repeated purchases from a seller accumulate points, working toward eventual redemption of a free reward. These programs…

机器学习 · 计算机科学 2025-06-05 Chamsi Hssaine , Yichun Hu , Ciara Pike-Burke

We present an algorithm to approximate the solutions to variational problems where set of admissible functions consists of convex functions. The main motivator behind this numerical method is estimating solutions to Adverse Selection…

最优化与控制 · 数学 2008-03-07 Ivar Ekeland , Santiago Moreno

We investigate the mechanism design problem faced by a principal who hires \emph{multiple} agents to gather and report costly information. Then, the principal exploits the information to make an informed decision. We model this problem as a…

计算机科学与博弈论 · 计算机科学 2023-07-13 Federico Cacciamani , Matteo Castiglioni , Nicola Gatti

Incentives are more likely to elicit desired outcomes when they are designed based on accurate models of agents' strategic behavior. A growing literature, however, suggests that people do not quite behave like standard economic agents in a…

计算机科学与博弈论 · 计算机科学 2014-06-09 Arpita Ghosh , Robert Kleinberg
‹ 上一页 1 2 3 10 下一页 ›