中文
相关论文

相关论文: General Discounting versus Average Reward

200 篇论文

We study the effect of interim feedback policies in a dynamic all-pay auction where two players bid over two stages to win a common-value prize. We show that sequential equilibrium outcomes are characterized by Cheapest Signal Equilibria,…

理论经济学 · 经济学 2025-10-28 Sumit Goel , Yiqing Yan , Jeffrey Zeidel

When allocating indivisible items, there are various ways to use monetary transfers for eliminating envy. Particularly, one can apply a balanced vector of transfer payments, or charge each agent a positive amount, or -- contrarily -- give…

计算机科学与博弈论 · 计算机科学 2025-06-24 Noga Klein Elmalem , Rica Gonen , Erel Segal-Halevi

Mechanism design is a well-established game-theoretic paradigm for designing games to achieve desired outcomes. This paper addresses a closely related but distinct concept, equilibrium design. Unlike mechanism design, the designer's…

计算机科学与博弈论 · 计算机科学 2024-08-20 Muhammad Najib , Giuseppe Perelli

The difference between the speed of the actions of different processes is typically considered as an obstacle that makes the achievement of cooperative goals more difficult. In this work, we aim to highlight potential benefits of such…

分布式、并行与集群计算 · 计算机科学 2015-09-15 Ofer Feinerman , Amos Korman , Shay Kutten , Yoav Rodeh

Mean Field Game (MFG) systems describe equilibrium configurations in games with infinitely many interacting controllers. We are interested in the behavior of this system as the horizon becomes large, or as the discount factor tends to $0$.…

最优化与控制 · 数学 2019-03-13 Pierre Cardaliaguet , Alessio Porretta

We consider the expressivity of Markov rewards in sequential decision making under uncertainty. We view reward functions in Markov Decision Processes (MDPs) as a means to characterize desired behaviors of agents. Assuming desired behaviors…

人工智能 · 计算机科学 2023-07-25 Shuwa Miura

A network of agents cooperate on a given area. Time evolution of their power is described within a set of nonlinear equations. The limitation of resources is introduced via the Verhulst term, equivalent to a global coupling. Each agent is…

凝聚态物理 · 物理学 2007-05-23 K. Malarz , K. Kulakowski

We study population dynamics under which each revising agent tests each strategy k times, with each trial being against a newly drawn opponent, and chooses the strategy whose mean payoff was highest. When k = 1, defection is globally stable…

理论经济学 · 经济学 2021-01-05 Srinivas Arigapudi , Yuval Heller , Igal Milchtaich

Reinforcement learning (RL) has traditionally been understood from an episodic perspective; the concept of non-episodic RL, where there is no restart and therefore no reliable recovery, remains elusive. A fundamental question in…

机器学习 · 计算机科学 2021-05-31 Shuang Liu , Hao Su

In the standard minority game, each agent in the minority group receives the same payoff regardless of the size of the minority group. Of great interest for real social and biological systems are cases in which the payoffs to members of the…

适应与自组织系统 · 物理学 2009-10-31 Yi Li , Adrian VanDeemen , Robert Savit

Partially observable Markov decision processes (POMDPs) with stage duration provide a framework for approximating continuous-time behavior by scaling transition probabilities with a stage duration parameter $h \in (0,1]$. While previous…

最优化与控制 · 数学 2026-03-18 Ivan Novikov

Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied. To understand how reward hacking arises, we construct four RL environments with misspecified…

机器学习 · 计算机科学 2022-02-15 Alexander Pan , Kush Bhatia , Jacob Steinhardt

In the last decade, a large body of literature has been developed to explain the universal features of inequality in terms of income and wealth. By now, it is established that the distributions of income and wealth in various economies show…

综合金融 · 定量金融 2016-11-25 Anindya S. Chakrabarti , Bikas K. Chakrabarti

We consider a renewal-reward process with multivariate rewards. Such a process is constructed from an i.i.d.\ sequence of time periods, to each of which there is associated a multivariate reward vector. The rewards in each time period may…

概率论 · 数学 2014-08-08 Brendan Patch , Yoni Nazarathy , Thomas Taimre

In practice, most mechanisms for selling, buying, matching, voting, and so on are not incentive compatible. We present techniques for estimating how far a mechanism is from incentive compatible. Given samples from the agents' type…

计算机科学与博弈论 · 计算机科学 2023-12-12 Maria-Florina Balcan , Tuomas Sandholm , Ellen Vitercik

A set of many identical interacting agents obeying a global additive constraint is considered. Under the hypothesis of equiprobability in the high-dimensional volume delimited in phase space by the constraint, the statistical behavior of a…

混沌动力学 · 物理学 2007-09-03 Ricardo Lopez-Ruiz , Jaime Sanudo , Xavier Calbet

We introduce a novel approach to hierarchical reinforcement learning for Linearly-solvable Markov Decision Processes (LMDPs) in the infinite-horizon average-reward setting. Unlike previous work, our approach allows learning low-level and…

机器学习 · 计算机科学 2024-07-10 Guillermo Infante , Anders Jonsson , Vicenç Gómez

In constant-payoff finite population games, when selection is weak and population size is large, the one-third law serves as the condition for a strategy to be advantageous. We generalize the result to the case where payoff matrices are…

种群与进化 · 定量生物学 2013-12-16 Weihong Xu , Yanling Zhang , Guangming Xie

Auction is applied for trade with various mechanisms. A simple but practical question is which mechanism, typically first-price or second-price auctions, is preferred from the perspective of bidders or sellers. A celebrated answer is…

计算机科学与博弈论 · 计算机科学 2026-02-20 Yuma Fujimoto , Kaito Ariu , Kenshi Abe

We consider average-energy games, where the goal is to minimize the long-run average of the accumulated energy. While several results have been obtained on these games recently, decidability of average-energy games with a lower-bound…

计算机科学中的逻辑 · 计算机科学 2017-01-16 Patricia Bouyer , Piotr Hofman , Nicolas Markey , Mickael Randour , Martin Zimmermann