English
Related papers

Related papers: General Discounting versus Average Reward

200 papers

In this work we present and analyze a fluid-mechanical model of competition (scavenging) amongst $N$ liquid droplets (individual competitors). The eventual outcome of this competition depends sensitively on the average resource (volume) per…

Fluid Dynamics · Physics 2019-02-18 Thomas C. Hagen , Paul H. Steen

We consider a novel setting where a set of items are matched to the same set of agents repeatedly over multiple rounds. Each agent gets exactly one item per round, which brings interesting challenges to finding efficient and/or fair {\em…

Computer Science and Game Theory · Computer Science 2022-07-05 Ioannis Caragiannis , Shivika Narang

We introduce a mean field game with rank-based reward: competing agents optimize their effort to achieve a goal, are ranked according to their completion time, and paid a reward based on their relative rank. First, we propose a tractable…

Optimization and Control · Mathematics 2017-08-07 Marcel Nutz , Yuchong Zhang

The value function plays a crucial role as a measure for the cumulative future reward an agent receives in both reinforcement learning and optimal control. It is therefore of interest to study how similar the values of neighboring states…

Systems and Control · Electrical Eng. & Systems 2024-03-22 Hans Harder , Sebastian Peitz

We address the problem of reinforcement learning in which observations may exhibit an arbitrary form of stochastic dependence on past observations and actions, i.e. environments more general than (PO)MDPs. The task for an agent is to attain…

Machine Learning · Computer Science 2009-12-30 Daniil Ryabko , Marcus Hutter

What does it mean to fully understand the behavior of a network of adaptive agents? The golden standard typically is the behavior of learning dynamics in potential games, where many evolutionary dynamics, e.g., replicator, are known to…

Computer Science and Game Theory · Computer Science 2016-10-04 Ioannis Panageas , Georgios Piliouras

In this paper, we investigate the concentration properties of cumulative reward in Markov Decision Processes (MDPs), focusing on both asymptotic and non-asymptotic settings. We introduce a unified approach to characterize reward…

Machine Learning · Computer Science 2025-12-04 Borna Sayedana , Peter E. Caines , Aditya Mahajan

We study a continuous time economy where agents have asymmetric information. The informed agent (``$I$''), at time zero, receives a private signal about the risky assets' terminal payoff $\Psi(X_T)$, while the uninformed agent (``$U$'') has…

Mathematical Finance · Quantitative Finance 2024-03-19 Jerome Detemple , Scott Robertson

We present two Policy Gradient-based algorithms with general parametrization in the context of infinite-horizon average reward Markov Decision Process (MDP). The first one employs Implicit Gradient Transport for variance reduction, ensuring…

Machine Learning · Computer Science 2025-05-13 Swetha Ganesh , Washim Uddin Mondal , Vaneet Aggarwal

The ability to learn reward functions plays an important role in enabling the deployment of intelligent agents in the real world. However, comparing reward functions, for example as a means of evaluating reward learning methods, presents a…

Machine Learning · Computer Science 2022-01-26 Blake Wulfe , Ashwin Balakrishna , Logan Ellis , Jean Mercat , Rowan McAllister , Adrien Gaidon

We obtain revenue guarantees for the simple pricing mechanism of a single posted price, in terms of a natural parameter of the distribution of buyers' valuations. Our revenue guarantee applies to the single item n buyers setting, with…

Computer Science and Game Theory · Computer Science 2015-06-02 Balasubramanian Sivan , Vasilis Syrgkanis , Omer Tamuz

We show that a simple evolutionary scheme, when applied to the minority game (MG), changes the phase structure of the game. In this scheme each agent evolves individually whenever his wealth reaches the specified bankruptcy level, in…

Statistical Mechanics · Physics 2009-11-10 Baosheng Yuan , Kan Chen

In the last decade quantum machine learning has provided fascinating and fundamental improvements to supervised, unsupervised and reinforcement learning. In reinforcement learning, a so-called agent is challenged to solve a task given by…

Quantum Physics · Physics 2022-04-13 Arne Hamann , Sabine Wölk

People often face trade-offs between costs and benefits occurring at various points in time. The predominant discounting approach is to use the exponential form. Central to this approach is the discount rate, a unique parameter that…

Theoretical Economics · Economics 2024-08-13 Bach Dong-Xuan , Philippe Bich

The interaction between an artificial agent and its environment is bi-directional. The agent extracts relevant information from the environment, and affects the environment by its actions in return to accumulate high expected reward.…

Systems and Control · Computer Science 2018-06-06 Stas Tiomkin , Naftali Tishby

In this paper, we investigate the robustness of stationary mean-field equilibria in the presence of model uncertainties, specifically focusing on infinite-horizon discounted cost functions. To achieve this, we initially establish…

Systems and Control · Electrical Eng. & Systems 2026-04-10 Uğur Aydın , Naci Saldi

This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs). For MDPs that are ergodic and linear (i.e. where rewards and dynamics are linear in some known…

Machine Learning · Computer Science 2020-06-24 Nevena Lazic , Dong Yin , Mehrdad Farajtabar , Nir Levine , Dilan Gorur , Chris Harris , Dale Schuurmans

We compare the profit of the optimal third-degree price discrimination policy against a uniform pricing policy. A uniform pricing policy offers the same price to all segments of the market. Our main result establishes that for a broad class…

General Economics · Economics 2021-11-16 Dirk Bergemann , Francisco Castro , Gabriel Weintraub

We analyze the asymptotic behavior for a system of fully nonlinear parabolic and elliptic quasi variational inequalities. These equations are related to robust switching control problems introduced in [3]. We prove that, as time horizon…

Probability · Mathematics 2017-02-07 Erhan Bayraktar , Andrea Cosso , Huyên Pham

We begin by formulating and characterizing a dominance criterion for prize sequences: $x$ dominates $y$ if any impatient agent prefers $x$ to $y$. With this in hand, we define a notion of comparative patience. Alice is more patient than Bob…

Theoretical Economics · Economics 2024-07-03 Mark Whitmeyer