English
Related papers

Related papers: General Discounting versus Average Reward

200 papers

We consider reinforcement learning for continuous-time Markov decision processes (MDPs) in the infinite-horizon, average-reward setting. In contrast to discrete-time MDPs, a continuous-time process moves to a state and stays there for a…

Machine Learning · Computer Science 2024-07-03 Xuefeng Gao , Xun Yu Zhou

We consider an ordinary differential equation with a unique hyperbolic attractor at the origin, to which we add a small random perturbation. It is known that under general conditions, the solution of this stochastic differential equation…

Probability · Mathematics 2023-05-05 Gerardo Barrera , Milton Jara

In many professons employees are rewarded according to their relative performance. Corresponding economy can be modeled by taking $N$ independent agents who gain from the market with a rate which depends on their current gain. We argue that…

Popular Physics · Physics 2010-09-03 P. K. Mohanty

We show that, with indivisible goods, the existence of competitive equilibrium fundamentally depends on agents' substitution effects, not their income effects. Our Equilibrium Existence Duality allows us to transport results on the…

Theoretical Economics · Economics 2020-07-01 Elizabeth Baldwin , Omer Edhan , Ravi Jagadeesan , Paul Klemperer , Alexander Teytelboym

There are situations in which an agent should receive rewards only after having accomplished a series of previous tasks. In other words, the reward that the agent receives is non-Markovian. One natural and quite general way to represent…

Artificial Intelligence · Computer Science 2020-01-28 Gavin Rens , Jean-François Raskin

In the framework of evolutionary games with institutional reciprocity, limited incentives are at disposal for rewarding cooperators and punishing defectors. In the simplest case, it can be assumed that, depending on their strategies, all…

Populations and Evolution · Quantitative Biology 2014-09-05 Xiaojie Chen , Matjaz Perc

Standard reinforcement learning (RL) assumes that an agent can observe a reward for each state-action pair. However, in practical applications, it is often difficult and costly to collect a reward for each state-action pair. While there…

Machine Learning · Computer Science 2025-06-18 Yihan Du , Anna Winnicki , Gal Dalal , Shie Mannor , R. Srikant

Reinforcement learning (RL) agents have traditionally been tasked with maximizing the value function of a Markov decision process (MDP), either in continuous settings, with fixed discount factor $\gamma < 1$, or in episodic settings, with…

Machine Learning · Computer Science 2019-02-11 Silviu Pitis

Consider a game where Alice generates an integer and Bob wins if he can factor that integer. Traditional game theory tells us that Bob will always win this game even though in practice Alice will win given our usual assumptions about the…

Computer Science and Game Theory · Computer Science 2009-11-18 Lance Fortnow , Rahul Santhanam

In this paper, we investigate the Merton portfolio management problem in the context of non-exponential discounting. This gives rise to time-inconsistency of the decision-maker. If the decision-maker at time t=0 can commit his/her…

Portfolio Management · Quantitative Finance 2008-12-02 Ivar Ekeland , Traian A. Pirvu

We study $n$-dimensional contests between two players with heterogeneous effort costs, where each dimension (battle) is modeled as a Tullock contest. Prize-allocation rules are identity-independent, budget-balanced, and weakly increasing in…

Theoretical Economics · Economics 2026-03-31 Siyuan Fan , Zhonghong Kuang , Jingfeng Lu

The classical policy gradient method is the theoretical and conceptual foundation of modern policy-based reinforcement learning (RL) algorithms. Most rigorous analyses of such methods, particularly those establishing convergence guarantees,…

Machine Learning · Computer Science 2026-02-11 Jongmin Lee , Ernest K. Ryu

Exploiting others is beneficial individually but it could also be detrimental globally. The reverse is also true: a higher cooperation level may change the environment in a way that is beneficial for all competitors. To explore the possible…

Physics and Society · Physics 2018-02-23 Attila Szolnoki , Xiaojie Chen

Once failure is irreversible, continuation payoffs cannot be meaningfully aggregated across strategies that differ in their survival properties. Standard scalar evaluation sidesteps this by arbitrarily completing payoffs beyond termination,…

Theoretical Economics · Economics 2026-02-10 Nicholas H. Kirk

Several practical applications of reinforcement learning involve an agent learning from past data without the possibility of further exploration. Often these applications require us to 1) identify a near optimal policy or to 2) estimate the…

Machine Learning · Computer Science 2021-06-22 Andrea Zanette

We study approachability theory in the presence of constraints. Given a repeated game with vector payoffs, we characterize the pairs of sets (A,D) in the payoff space such that Player 1 can guarantee that the long-run average payoff…

Optimization and Control · Mathematics 2017-12-05 Gaëtan Fournier , Eden Kuperwasser , Orin Munk , Eilon Solan , Avishay Weinbaum

We consider a number of questions related to tradeoffs between reward and regret in repeated gameplay between two agents. To facilitate this, we introduce a notion of $\textit{generalized equilibrium}$ which allows for asymmetric regret…

Computer Science and Game Theory · Computer Science 2023-12-19 William Brown , Jon Schneider , Kiran Vodrahalli

We develop theory and algorithms for average-reward on-policy Reinforcement Learning (RL). We first consider bounding the difference of the long-term average reward for two policies. We show that previous work based on the discounted return…

Machine Learning · Computer Science 2021-06-15 Yiming Zhang , Keith W. Ross

We study the problem of Inverse Reinforcement Learning (IRL) with an average-reward criterion. The goal is to recover an unknown policy and a reward function when the agent only has samples of states and actions from an experienced agent.…

Machine Learning · Computer Science 2023-05-25 Feiyang Wu , Jingyang Ke , Anqi Wu

As is well known, average-cost optimality inequalities imply the existence of stationary optimal policies for Markov Decision Processes with average costs per unit time, and these inequalities hold under broad natural conditions. This paper…

Optimization and Control · Mathematics 2016-10-04 Eugene A. Feinberg , Yan Liang