Related papers: General Discounting versus Average Reward
In this work we present and analyze a fluid-mechanical model of competition (scavenging) amongst $N$ liquid droplets (individual competitors). The eventual outcome of this competition depends sensitively on the average resource (volume) per…
We consider a novel setting where a set of items are matched to the same set of agents repeatedly over multiple rounds. Each agent gets exactly one item per round, which brings interesting challenges to finding efficient and/or fair {\em…
We introduce a mean field game with rank-based reward: competing agents optimize their effort to achieve a goal, are ranked according to their completion time, and paid a reward based on their relative rank. First, we propose a tractable…
The value function plays a crucial role as a measure for the cumulative future reward an agent receives in both reinforcement learning and optimal control. It is therefore of interest to study how similar the values of neighboring states…
We address the problem of reinforcement learning in which observations may exhibit an arbitrary form of stochastic dependence on past observations and actions, i.e. environments more general than (PO)MDPs. The task for an agent is to attain…
What does it mean to fully understand the behavior of a network of adaptive agents? The golden standard typically is the behavior of learning dynamics in potential games, where many evolutionary dynamics, e.g., replicator, are known to…
In this paper, we investigate the concentration properties of cumulative reward in Markov Decision Processes (MDPs), focusing on both asymptotic and non-asymptotic settings. We introduce a unified approach to characterize reward…
We study a continuous time economy where agents have asymmetric information. The informed agent (``$I$''), at time zero, receives a private signal about the risky assets' terminal payoff $\Psi(X_T)$, while the uninformed agent (``$U$'') has…
We present two Policy Gradient-based algorithms with general parametrization in the context of infinite-horizon average reward Markov Decision Process (MDP). The first one employs Implicit Gradient Transport for variance reduction, ensuring…
The ability to learn reward functions plays an important role in enabling the deployment of intelligent agents in the real world. However, comparing reward functions, for example as a means of evaluating reward learning methods, presents a…
We obtain revenue guarantees for the simple pricing mechanism of a single posted price, in terms of a natural parameter of the distribution of buyers' valuations. Our revenue guarantee applies to the single item n buyers setting, with…
We show that a simple evolutionary scheme, when applied to the minority game (MG), changes the phase structure of the game. In this scheme each agent evolves individually whenever his wealth reaches the specified bankruptcy level, in…
In the last decade quantum machine learning has provided fascinating and fundamental improvements to supervised, unsupervised and reinforcement learning. In reinforcement learning, a so-called agent is challenged to solve a task given by…
People often face trade-offs between costs and benefits occurring at various points in time. The predominant discounting approach is to use the exponential form. Central to this approach is the discount rate, a unique parameter that…
The interaction between an artificial agent and its environment is bi-directional. The agent extracts relevant information from the environment, and affects the environment by its actions in return to accumulate high expected reward.…
In this paper, we investigate the robustness of stationary mean-field equilibria in the presence of model uncertainties, specifically focusing on infinite-horizon discounted cost functions. To achieve this, we initially establish…
This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs). For MDPs that are ergodic and linear (i.e. where rewards and dynamics are linear in some known…
We compare the profit of the optimal third-degree price discrimination policy against a uniform pricing policy. A uniform pricing policy offers the same price to all segments of the market. Our main result establishes that for a broad class…
We analyze the asymptotic behavior for a system of fully nonlinear parabolic and elliptic quasi variational inequalities. These equations are related to robust switching control problems introduced in [3]. We prove that, as time horizon…
We begin by formulating and characterizing a dominance criterion for prize sequences: $x$ dominates $y$ if any impatient agent prefers $x$ to $y$. With this in hand, we define a notion of comparative patience. Alice is more patient than Bob…