中文
相关论文

相关论文: General Discounting versus Average Reward

200 篇论文

In this work we present and analyze a fluid-mechanical model of competition (scavenging) amongst $N$ liquid droplets (individual competitors). The eventual outcome of this competition depends sensitively on the average resource (volume) per…

流体动力学 · 物理学 2019-02-18 Thomas C. Hagen , Paul H. Steen

We consider a novel setting where a set of items are matched to the same set of agents repeatedly over multiple rounds. Each agent gets exactly one item per round, which brings interesting challenges to finding efficient and/or fair {\em…

计算机科学与博弈论 · 计算机科学 2022-07-05 Ioannis Caragiannis , Shivika Narang

We introduce a mean field game with rank-based reward: competing agents optimize their effort to achieve a goal, are ranked according to their completion time, and paid a reward based on their relative rank. First, we propose a tractable…

最优化与控制 · 数学 2017-08-07 Marcel Nutz , Yuchong Zhang

The value function plays a crucial role as a measure for the cumulative future reward an agent receives in both reinforcement learning and optimal control. It is therefore of interest to study how similar the values of neighboring states…

系统与控制 · 电气工程与系统科学 2024-03-22 Hans Harder , Sebastian Peitz

We address the problem of reinforcement learning in which observations may exhibit an arbitrary form of stochastic dependence on past observations and actions, i.e. environments more general than (PO)MDPs. The task for an agent is to attain…

机器学习 · 计算机科学 2009-12-30 Daniil Ryabko , Marcus Hutter

What does it mean to fully understand the behavior of a network of adaptive agents? The golden standard typically is the behavior of learning dynamics in potential games, where many evolutionary dynamics, e.g., replicator, are known to…

计算机科学与博弈论 · 计算机科学 2016-10-04 Ioannis Panageas , Georgios Piliouras

In this paper, we investigate the concentration properties of cumulative reward in Markov Decision Processes (MDPs), focusing on both asymptotic and non-asymptotic settings. We introduce a unified approach to characterize reward…

机器学习 · 计算机科学 2025-12-04 Borna Sayedana , Peter E. Caines , Aditya Mahajan

We study a continuous time economy where agents have asymmetric information. The informed agent (``$I$''), at time zero, receives a private signal about the risky assets' terminal payoff $\Psi(X_T)$, while the uninformed agent (``$U$'') has…

数理金融 · 定量金融 2024-03-19 Jerome Detemple , Scott Robertson

We present two Policy Gradient-based algorithms with general parametrization in the context of infinite-horizon average reward Markov Decision Process (MDP). The first one employs Implicit Gradient Transport for variance reduction, ensuring…

机器学习 · 计算机科学 2025-05-13 Swetha Ganesh , Washim Uddin Mondal , Vaneet Aggarwal

The ability to learn reward functions plays an important role in enabling the deployment of intelligent agents in the real world. However, comparing reward functions, for example as a means of evaluating reward learning methods, presents a…

机器学习 · 计算机科学 2022-01-26 Blake Wulfe , Ashwin Balakrishna , Logan Ellis , Jean Mercat , Rowan McAllister , Adrien Gaidon

We obtain revenue guarantees for the simple pricing mechanism of a single posted price, in terms of a natural parameter of the distribution of buyers' valuations. Our revenue guarantee applies to the single item n buyers setting, with…

计算机科学与博弈论 · 计算机科学 2015-06-02 Balasubramanian Sivan , Vasilis Syrgkanis , Omer Tamuz

We show that a simple evolutionary scheme, when applied to the minority game (MG), changes the phase structure of the game. In this scheme each agent evolves individually whenever his wealth reaches the specified bankruptcy level, in…

统计力学 · 物理学 2009-11-10 Baosheng Yuan , Kan Chen

In the last decade quantum machine learning has provided fascinating and fundamental improvements to supervised, unsupervised and reinforcement learning. In reinforcement learning, a so-called agent is challenged to solve a task given by…

量子物理 · 物理学 2022-04-13 Arne Hamann , Sabine Wölk

People often face trade-offs between costs and benefits occurring at various points in time. The predominant discounting approach is to use the exponential form. Central to this approach is the discount rate, a unique parameter that…

理论经济学 · 经济学 2024-08-13 Bach Dong-Xuan , Philippe Bich

The interaction between an artificial agent and its environment is bi-directional. The agent extracts relevant information from the environment, and affects the environment by its actions in return to accumulate high expected reward.…

系统与控制 · 计算机科学 2018-06-06 Stas Tiomkin , Naftali Tishby

In this paper, we investigate the robustness of stationary mean-field equilibria in the presence of model uncertainties, specifically focusing on infinite-horizon discounted cost functions. To achieve this, we initially establish…

系统与控制 · 电气工程与系统科学 2026-04-10 Uğur Aydın , Naci Saldi

This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs). For MDPs that are ergodic and linear (i.e. where rewards and dynamics are linear in some known…

机器学习 · 计算机科学 2020-06-24 Nevena Lazic , Dong Yin , Mehrdad Farajtabar , Nir Levine , Dilan Gorur , Chris Harris , Dale Schuurmans

We compare the profit of the optimal third-degree price discrimination policy against a uniform pricing policy. A uniform pricing policy offers the same price to all segments of the market. Our main result establishes that for a broad class…

综合经济学 · 经济学 2021-11-16 Dirk Bergemann , Francisco Castro , Gabriel Weintraub

We analyze the asymptotic behavior for a system of fully nonlinear parabolic and elliptic quasi variational inequalities. These equations are related to robust switching control problems introduced in [3]. We prove that, as time horizon…

概率论 · 数学 2017-02-07 Erhan Bayraktar , Andrea Cosso , Huyên Pham

We begin by formulating and characterizing a dominance criterion for prize sequences: $x$ dominates $y$ if any impatient agent prefers $x$ to $y$. With this in hand, we define a notion of comparative patience. Alice is more patient than Bob…

理论经济学 · 经济学 2024-07-03 Mark Whitmeyer