中文
相关论文

相关论文: General Discounting versus Average Reward

200 篇论文

What are the functionals of the reward that can be computed and optimized exactly in Markov Decision Processes?In the finite-horizon, undiscounted setting, Dynamic Programming (DP) can only handle these operations efficiently for certain…

人工智能 · 计算机科学 2024-02-20 Alexandre Marthe , Aurélien Garivier , Claire Vernade

We consider a dynamic version of sender-receiver games, where the sequence of states follows an irreducible Markov chain observed by the sender. Under mild assumptions, we provide a simple characterization of the limit set of equilibrium…

概率论 · 数学 2012-04-03 Jerome Renault , Eilon Solan , Nicolas Vieille

We study infinite horizon control of continuous-time non-linear branching processes with almost sure extinction for general (positive or negative) discount. Our main goal is to study the link between infinite horizon control of these…

概率论 · 数学 2016-07-28 Julien Claisse , Nicolas Champagnat

This paper examines the impact of agents' myopic optimization on the efficiency of systems comprised by many selfish agents. In contrast to standard congestion games where agents interact in a one-shot fashion, in our model each agent…

计算机科学与博弈论 · 计算机科学 2025-04-30 Yunpeng Li , Antonis Dimakis , Costas A. Courcoubetis

We investigate the survivor distributions of a spatially extended model of competitive dynamics in different geometries. The model consists of a deterministic dynamical system of individual agents at specified nodes, which might or might…

定量方法 · 定量生物学 2015-11-24 J. M. Luck , A. Mehta

We present a simple game model where agents with different memory lengths compete for finite resources. We show by simulation and analytically that an instability exists at a critical memory length, and as a result, different memory lengths…

适应与自组织系统 · 物理学 2015-05-12 James Burridge , Yu Gao , Yong Mao

A basic assumption of traditional reinforcement learning is that the value of a reward does not change once it is received by an agent. The present work forgoes this assumption and considers the situation where the value of a reward decays…

人工智能 · 计算机科学 2023-03-01 Taylor Dohmen , Ashutosh Trivedi

Principal-agent problems arise when one party acts on behalf of another, leading to conflicts of interest. The economic literature has extensively studied principal-agent problems, and recent work has extended this to more complex scenarios…

人工智能 · 计算机科学 2024-01-02 Omer Ben-Porat , Yishay Mansour , Michal Moshkovitz , Boaz Taitler

Deep reinforcement learning (RL) works impressively in some environments and fails catastrophically in others. Ideally, RL theory should be able to provide an understanding of why this is, i.e. bounds predictive of practical performance.…

机器学习 · 计算机科学 2024-01-15 Cassidy Laidlaw , Stuart Russell , Anca Dragan

Commonly in reinforcement learning (RL), rewards are discounted over time using an exponential function to model time preference, thereby bounding the expected long-term reward. In contrast, in economics and psychology, it has been shown…

机器学习 · 计算机科学 2022-12-08 Matthias Schultheis , Constantin A. Rothkopf , Heinz Koeppl

This note re-visits the rolling-horizon control approach to the problem of a Markov decision process (MDP) with infinite-horizon discounted expected reward criterion. Distinguished from the classical value-iteration approach, we develop an…

最优化与控制 · 数学 2022-06-07 Hyeong Soo Chang

We study the Improving Multi-Armed Bandit (IMAB) problem, where the reward obtained from an arm increases with the number of pulls it receives. This model provides an elegant abstraction for many real-world problems in domains such as…

机器学习 · 计算机科学 2022-08-22 Vishakha Patil , Vineet Nair , Ganesh Ghalme , Arindam Khan

Natural language is an intuitive and expressive way to communicate reward information to autonomous agents. It encompasses everything from concrete instructions to abstract descriptions of the world. Despite this, natural language is often…

人工智能 · 计算机科学 2022-04-12 Theodore R. Sumers , Robert D. Hawkins , Mark K. Ho , Thomas L. Griffiths , Dylan Hadfield-Menell

In this work, the dynamics of agents below a \textit{threshold line} in some modified CCM type kinetic wealth exchange models are studied. These agents are eligible for subsidy as can be seen in any real economy. An interaction is…

物理与社会 · 物理学 2022-06-29 Sanchari Goswami

We investigate whether fairness is compatible with efficiency in economies with multi-self agents, who may not be able to integrate their multiple objectives into a single complete and transitive ranking. We adapt envy-freeness,…

理论经济学 · 经济学 2022-04-15 Sophie Bade , Erel Segal-Halevi

This paper is the second of two papers devoted to the study of the evolution of the cosmological horizons (particle and event horizons). Specifically, in this paper we consider the extremely general case of an accelerated universe with…

宇宙学与河外天体物理 · 物理学 2015-09-28 Berta Margalef-Bentabol , Juan Margalef-Bentabol , Jordi Cepa

We use the Minority Game as a testing frame for the problem of the emergence of diversity in socio-economic systems. For the MG with heterogeneous impacts, we show that the direct generalization of the usual agents' profit does not fit some…

交易与市场微观结构 · 定量金融 2014-01-20 Miroslav Pištěk , Frantisek Slanina

Under non-exponential discounting, we develop a dynamic theory for stopping problems in continuous time. Our framework covers discount functions that induce decreasing impatience. Due to the inherent time inconsistency, we look for…

最优化与控制 · 数学 2017-03-13 Yu-Jui Huang , Adrien Nguyen-Huu

In fair division of indivisible goods, using sequences of sincere choices (or picking sequences) is a natural way to allocate the objects. The idea is as follows: at each stage, a designated agent picks one object among those that remain.…

人工智能 · 计算机科学 2018-08-01 Aurélie Beynier , Sylvain Bouveret , Michel Lemaître , Nicolas Maudet , Simon Rey

In mean-payoff games, the objective of the protagonist is to ensure that the limit average of an infinite sequence of numeric weights is nonnegative. In energy games, the objective is to ensure that the running sum of weights is always…

计算机科学中的逻辑 · 计算机科学 2010-10-05 Krishnendu Chatterjee , Laurent Doyen , Thomas A. Henzinger , Jean-Francois Raskin
‹ 上一页 1 8 9 10 下一页 ›