中文
相关论文

相关论文: General Discounting versus Average Reward

200 篇论文

We present an alternative methodology for the analysis of algorithms, based on the concept of expected discounted reward. This methodology naturally handles algorithms that do not always terminate, so it can (theoretically) be used with…

人工智能 · 计算机科学 2017-08-08 Andrew MacFie

We study reward-free and reward-agnostic exploration in episodic finite-horizon Markov decision processes (MDPs), where an agent explores an unknown environment without observing external rewards. Reward-free exploration aims to enable…

机器学习 · 计算机科学 2026-05-18 Oran Ridel , Alon Cohen

We study model-based reinforcement learning with non-linear function approximation where the transition function of the underlying Markov decision process (MDP) is given by a multinomial logistic (MNL) model. We develop a provably efficient…

机器学习 · 计算机科学 2024-10-15 Jaehyun Park , Junyeop Kwon , Dabeen Lee

We show that discounted methods for solving continuing reinforcement learning problems can perform significantly better if they center their rewards by subtracting out the rewards' empirical average. The improvement is substantial at…

机器学习 · 计算机科学 2024-10-31 Abhishek Naik , Yi Wan , Manan Tomar , Richard S. Sutton

Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman…

机器学习 · 统计学 2019-03-01 William Fedus , Carles Gelada , Yoshua Bengio , Marc G. Bellemare , Hugo Larochelle

We study reward-free reinforcement learning (RL) with linear function approximation, where the agent works in two phases: (1) in the exploration phase, the agent interacts with the environment but cannot access the reward; and (2) in the…

机器学习 · 计算机科学 2024-02-15 Junkai Zhang , Weitong Zhang , Quanquan Gu

We consider infinite horizon dynamic programming problems, where the control at each stage consists of several distinct decisions, each one made by one of several agents. In an earlier work we introduced a policy iteration algorithm, where…

最优化与控制 · 数学 2020-05-05 Dimitri Bertsekas

The policy improvement bound on the difference of the discounted returns plays a crucial role in the theoretical justification of the trust-region policy optimization (TRPO) algorithm. The existing bound leads to a degenerate bound when the…

机器学习 · 计算机科学 2021-07-20 J. G. Dai , Mark Gluzman

Maximum entropy reinforcement learning motivates agents to explore states and actions to maximize the entropy of some distribution, typically by providing additional intrinsic rewards proportional to that entropy function. In this paper, we…

机器学习 · 计算机科学 2026-03-20 Adrien Bolland , Gaspard Lambrechts , Damien Ernst

In the incentivized exploration model, a principal aims to explore and learn over time by interacting with a sequence of self-interested agents. It has been recently understood that the main challenge in designing incentive-compatible…

计算机科学与博弈论 · 计算机科学 2025-06-03 Benjamin Schiffer , Mark Sellke

In recent years, work has been done to develop the theory of General Reinforcement Learning (GRL). However, there are few examples demonstrating these results in a concrete way. In particular, there are no examples demonstrating the known…

人工智能 · 计算机科学 2017-03-07 Sean Lamont , John Aslanides , Jan Leike , Marcus Hutter

For two-person dynamic zero-sum games (both discrete and continuous settings), we investigate the limit of value functions of finite horizon games with long run average cost as the time horizon tends to infinity and the limit of value…

最优化与控制 · 数学 2017-09-26 Dmitry Khlopin

We study infinite-horizon Constrained Markov Decision Processes (CMDPs) with general policy parameterizations and multi-layer neural network critics. Existing theoretical analyses for constrained reinforcement learning largely rely on…

机器学习 · 计算机科学 2026-03-10 Anirudh Satheesh , Pankaj Kumar Barman , Washim Uddin Mondal , Vaneet Aggarwal

We provide an original theoretical study of Inverse Reinforcement Learning (IRL) through the lens of reward compatibility, a novel framework to quantify the compatibility of a reward with the given expert's demonstrations. Intuitively, a…

机器学习 · 计算机科学 2025-01-15 Filippo Lazzati , Mirco Mutti , Alberto Metelli

The uniform distribution is an important counterexample in game theory as many of the canonical game dynamics have been shown not to converge to the equilibrium in certain cases. In particular none of the canonical game dynamics converge to…

计算机科学与博弈论 · 计算机科学 2012-07-03 Dashiell E. A. Fryer

We study the design of optimal incentives in sequential processes. To do so, we consider a basic and fundamental model in which an agent initiates a value-creating sequential process through costly investment with random success. If…

理论经济学 · 经济学 2023-11-22 Jens Gudmundsson , Jens Leth Hougaard , Juan D. Moreno-Ternero , Lars Peter Østerdal

Motivated by applications in online marketplaces such as ride-hailing platforms and payment channel networks, we study a single-server queue with state-dependent arrival control. The service operator dynamically chooses the arrival rate as…

最优化与控制 · 数学 2026-02-03 Tianze Qu , Sushil Mahavir Varma

We study mechanisms for an allocation of goods among agents, where agents have no incentive to lie about their true values (incentive compatible) and for which no agent will seek to exchange outcomes with another (envy-free). Mechanisms…

计算机科学与博弈论 · 计算机科学 2010-03-30 Edith Cohen , Michal Feldman , Amos Fiat , Haim Kaplan , Svetlana Olonetsky

Algorithms developed under stationary Markov Decision Processes (MDPs) often face challenges in non-stationary environments, and infinite-horizon formulations may not directly apply to finite-horizon tasks. To address these limitations, we…

机器学习 · 计算机科学 2025-12-03 Zhizuo Chen , Theodore T. Allen

How much should you receive in a week to be indifferent to \$ 100 in six months? Note that the indifference requires a rule to ensure the similarity between early and late payments. Assuming that rational individuals have low accuracy, then…

综合经济学 · 经济学 2020-03-02 José Cláudio do Nascimento