中文
相关论文

相关论文: Monte Carlo Rollout Policy for Recommendation Syst…

200 篇论文

Restless multi-armed bandits with partially observable states has applications in communication systems, age of information and recommendation systems. In this paper, we study multi-state partially observable restless bandit models. We…

机器学习 · 计算机科学 2021-08-03 Rahul Meshram , Kesav Kaza

We consider multi-dimensional Markov decision processes and formulate a long term discounted reward optimization problem. Two simulation based algorithms---Monte Carlo rollout policy and parallel rollout policy are studied, and various…

系统与控制 · 电气工程与系统科学 2020-07-28 Rahul Meshram , Kesav Kaza

The problem of rested and restless multi-armed bandits with constrained availability of arms is considered. The states of arms evolve in Markovian manner and the exact states are hidden from the decision maker. First, some structural…

系统与控制 · 计算机科学 2017-10-20 Varun Mehta , Rahul Meshram , Kesav Kaza , S. N. Merchant

This paper investigates methods for estimating the optimal stochastic control policy for a Markov Decision Process with unknown transition dynamics and an unknown reward function. This form of model-free reinforcement learning comprises…

机器学习 · 计算机科学 2019-12-06 Brandon Trabucco , Albert Qu , Simon Li , Ganeshkumar Ashokavardhanan

Multi-step agentic reinforcement learning benefits from fine-grained credit assignment, yet existing approaches offer limited options: critic-free methods like GRPO assign a uniform advantage to every action in a trajectory, while learned…

机器学习 · 计算机科学 2026-04-14 Tao Wang , Suhang Zheng , Xiaoxiao Xu

We consider the scheduling problem concerning N projects. Each project evolves as a multi-state Markov process. At each time instant, one project is scheduled to work, and some reward depending on the state of the chosen project is…

最优化与控制 · 数学 2016-02-02 Kehao Wang

We consider a Markov decision process with deterministic state transition dynamics, adversarially generated rewards that change arbitrarily from round to round, and a bandit feedback model in which the decision maker only observes the…

计算机科学与博弈论 · 计算机科学 2012-10-19 Raman Arora , Ofer Dekel , Ambuj Tewari

Maneuver decision-making can be regarded as a Markov decision process and can be address by reinforcement learning. However, original reinforcement learning algorithms can hardly solve the maneuvering decision-making problem. One reason is…

人工智能 · 计算机科学 2023-09-19 Zhang Hong-Peng

Traditionally, when recommender systems are formalized as multi-armed bandits, the policy of the recommender system influences the rewards accrued, but not the length of interaction. However, in real-world systems, dissatisfied users may…

机器学习 · 计算机科学 2024-02-19 Omer Ben-Porat , Lee Cohen , Liu Leqi , Zachary C. Lipton , Yishay Mansour

Policy-guided Monte Carlo is an adaptive method to simulate classical interacting systems. It adjusts the proposal distribution of the Metropolis-Hastings algorithm to maximize the sampling efficiency, using a formalism inspired by…

软凝聚态物质 · 物理学 2024-08-23 Leonardo Galliano , Riccardo Rende , Daniele Coslovich

In this paper we propose a flexible and efficient framework for handling multi-armed bandits, combining sequential Monte Carlo algorithms with hierarchical Bayesian modeling techniques. The framework naturally encompasses restless bandits,…

机器学习 · 统计学 2013-10-08 Michael Cherkassky , Luke Bornn

We consider the restless multi-armed bandit (RMAB) problem with unknown dynamics in which a player chooses M out of N arms to play at each time. The reward state of each arm transits according to an unknown Markovian rule when it is played…

最优化与控制 · 数学 2011-12-30 Haoyang Liu , Keqin Liu , Qing Zhao

A restless multi-armed bandit problem that arises in multichannel opportunistic communications is considered, where channels are modeled as independent and identical Gilbert-Elliot channels and channel state observations are subject to…

网络与互联网体系结构 · 计算机科学 2008-11-13 Qing Zhao , Bhaskar Krishnamachari

We consider the classical multi-armed bandit problem with Markovian rewards. When played an arm changes its state in a Markovian fashion while it remains frozen when not played. The player receives a state-dependent reward each time it…

最优化与控制 · 数学 2022-11-15 Cem Tekin , Mingyan Liu

We consider the channel access problem in a multi-channel opportunistic communication system with imperfect channel sensing, where the state of each channel evolves as a non independent and identically distributed Markov process. This…

系统与控制 · 计算机科学 2015-06-05 Kehao Wang , Lin Chen , Quan Liu , Khaldoun Al Agha

Most reinforcement learning practitioners evaluate their policies with online Monte Carlo estimators for either hyperparameter tuning or testing different algorithmic design choices, where the policy is repeatedly executed in the…

机器学习 · 计算机科学 2024-10-03 Shuze Liu , Shangtong Zhang

We consider finite state restless multi-armed bandit problem. The decision maker can act on M bandits out of N bandits in each time step. The play of arm (active arm) yields state dependent rewards based on action and when the arm is not…

机器学习 · 计算机科学 2023-05-02 Vishesh Mittal , Rahul Meshram , Deepak Dev , Surya Prakash

The early sections of this paper present an analysis of a Markov decision model that is known as the multi-armed bandit under the assumption that the utility function of the decision maker is either linear or exponential. The analysis…

最优化与控制 · 数学 2012-03-22 Eric V. Denardo , Eugene A. Feinberg , Uriel G. Rothblum

We present a Monte-Carlo simulation algorithm for real-time policy improvement of an adaptive controller. In the Monte-Carlo simulation, the long-term expected reward of each possible action is statistically measured, using the initial…

机器学习 · 计算机科学 2025-04-07 Gerald Tesauro , Gregory R. Galperin

We describe and study a model for an Automated Online Recommendation System (AORS) in which a user's preferences can be time-dependent and can also depend on the history of past recommendations and play-outs. The three key features of the…

机器学习 · 计算机科学 2016-03-31 Rahul Meshram , Aditya Gopalan , D. Manjunath
‹ 上一页 1 2 3 10 下一页 ›