English
Related papers

Related papers: The Gittins Policy in the M/G/1 Queue

200 papers

In the stochastic contextual low-rank matrix bandit problem, the expected reward of an action is given by the inner product between the action's feature matrix and some fixed, but initially unknown $d_1$ by $d_2$ matrix $\Theta^*$ with rank…

Machine Learning · Statistics 2024-01-17 Yue Kang , Cho-Jui Hsieh , Thomas C. M. Lee

We study new types of dynamic allocation problems the {\sl Halting Bandit} models. As an application, we obtain new proofs for the classic Gittins index decomposition result and recent results of the authors in `Multi-armed bandits under…

Machine Learning · Statistics 2023-04-21 Wesley Cowan , Michael N. Katehakis , Sheldon M. Ross

We consider an extremely broad class of M/G/1 scheduling policies called SOAP: Schedule Ordered by Age-based Priority. The SOAP policies include almost all scheduling policies in the literature as well as an infinite number of variants…

Performance · Computer Science 2018-02-20 Ziv Scully , Mor Harchol-Balter , Alan Scheller-Wolf

For a GI/GI/1 queue, we show that the average sojourn time under the (blind) Randomized Multilevel Feedback algorithm is no worse than that under the Shortest Remaining Processing Time algorithm times a logarithmic function of the system…

Probability · Mathematics 2017-02-06 Nikhil Bansal , Bart Kamphorst , Bert Zwart

In many bandit problems, the maximal reward achievable by a policy is often unknown in advance. We consider the problem of estimating the optimal policy value in the sublinear data regime before the optimal policy is even learnable. We…

Machine Learning · Computer Science 2023-02-21 Jonathan N. Lee , Weihao Kong , Aldo Pacchiano , Vidya Muthukumar , Emma Brunskill

This paper studies a multiclass queueing system with an associated risk- sensitive cost observed in heavy traffic at the moderate deviation scale, accounting for convex queue length penalties. The main result is the asymptotic optimality of…

Optimization and Control · Mathematics 2017-04-11 Rami Atar , Subhamay Saha

We study the stochastic multi-armed bandit problem and design new policies that enjoy both worst-case optimality for expected regret and light-tailed risk for regret distribution. Specifically, our policy design (i) enjoys the worst-case…

Machine Learning · Statistics 2024-07-23 David Simchi-Levi , Zeyu Zheng , Feng Zhu

We consider a version of the continuous-time multi-armed bandit problem where decision opportunities arrive at Poisson arrival times, and study its Gittins index policy. When driven by spectrally one-sided L\'evy processes, the Gittins…

Probability · Mathematics 2023-01-20 José-Luis Pérez , Kazutoshi Yamazaki

We consider the distributed SGD problem, where a main node distributes gradient calculations among $n$ workers. By assigning tasks to all the workers and waiting only for the $k$ fastest ones, the main node can trade-off the algorithm's…

Information Theory · Computer Science 2022-06-29 Maximilian Egger , Rawad Bitar , Antonia Wachter-Zeh , Deniz Gündüz

In this paper, we study optimal control problems for multiclass GI/M/n+M queues in an alternating renewal (up-down) random environment in the Halfin-Whitt regime. Assuming that the downtimes are asymptotically negligible and only the…

Optimization and Control · Mathematics 2019-08-20 Ari Arapostathis , Guodong Pang , Yi Zheng

We consider the infinite-horizon, average-reward restless bandit problem in discrete time. We propose a new class of policies that are designed to drive a progressively larger subset of arms toward the optimal distribution. We show that our…

Machine Learning · Computer Science 2026-03-31 Yige Hong , Qiaomin Xie , Yudong Chen , Weina Wang

This paper studies optimal switching on and o? of the entire service capacity of an M/M/Infinity queue with holding, running and switching costs where the running costs depend only on whether the system is running or not. The goal is to…

Optimization and Control · Mathematics 2015-10-21 Eugene Feinberg , Xiaoxuan Zhang

We study the impact of service-time distributions on the distribution of the maximum queue length during a busy period for the M^X/G/1 queue. The maximum queue length is an important random variable to understand when designing the buffer…

Probability · Mathematics 2007-05-23 Ger Koole , Misja Nuyens , Rhonda Righter

Time-constrained decision processes have been ubiquitous in many fundamental applications in physics, biology and computer science. Recently, restart strategies have gained significant attention for boosting the efficiency of…

Machine Learning · Computer Science 2020-07-02 Semih Cayci , Atilla Eryilmaz , R. Srikant

Policy regret is a well established notion of measuring the performance of an online learning algorithm against an adaptive adversary. We study restrictions on the adversary that enable efficient minimization of the \emph{complete policy…

Machine Learning · Statistics 2022-04-26 Dhruv Malik , Yuanzhi Li , Aarti Singh

This paper studies the deviations of the regret in a stochastic multi-armed bandit problem. When the total number of plays n is known beforehand by the agent, Audibert et al. (2009) exhibit a policy such that with probability at least…

Machine Learning · Statistics 2011-07-26 Antoine Salomon , Jean-Yves Audibert

We propose minimum empirical divergence (MED) policy for the multiarmed bandit problem. We prove asymptotic optimality of the proposed policy for the case of finite support models. In our setting, Burnetas and Katehakis has already proposed…

Statistics Theory · Mathematics 2011-11-21 Junya Honda , Akimichi Takemura

We consider the age of information in G/G/1/1 systems under two service discipline models. In the first model, if a new update arrives when the service is busy, it is blocked; in the second model, a new update preempts the current update in…

Information Theory · Computer Science 2019-06-03 Alkan Soysal , Sennur Ulukus

We study the problem of planning restless multi-armed bandits (RMABs) with multiple actions. This is a popular model for multi-agent systems with applications like multi-channel communication, monitoring and machine maintenance tasks, and…

Multiagent Systems · Computer Science 2023-03-01 Abheek Ghosh , Dheeraj Nagaraj , Manish Jain , Milind Tambe

Most of the early input-queued switch research focused on establishing throughput optimality of the max-weight scheduling policy, with some recent research showing that max-weight scheduling is optimal with respect to total expected delay…

Optimization and Control · Mathematics 2020-10-13 Yingdong Lu , Mark S. Squillante , Tonghoon Suk