English
Related papers

Related papers: Optimal Epidemic Control as a Contextual Combinato…

200 papers

We study the problem of stochastic contextual bandits in the agnostic setting, where the goal is to compete with the best policy in a given class without assuming realizability or imposing model restrictions on losses or rewards. In this…

Machine Learning · Statistics 2026-04-06 Samuel Girard , Aurelien Bibaut , Arthur Gretton , Nathan Kallus , Houssam Zenati

We develop asymptotically optimal policies for the multi armed bandit (MAB), problem, under a cost constraint. This model is applicable in situations where each sample (or activation) from a population (bandit) incurs a known bandit…

Machine Learning · Statistics 2015-12-18 Apostolos N. Burnetas , Odysseas Kanavetas , Michael N. Katehakis

Algorithms for hyperparameter optimization abound, all of which work well under different and often unverifiable assumptions. Motivated by the general challenge of sequentially choosing which algorithm to use, we study the more specific…

Machine Learning · Statistics 2016-04-12 Robert Nishihara , David Lopez-Paz , Léon Bottou

In a conventional contextual multi-armed bandit problem, the feedback (or reward) is immediately observable after an action. Nevertheless, delayed feedback arises in numerous real-life situations and is particularly crucial in…

Machine Learning · Computer Science 2024-05-21 Kweiguu Liu , Setareh Maghsudi

In this paper we consider the adversarial contextual bandit problem in metric spaces. The paper "Nearest neighbour with bandit feedback" tackled this problem but when there are many contexts near the decision boundary of the comparator…

Machine Learning · Computer Science 2023-12-18 Stephen Pasteris , Chris Hicks , Vasilios Mavroudis

Compartmental models have long served as important tools in mathematical epidemiology, with their usefulness highlighted by the recent COVID-19 pandemic. However, most of the classical models fail to account for certain features of this…

Dynamical Systems · Mathematics 2023-08-21 Monique Chyba , Taylor Klotz , Yuriy Mileyko , Corey Shanbrom

Balancing common disease treatment and epidemic control is a key objective of medical supplies procurement in hospitals during a pandemic such as COVID-19. This problem can be formulated as a bi-objective optimization problem for…

Neural and Evolutionary Computing · Computer Science 2020-08-04 Yu-Jun Zheng , Xin Chen , Tie-Er Gan , Min-Xia Zhang , Wei-Guo Sheng , Ling Wang

We introduce the "inverse bandit" problem of estimating the rewards of a multi-armed bandit instance from observing the learning process of a low-regret demonstrator. Existing approaches to the related problem of inverse reinforcement…

Machine Learning · Statistics 2022-02-23 Wenshuo Guo , Kumar Krishna Agrawal , Aditya Grover , Vidya Muthukumar , Ashwin Pananjady

In recent years, preference-based human feedback mechanisms have become essential for enhancing model performance across diverse applications, including conversational AI systems such as ChatGPT. However, existing approaches often neglect…

Artificial Intelligence · Computer Science 2025-02-14 Raihan Seraj , Lili Meng , Tristan Sylvain

Motivated by practical needs such as large-scale learning, we study the impact of adaptivity constraints to linear contextual bandits, a central problem in online active learning. We consider two popular limited adaptivity models in…

Machine Learning · Computer Science 2021-04-26 Yufei Ruan , Jiaqi Yang , Yuan Zhou

This paper investigates stochastic and adversarial combinatorial multi-armed bandit problems. In the stochastic setting under semi-bandit feedback, we derive a problem-specific regret lower bound, and discuss its scaling with the dimension…

Machine Learning · Computer Science 2015-11-09 Richard Combes , M. Sadegh Talebi , Alexandre Proutiere , Marc Lelarge

Multi-armed bandit problems are the predominant theoretical model of exploration-exploitation tradeoffs in learning, and they have countless applications ranging from medical trials, to communication networks, to Web search and advertising.…

Data Structures and Algorithms · Computer Science 2017-09-06 Ashwinkumar Badanidiyuru , Robert Kleinberg , Aleksandrs Slivkins

We consider contextual bandit learning under distribution shift when reward vectors are ordered according to a given preference cone. We propose an adaptive-discretization and optimistic elimination based policy that self-tunes to the…

Machine Learning · Computer Science 2025-08-25 Apurv Shukla , P. R. Kumar

Several non-pharmaceutical interventions have been proposed to control the spread of the COVID-19 pandemic. On the large scale, these empirical solutions, often associated with extended and complete lockdowns, attempt to minimize the costs…

Physics and Society · Physics 2020-07-23 M. Serra , S. al-Mosleh , S. Ganga Prasath , V. Raju , S. Mantena , J. Chandra , S. Iams , L. Mahadevan

Bandits with feedback graphs are powerful online learning models that interpolate between the full information and classic bandit problems, capturing many real-life applications. A recent work by Zhang et al. (2023) studies the contextual…

Machine Learning · Computer Science 2024-02-14 Mengxiao Zhang , Yuheng Zhang , Haipeng Luo , Paul Mineiro

While sequential task assignment for a single agent has been widely studied, such problems in a multi-agent setting, where the agents have heterogeneous task preferences or capabilities, remain less well-characterized. We study a…

Multiagent Systems · Computer Science 2025-10-21 Qinshuang Wei , Vaibhav Srivastava , Vijay Gupta

This paper presents a new contextual bandit algorithm, NeuralBandit, which does not need hypothesis on stationarity of contexts and rewards. Several neural networks are trained to modelize the value of rewards knowing the context. Two…

Neural and Evolutionary Computing · Computer Science 2014-09-30 Robin Allesiardo , Raphael Feraud , Djallel Bouneffouf

We consider the problem of stochastic $K$-armed dueling bandit in the contextual setting, where at each round the learner is presented with a context set of $K$ items, each represented by a $d$-dimensional feature vector, and the goal of…

Machine Learning · Computer Science 2021-05-11 Aadirupa Saha , Aditya Gopalan

Controlling antenna tilts in cellular networks is imperative to reach an efficient trade-off between network coverage and capacity. In this paper, we devise algorithms learning optimal tilt control policies from existing data (in the…

Machine Learning · Computer Science 2022-01-07 Filippo Vannella , Alexandre Proutiere , Yassir Jedra , Jaeseong Jeong

The spread of COVID-19 has been thwarted in most countries through non-pharmaceutical interventions. In particular, the most effective measures in this direction have been the stay-at-home and closure strategies of businesses and schools.…

Physics and Society · Physics 2022-05-04 G. Dimarco , G. Toscani , M. Zanella
‹ Prev 1 8 9 10 Next ›