中文
相关论文

相关论文: A Unified Analysis of Nonstochastic Delayed Feedba…

200 篇论文

We propose a new algorithm for adversarial multi-armed bandits with unrestricted delays. The algorithm is based on a novel hybrid regularizer applied in the Follow the Regularized Leader (FTRL) framework. It achieves…

机器学习 · 计算机科学 2020-06-17 Julian Zimmert , Yevgeny Seldin

We consider the nonstochastic multi-agent multi-armed bandit problem with agents collaborating via a communication network with delays. We show a lower bound for individual regret of all agents. We show that with suitable regularizers and…

机器学习 · 统计学 2023-10-24 Jialin Yi , Milan Vojnović

The linear bandit problem has been studied for many years in both stochastic and adversarial settings. Designing an algorithm that can optimize the environment without knowing the loss type attracts lots of interest. \citet{LeeLWZ021}…

机器学习 · 计算机科学 2023-07-19 Fang Kong , Canzhe Zhao , Shuai Li

We consider a common case of the combinatorial semi-bandit problem, the $m$-set semi-bandit, where the learner exactly selects $m$ arms from the total $d$ arms. In the adversarial setting, the best regret bound, known to be…

机器学习 · 计算机科学 2025-07-08 Jingxin Zhan , Yuchen Xin , Chenjie Sun , Zhihua Zhang

We consider the problem of online combinatorial optimization under semi-bandit feedback. The goal of the learner is to sequentially select its actions from a combinatorial decision set so as to minimize its cumulative loss. We propose a…

机器学习 · 计算机科学 2013-05-14 Gergely Neu , Gábor Bartók

We consider regret minimization for Adversarial Markov Decision Processes (AMDPs), where the loss functions are changing over time and adversarially chosen, and the learner only observes the losses for the visited state-action pairs (i.e.,…

机器学习 · 计算机科学 2022-09-20 Yan Dai , Haipeng Luo , Liyu Chen

The stochastic generalised linear bandit is a well-understood model for sequential decision-making problems, with many algorithms achieving near-optimal regret guarantees under immediate feedback. However, the stringent requirement for…

机器学习 · 计算机科学 2023-04-12 Benjamin Howson , Ciara Pike-Burke , Sarah Filippi

This paper studies the optimality and complexity of Follow-the-Perturbed-Leader (FTPL) policy in size-invariant combinatorial semi-bandit problems. Recently, Honda et al. (2023) and Lee et al. (2024) showed that FTPL achieves…

机器学习 · 计算机科学 2025-07-23 Botao Chen , Junya Honda

We investigate a nonstochastic bandit setting in which the loss of an action is not immediately charged to the player, but rather spread over the subsequent rounds in an adversarial way. The instantaneous loss observed by the player at the…

机器学习 · 计算机科学 2022-09-27 Nicolò Cesa-Bianchi , Tommaso Cesari , Roberto Colomboni , Claudio Gentile , Yishay Mansour

Follow-The-Regularized-Leader (FTRL) is known as an effective and versatile approach in online learning, where appropriate choice of the learning rate is crucial for smaller regret. To this end, we formulate the problem of adjusting FTRL's…

机器学习 · 计算机科学 2024-03-12 Shinji Ito , Taira Tsuchiya , Junya Honda

Follow-the-Regularized-Leader (FTRL) algorithms are a popular class of learning algorithms for online linear optimization (OLO) that guarantee sub-linear regret, but the choice of regularizer can significantly impact dimension-dependent…

机器学习 · 计算机科学 2024-10-24 Khashayar Gatmiry , Jon Schneider , Stefanie Jegelka

We study the problem of online generalized linear regression in the stochastic setting, where the label is generated from a generalized linear model with possibly unbounded additive noise. We provide a sharp analysis of the classical…

机器学习 · 计算机科学 2023-03-28 Heyang Zhao , Dongruo Zhou , Jiafan He , Quanquan Gu

This paper investigates the problem of combinatorial multiarmed bandits with stochastic submodular (in expectation) rewards and full-bandit delayed feedback, where the delayed feedback is assumed to be composite and anonymous. In other…

机器学习 · 计算机科学 2025-01-23 Mohammad Pedramfar , Vaneet Aggarwal

This paper studies the optimality and complexity of Follow-the-Perturbed-Leader (FTPL) policy in $m$-set semi-bandit problems. FTPL has been studied extensively as a promising candidate of an efficient algorithm with favorable regret for…

机器学习 · 计算机科学 2026-03-13 Botao Chen , Jongyeong Lee , Chansoo Kim , Junya Honda

We present tools for the analysis of Follow-The-Regularized-Leader (FTRL), Dual Averaging, and Mirror Descent algorithms when the regularizer (equivalently, prox-function or learning rate schedule) is chosen adaptively based on the data.…

机器学习 · 计算机科学 2015-11-10 H. Brendan McMahan

Follow-The-Regularized-Leader (FTRL) algorithms often enjoy optimal regret for adversarial as well as stochastic bandit problems and allow for a streamlined analysis. Nonetheless, FTRL algorithms require the solution of an optimization…

机器学习 · 计算机科学 2025-02-14 Mengmeng Li , Daniel Kuhn , Bahar Taşkesen

The standard assumption in reinforcement learning (RL) is that agents observe feedback for their actions immediately. However, in practice feedback is often observed in delay. This paper studies online learning in episodic Markov decision…

机器学习 · 计算机科学 2023-01-24 Tiancheng Jin , Tal Lancewicki , Haipeng Luo , Yishay Mansour , Aviv Rosenberg

We study online convex optimization on $\ell_p$-balls in $\mathbb{R}^d$ for $p > 2$. While always sub-linear, the optimal regret exhibits a shift between the high-dimensional setting ($d > T$), when the dimension $d$ is greater than the…

机器学习 · 计算机科学 2025-12-01 Emmeran Johnson , David Martínez-Rubio , Ciara Pike-Burke , Patrick Rebeschini

We introduce a novel online learning framework that unifies and generalizes pre-established models, such as delayed and corrupted feedback, to encompass adversarial environments where action feedback evolves over time. In this setting, the…

机器学习 · 计算机科学 2024-05-28 Yogev Bar-On , Yishay Mansour

We consider the problem of asynchronous online combinatorial optimization on a network of communicating agents. At each time step, some of the agents are stochastically activated, requested to make a prediction, and the system pays the…

机器学习 · 计算机科学 2021-02-10 Riccardo Della Vecchia , Tommaso Cesari
‹ 上一页 1 2 3 10 下一页 ›