中文
相关论文

相关论文: Continuous-in-time Limit for Bayesian Bandits

200 篇论文

We investigate the adversarial bandit problem with multiple plays under semi-bandit feedback. We introduce a highly efficient algorithm that asymptotically achieves the performance of the best switching $m$-arm strategy with minimax optimal…

机器学习 · 计算机科学 2019-12-02 N. Mert Vural , Hakan Gokcesu , Kaan Gokcesu , Suleyman S. Kozat

This paper, which is the natural continuation of a previous paper by the same authors, studies a class of optimal control problems with state constraints where the state equation is a differential equation with delays. This class includes…

最优化与控制 · 数学 2009-07-10 Salvatore Federico , Ben Goldys , Fausto Gozzi

This paper is devoted to regret lower bounds in the classical model of stochastic multi-armed bandit. A well-known result of Lai and Robbins, which has then been extended by Burnetas and Katehakis, has established the presence of a…

机器学习 · 统计学 2011-12-19 Antoine Salomon , Jean-Yves Audibert , Issam El Alaoui

Bayesian bandit algorithms with approximate Bayesian inference have been widely used in real-world applications. However, there is a large discrepancy between the superior practical performance of these approaches and their theoretical…

机器学习 · 计算机科学 2023-11-13 Ziyi Huang , Henry Lam , Amirhossein Meisami , Haofeng Zhang

We consider a stochastic multi-armed bandit problem with i.i.d. rewards where the expected reward function is multimodal with at most m modes. We propose the first known computationally tractable algorithm for computing the solution to the…

机器学习 · 统计学 2025-10-31 William Réveillard , Richard Combes

A new framework for formulating reachability problems with competing inputs, nonlinear dynamics and state constraints as optimal control problems is developed. Such reach-avoid problems arise in, among others, the study of safety problems…

最优化与控制 · 数学 2009-11-25 Kostas Margellos , John Lygeros

The paper concerns the infinite dimensional Hamilton-Jacobi-Bellman equation related to optimal control problem regulated by a transport equation with boundary control. A suitable viscosity solution approach is needed in view of the…

最优化与控制 · 数学 2007-05-23 Giorgio Fabbri

We develop a new approach to obtaining high probability regret bounds for online learning with bandit feedback against an adaptive adversary. While existing approaches all require carefully constructing optimistic and biased loss…

机器学习 · 计算机科学 2020-11-02 Chung-Wei Lee , Haipeng Luo , Chen-Yu Wei , Mengxiao Zhang

This paper investigates the convergence properties of the upwind difference scheme for the Hamilton--Jacobi--Bellman (HJB) equation, a central partial differential equation in optimal control theory. First, assuming the existence of a…

数值分析 · 数学 2026-02-05 Daisuke Inoue , Yuji Ito , Takahito Kashiwabara , Norikazu Saito , Hiroaki Yoshida

This paper studies the time-inconsistent MV optimal stopping problem via a game-theoretic approach to find equilibrium strategies. To overcome the mathematical intractability of direct equilibrium analysis, we propose a vanishing…

最优化与控制 · 数学 2025-10-29 Yuchao Dong , Harry Zheng

Taking advantage of contextual information can potentially boost the performance of recommender systems. In the era of big data, such side information often has several dimensions. Thus, developing decision-making algorithms to cope with…

机器学习 · 计算机科学 2023-07-26 Saeed Ghoorchian , Evgenii Kortukov , Setareh Maghsudi

We investigate feedback control for infinite horizon optimal control problems for partial differential equations. The method is based on the coupling between Hamilton-Jacobi-Bellman (HJB) equations and model reduction techniques. It is…

最优化与控制 · 数学 2016-07-11 Alessandro Alla , Andreas Schmidt , Bernard Haasdonk

This paper investigates stochastic and adversarial combinatorial multi-armed bandit problems. In the stochastic setting under semi-bandit feedback, we derive a problem-specific regret lower bound, and discuss its scaling with the dimension…

机器学习 · 计算机科学 2015-11-09 Richard Combes , M. Sadegh Talebi , Alexandre Proutiere , Marc Lelarge

In this paper we consider the problem of learning the optimal policy for uncontrolled restless bandit problems. In an uncontrolled restless bandit problem, there is a finite set of arms, each of which when pulled yields a positive reward.…

最优化与控制 · 数学 2015-01-30 Cem Tekin , Mingyan Liu

This paper studies an optimal dividend problem with a drawdown constraint in a Brownian motion model, requiring the dividend payout rate to remain above a fixed proportion of its historical maximum. This leads to a path-dependent stochastic…

数理金融 · 定量金融 2026-01-08 Chonghu Guan , Jiacheng Fan , Zuo Quan Xu

Multi-armed bandit problems are considered as a paradigm of the trade-off between exploring the environment to find profitable actions and exploiting what is already known. In the stationary case, the distributions of the rewards do not…

统计理论 · 数学 2008-12-18 Aurélien Garivier , Eric Moulines

We revisit the classic regret-minimization problem in the stochastic multi-armed bandit setting when the arm-distributions are allowed to be heavy-tailed. Regret minimization has been well studied in simpler settings of either bounded…

机器学习 · 计算机科学 2021-02-09 Shubhada Agrawal , Sandeep Juneja , Wouter M. Koolen

A general time-inconsistent optimal control problem is considered for stochastic differential equations with deterministic coefficients. Under suitable conditions, a Hamilton-Jacobi-Bellman type equation is derived for the equilibrium value…

最优化与控制 · 数学 2012-04-04 Jiongmin Yong

We consider a sequential decision-making problem where an agent can take one action at a time and each action has a stochastic temporal extent, i.e., a new action cannot be taken until the previous one is finished. Upon completion, the…

机器学习 · 计算机科学 2020-03-26 P Sharoff , Nishant A. Mehta , Ravi Ganti

Exploration policies in Bayesian bandits maximize the average reward over problem instances drawn from some distribution $\mathcal{P}$. In this work, we learn such policies for an unknown distribution $\mathcal{P}$ using samples from…

机器学习 · 计算机科学 2020-06-11 Craig Boutilier , Chih-Wei Hsu , Branislav Kveton , Martin Mladenov , Csaba Szepesvari , Manzil Zaheer