中文
相关论文

相关论文: An Asymptotic Law of the Iterated Logarithm for $\…

200 篇论文

M-estimators are ubiquitous in machine learning and statistical learning theory. They are used both for defining prediction strategies and for evaluating their precision. In this paper, we propose the first non-asymptotic "any-time"…

统计理论 · 数学 2019-05-27 Victor-Emmanuel Brunel , Arnak S. Dalalyan , Nicolas Schreuder

Stochastic linear bandits are a natural and simple generalisation of finite-armed bandits with numerous practical applications. Current approaches focus on generalising existing techniques for finite-armed bandits, notably the optimism…

机器学习 · 统计学 2016-10-17 Tor Lattimore , Csaba Szepesvari

We propose the kl-UCB ++ algorithm for regret minimization in stochastic bandit models with exponential families of distributions. We prove that it is simultaneously asymptotically optimal (in the sense of Lai and Robbins' lower bound) and…

机器学习 · 统计学 2017-09-21 Pierre Ménard , Aurélien Garivier

Consider the problem of a controller sampling sequentially from a finite number of $N \geq 2$ populations, specified by random variables $X^i_k$, $ i = 1,\ldots , N,$ and $k = 1, 2, \ldots$; where $X^i_k$ denotes the outcome from population…

机器学习 · 统计学 2015-09-25 Wesley Cowan , Michael N. Katehakis

We introduce a simple and efficient algorithm for stochastic linear bandits with finitely many actions that is asymptotically optimal and (nearly) worst-case optimal in finite time. The approach is based on the frequentist…

机器学习 · 统计学 2021-07-05 Johannes Kirschner , Tor Lattimore , Claire Vernade , Csaba Szepesvári

Recent progress in reinforcement learning has led to remarkable performance in a range of applications, but its deployment in high-stakes settings remains quite rare. One reason is a limited understanding of the behavior of reinforcement…

机器学习 · 计算机科学 2020-11-04 Feicheng Wang , Lucas Janson

We study the central limit theorem in the non-normal domain of attraction to symmetric $\alpha$-stable laws for $0<\alpha\leq2$. We show that for i.i.d. random variables $X_i$, the convergence rate in $L^\infty$ of both the densities and…

概率论 · 数学 2018-04-24 Christoph Börgers , Claude Greengard

Recent studies have shown that reinforcement learning with KL-regularized objectives can enjoy faster rates of convergence or logarithmic regret, in contrast to the classical $\sqrt{T}$-type regret in the unregularized setting. However, the…

机器学习 · 计算机科学 2026-03-03 Kaixuan Ji , Qingyue Zhao , Heyang Zhao , Qiwei Di , Quanquan Gu

The article studies the almost surely asymptotics of extreme values $\bar{\xi}_n = \max_{1\leq i \leq n} \xi_i$, where $ \xi , \xi_1 , \xi_2 , \ldots$ are discrete identically distributed random variables. One of the main results on this…

概率论 · 数学 2025-03-27 Kateryna Akbash , Ivan Matsak

We propose a new analysis framework for clustering $M$ items into an unknown number of $K$ distinct groups using noisy and actively collected responses. At each time step, an agent is allowed to query pairs of items and observe bandit…

机器学习 · 计算机科学 2026-02-06 Rachel S. Y. Teo , P. N. Karthik , Ramya Korlakai Vinayak , Vincent Y. F. Tan

The law of the iterated logarithm (LIL) for the time-homogeneous Markov process with a unique invariant measure characterizes the almost sure maximum possible fluctuation of time averages around the ergodic limit. Whether a numerical…

数值分析 · 数学 2025-11-10 Chuchu Chen , Xinyu Chen , Jialin Hong

Two new test statistics are introduced to test the null hypotheses that the sampling distribution has an increasing hazard rate on a specified interval [0,a]. These statistics are empirical L_1-type distances between the isotonic estimates,…

统计理论 · 数学 2015-03-17 Piet Groeneboom , Geurt Jongbloed

We provide a framework to analyse control policies for the restless Markovian bandit model, under both finite and infinite time horizon. We show that when the population of arms goes to infinity, the value of the optimal control policy…

最优化与控制 · 数学 2023-12-25 Nicolas Gast , Bruno Gaujal , Chen Yan

We study pure exploration with infinitely many bandit arms generated i.i.d. from an unknown distribution. Our goal is to efficiently select a single high quality arm whose average reward is, with probability $1-\delta$, within $\varepsilon$…

机器学习 · 计算机科学 2023-06-06 Xiao-Yue Gong , Mark Sellke

Recent advances in Reinforcement Learning from Human Feedback (RLHF) have shown that KL-regularization plays a pivotal role in improving the efficiency of RL fine-tuning for large language models (LLMs). Despite its empirical advantage, the…

机器学习 · 计算机科学 2026-03-12 Heyang Zhao , Chenlu Ye , Wei Xiong , Quanquan Gu , Tong Zhang

We consider the \mnk{classical} problem of a controller activating (or sampling) sequentially from a finite number of $N \geq 2$ populations, specified by unknown distributions. Over some time horizon, at each time $n = 1, 2, \ldots$, the…

机器学习 · 统计学 2015-12-18 Wesley Cowan , Michael N. Katehakis

A classic setting of the stochastic K-armed bandit problem is considered in this note. In this problem it has been known that KL-UCB policy achieves the asymptotically optimal regret bound and KL-UCB+ policy empirically performs better than…

机器学习 · 计算机科学 2019-03-21 Junya Honda

The superiority of stochastic symplectic methods over non-symplectic counterparts has been verified by plenty of numerical experiments, especially in capturing the asymptotic behaviour of the underlying solution process. How can one…

数值分析 · 数学 2024-04-24 Chuchu Chen , Xinyu Chen , Tonghe Dang , Jialin Hong

Ensemble of initial conditions for nonlinear maps can be described in terms of entropy. This ensemble entropy shows an asymptotic linear growth with rate K. The rate K matches the logarithm of the corresponding asymptotic sensitivity to…

统计力学 · 物理学 2011-01-04 Massmimo Coraddu , Marcello Lissia , Roberto Tonelli

We study the asymptotic optimal control of multi-class restless bandits. A restless bandit is a controllable stochastic process whose state evolution depends on whether or not the bandit is made active. Since finding the optimal control is…

概率论 · 数学 2016-09-05 I. M. Verloop
‹ 上一页 1 2 3 10 下一页 ›