中文
相关论文

相关论文: Delightful Exploration

200 篇论文

We address online learning in complex auction settings, such as sponsored search auctions, where the value of the bidder is unknown to her, evolving in an arbitrary manner and observed only if the bidder wins an allocation. We leverage the…

计算机科学与博弈论 · 计算机科学 2018-06-04 Zhe Feng , Chara Podimata , Vasilis Syrgkanis

Based on differential privacy (DP) framework, we introduce and unify privacy definitions for the multi-armed bandit algorithms. We represent the framework with a unified graphical model and use it to connect privacy definitions. We derive…

机器学习 · 计算机科学 2020-06-25 Debabrota Basu , Christos Dimitrakakis , Aristide Tossou

We extend the model of Multi-armed Bandit with unit switching cost to incorporate a metric between the actions. We consider the case where the metric over the actions can be modeled by a complete binary tree, and the distance between two…

机器学习 · 计算机科学 2017-02-27 Tomer Koren , Roi Livni , Yishay Mansour

We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maximum entropy exploration problem of two different types. The…

We discuss the relative merits of optimistic and randomized approaches to exploration in reinforcement learning. Optimistic approaches presented in the literature apply an optimistic boost to the value estimate at each state-action pair and…

机器学习 · 统计学 2017-06-15 Ian Osband , Benjamin Van Roy

In this paper, we consider the low rank structure of the reward sequence of the pure exploration problems. Firstly, we propose the separated setting in pure exploration problem, where the exploration strategy cannot receive the feedback of…

机器学习 · 计算机科学 2023-06-29 Yaxiong Liu , Atsuyoshi Nakamura , Kohei Hatano , Eiji Takimoto

Model-based reinforcement learning algorithms with probabilistic dynamical models are amongst the most data-efficient learning methods. This is often attributed to their ability to distinguish between epistemic and aleatoric uncertainty.…

机器学习 · 计算机科学 2020-12-02 Sebastian Curi , Felix Berkenkamp , Andreas Krause

We consider the problem of online fair division of indivisible goods to players when there are a finite number of types of goods and player values are drawn from distributions with unknown means. Our goal is to maximize social welfare…

计算机科学与博弈论 · 计算机科学 2024-12-10 Ariel D. Procaccia , Benjamin Schiffer , Shirley Zhang

We consider online variations of the Pandora's box problem (Weitzman. 1979), a standard model for understanding issues related to the cost of acquiring information for decision-making. Our problem generalizes both the classic Pandora's box…

数据结构与算法 · 计算机科学 2019-01-31 Hossein Esfandiari , MohammadTaghi Hajiaghayi , Brendan Lucier , Michael Mitzenmacher

We study adversarial multi-armed bandits with and without delayed feedback under a safety-aware goal: achieving minimax-optimal worst-case regret while keeping nearly constant regret relative to a designated "safe" baseline policy. Existing…

机器学习 · 计算机科学 2026-05-25 Ting Hu , Luanda Cai , Emmanouil-Vasileios Vlatakis-Gkaragkounis

Consider the sequential optimization of a continuous, possibly non-convex, and expensive to evaluate objective function $f$. The problem can be cast as a Gaussian Process (GP) bandit where $f$ lives in a reproducing kernel Hilbert space…

机器学习 · 统计学 2021-08-23 Sattar Vakili , Nacime Bouziani , Sepehr Jalali , Alberto Bernacchia , Da-shan Shiu

The Pandora's Box Problem, originally formalized by Weitzman in 1979, models selection from set of random, alternative options, when evaluation is costly. This includes, for example, the problem of hiring a skilled worker, where only one…

计算机科学与博弈论 · 计算机科学 2024-02-20 Shant Boodaghians , Federico Fusco , Philip Lazos , Stefano Leonardi

We consider a multi-armed bandit problem where the decision maker can explore and exploit different arms at every round. The exploited arm adds to the decision maker's cumulative reward (without necessarily observing the reward) while the…

机器学习 · 计算机科学 2012-07-03 Orly Avner , Shie Mannor , Ohad Shamir

We motivate and analyse a new Tree Search algorithm, GPTS, based on recent theoretical advances in the use of Gaussian Processes for Bandit problems. We consider tree paths as arms and we assume the target/reward function is drawn from a GP…

机器学习 · 计算机科学 2011-01-18 Louis Dorard , John Shawe-Taylor

In this paper, we study the problem of Gaussian process (GP) bandits under relaxed optimization criteria stating that any function value above a certain threshold is "good enough". On the theoretical side, we study various {\em lenient…

机器学习 · 统计学 2021-05-27 Xu Cai , Selwyn Gomes , Jonathan Scarlett

How to efficiently explore in reinforcement learning is an open problem. Many exploration algorithms employ the epistemic uncertainty of their own value predictions -- for instance to compute an exploration bonus or upper confidence bound.…

机器学习 · 计算机科学 2023-03-08 Simon Schmitt , John Shawe-Taylor , Hado van Hasselt

A sequential decision-making agent balances between exploring to gain new knowledge about an environment and exploiting current knowledge to maximize immediate reward. For environments studied in the traditional literature, optimal…

机器学习 · 计算机科学 2024-07-23 Dilip Arumugam , Wanqiao Xu , Benjamin Van Roy

Optimism about the poorly understood states and actions is the main driving force of exploration for many provably-efficient reinforcement learning algorithms. We propose optimism in the face of sensible value functions (OFVF)- a novel…

机器学习 · 计算机科学 2019-04-19 Reazul H. Russel , Tianyi Gu , Marek Petrik

We study online fair allocation of $T$ sequentially arriving items among $n$ agents with heterogeneous preferences, with the objective of maximizing generalized-mean welfare, defined as the $p$-mean of agents' time-averaged utilities, with…

计算机科学与博弈论 · 计算机科学 2026-02-12 Zongjun Yang , Rachitesh Kumar , Christian Kroer

The Upper Confidence Bounds For Trees (UCT) algorithm is not agnostic to the reward scale of the game it is applied to. For zero-sum games with the sparse rewards of $\{-1,0,1\}$ at the end of the game, this is not a problem, but many games…

人工智能 · 计算机科学 2025-10-27 Robin Schmöcker , Christoph Schnell , Alexander Dockhorn
‹ 上一页 1 8 9 10 下一页 ›