English
Related papers

Related papers: Pure Exploration with Infinite Answers

200 papers

This paper introduces the framework of multi-armed sampling, which serves as the sampling counterpart to the optimization problem of multi-armed bandits. Our primary motivation is to rigorously examine the exploration-exploitation trade-off…

Machine Learning · Computer Science 2026-05-14 Mohammad Pedramfar , Siamak Ravanbakhsh

Motivated by real-world applications such as fast fashion retailing and online advertising, the Multinomial Logit Bandit (MNL-bandit) is a popular model in online learning and operations research, and has attracted much attention in the…

Machine Learning · Computer Science 2021-08-17 Nikolai Karpov , Qin Zhang

In pure-exploration problems, information is gathered sequentially to answer a question on the stochastic environment. While best-arm identification for linear bandits has been extensively studied in recent years, few works have been…

Machine Learning · Statistics 2022-06-10 Marc Jourdan , Rémy Degenne

Using the 20 questions estimation framework with query-dependent noise, we study non-adaptive search strategies for a moving target over the unit cube with unknown initial location and velocities under a piecewise constant velocity model.…

Information Theory · Computer Science 2023-06-02 Lin Zhou , Alfred Hero

We study tracking-type optimal control problems that involve a non-affine, weak-to-weak continuous control-to-state mapping, a desired state $y_d$, and a desired control $u_d$. It is proved that such problems are always nonuniquely solvable…

Optimization and Control · Mathematics 2021-02-04 Constantin Christof , Dominik Hafemeyer

This paper investigates the behavior of sets and functions at infinity by introducing new concepts, namely directional normal cones at infinity for unbounded sets, along with limiting and singular subdifferentials at infinity in the…

Optimization and Control · Mathematics 2025-10-13 Le Ngoc Kien , Nguyen Van Tuyen , Tran Van Nghi

In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to identify an unknown ground-truth hypothesis with as few queries as possible. This problem, originally…

Machine Learning · Computer Science 2026-02-23 Alessio Russo , Yin-Ching Lee , Ryan Welch , Aldo Pacchiano

In this paper, we study the optimal stopping problem in the so-called exploratory framework, in which the agent takes actions randomly conditioning on current state and an entropy-regularized term is added to the reward functional. Such a…

Optimization and Control · Mathematics 2023-09-04 Yuchao Dong

Learning problems commonly exhibit an interesting feedback mechanism wherein the population data reacts to competing decision makers' actions. This paper formulates a new game theoretic framework for this phenomenon, called "multi-player…

Computer Science and Game Theory · Computer Science 2022-04-08 Adhyyan Narang , Evan Faulkner , Dmitriy Drusvyatskiy , Maryam Fazel , Lillian J. Ratliff

We study the Combinatorial Pure Exploration problem with Continuous and Separable reward functions (CPE-CS) in the stochastic multi-armed bandit setting. In a CPE-CS instance, we are given several stochastic arms with unknown distributions,…

Machine Learning · Computer Science 2018-05-07 Weiran Huang , Jungseul Ok , Liang Li , Wei Chen

We investigate stochastic differential games of optimal trading comprising a finite population. There are market frictions in the present framework, which take the form of stochastic permanent and temporary price impacts. Moreover,…

Mathematical Finance · Quantitative Finance 2021-02-09 David Evangelista , Yuri Thamsten

We study the problem of regression with interval targets, where only upper and lower bounds on target values are available in the form of intervals. This problem arises when the exact target label is expensive or impossible to obtain, due…

Machine Learning · Computer Science 2025-10-27 Rattana Pukdee , Ziqi Ke , Chirag Gupta

We introduce the "inverse bandit" problem of estimating the rewards of a multi-armed bandit instance from observing the learning process of a low-regret demonstrator. Existing approaches to the related problem of inverse reinforcement…

Machine Learning · Statistics 2022-02-23 Wenshuo Guo , Kumar Krishna Agrawal , Aditya Grover , Vidya Muthukumar , Ashwin Pananjady

Reinforcement Learning has emerged as a strong alternative to solve optimization tasks efficiently. The use of these algorithms highly depends on the feedback signals provided by the environment in charge of informing about how good (or…

Machine Learning · Computer Science 2022-12-01 Alain Andres , Esther Villar-Rodriguez , Javier Del Ser

Recent work on exploration in reinforcement learning (RL) has led to a series of increasingly complex solutions to the problem. This increase in complexity often comes at the expense of generality. Recent empirical studies suggest that,…

Machine Learning · Computer Science 2020-06-03 Will Dabney , Georg Ostrovski , André Barreto

This paper investigates the Nash equilibrium of a bi-objective optimal control problem governed by the Stokes equations. A multi-objective Nash strategy is formulated, and fundamental theoretical results are established, including the…

Optimization and Control · Mathematics 2025-12-16 Kedarnath Buda , B. V. Rathish Kumar , Anil Rathi

Suboptimal methods in optimal control arise due to a limited computational budget, unknown system dynamics, or a short prediction window among other reasons. Although these methods are ubiquitous, their transient performance remains…

Systems and Control · Electrical Eng. & Systems 2025-04-08 Aren Karapetyan , Efe C. Balta , Andrea Iannelli , John Lygeros

Pure exploration in multi-armed bandits has emerged as an important framework for modeling decision-making and search under uncertainty. In modern applications, however, one is often faced with a tremendously large number of options. Even…

Machine Learning · Computer Science 2022-11-22 Parth K. Thaker , Mohit Malu , Nikhil Rao , Gautam Dasarathy

In this two-part study we develop a general approach to the design and analysis of exact penalty functions for various optimal control problems, including problems with terminal and state constraints, problems involving differential…

Optimization and Control · Mathematics 2020-01-10 M. V. Dolgopolik , A. V. Fominyh

The problem of finding the sparsest solution to a linear underdetermined system of equations, often appearing, e.g., in data analysis, optimal control, system identification, or sensor selection problems, is considered. This non-convex…

Optimization and Control · Mathematics 2026-03-17 Maya V. Marmary , Christian Grussler