English
Related papers

Related papers: Pandora's Box Problem With Time Constraints

200 papers

Exploration in reinforcement learning (RL) remains an open challenge. RL algorithms rely on observing rewards to train the agent, and if informative rewards are sparse the agent learns slowly or may not learn at all. To improve exploration…

Machine Learning · Computer Science 2024-11-12 Simone Parisi , Alireza Kazemipour , Michael Bowling

Multi-armed bandit algorithms have become a reference solution for handling the explore/exploit dilemma in recommender systems, and many other important real-world problems, such as display advertisement. However, such algorithms usually…

Machine Learning · Computer Science 2018-05-25 Qingyun Wu , Naveen Iyer , Hongning Wang

Recommender systems are a ubiquitous feature of online platforms. Increasingly, they are explicitly tasked with increasing users' long-term satisfaction. In this context, we study a content exploration task, which we formalize as a…

Machine Learning · Computer Science 2023-07-21 Thomas M. McDonald , Lucas Maystre , Mounia Lalmas , Daniel Russo , Kamil Ciosek

People often deviate from expected utility theory when making risky and intertemporal choices. While the effects of probabilistic risk and time delay have been extensively studied in isolation, their interplay and underlying theoretical…

Theoretical Economics · Economics 2025-04-10 Ho Ka Chan , Taro Toyoizumi

Bandit algorithms are guaranteed to solve diverse sequential decision-making problems, provided that a sufficient exploration budget is available. However, learning from scratch is often too costly for personalization tasks where a single…

Machine Learning · Computer Science 2025-08-08 Newton Mwai , Emil Carlsson , Fredrik D. Johansson

The Prophet Inequality and Pandora's Box problems are fundamental stochastic problem with applications in Mechanism Design, Online Algorithms, Stochastic Optimization, Optimal Stopping, and Operations Research. A usual assumption in these…

Data Structures and Algorithms · Computer Science 2023-12-08 Khashayar Gatmiry , Thomas Kesselheim , Sahil Singla , Yifan Wang

In this article, we consider the problem of unconstrained time-varying convex optimization, where the cost function changes with time. We provide an in-depth technical analysis of the problem and argue why freezing the cost at each time…

Optimization and Control · Mathematics 2024-10-28 M. Rostami , S. S. Kia

We formalize the problem of selecting the optimal set of options for planning as that of computing the smallest set of options so that planning converges in less than a given maximum of value-iteration passes. We first show that the problem…

Artificial Intelligence · Computer Science 2019-03-19 Yuu Jinnai , David Abel , D Ellis Hershkowitz , Michael Littman , George Konidaris

There are many papers written on the Two Envelopes Problem that usually study some of its variations. In this paper we will study and compare the most significant variations of the problem. We will see the correct decisions for each player…

History and Overview · Mathematics 2014-11-12 Panagiotis Tsikogiannopoulos

Individuals are often faced with temptations that can lead them astray from long-term goals. We're interested in developing interventions that steer individuals toward making good initial decisions and then maintaining those decisions over…

Machine Learning · Computer Science 2022-03-15 Shruthi Sukumar , Adrian F. Ward , Camden Elliott-Williams , Shabnam Hakimi , Michael C. Mozer

This paper studies an open question in the warehouse problem where a merchant trading a commodity tries to find an optimal inventory-trading policy to decide on purchase and sale quantities during a fixed time horizon in order to maximize…

Data Structures and Algorithms · Computer Science 2023-02-24 Ishan Bansal , Oktay Günlük

We study a sequential estimation problem for an unknown reward in the presence of a random horizon. The reward takes one of two predetermined values which can be inferred from the drift of a Wiener process, which serves as a signal. The…

Probability · Mathematics 2025-03-11 Steven Campbell , Georgy Gaitsgori , Richard Groenewald , Ioannis Karatzas

We consider a class of zero-sum search games in which a Hider hides one or more target among a set of $n$ boxes. The boxes may require differing amount of time to search, and detection may be imperfect, so that there is a certain…

Optimization and Control · Mathematics 2025-05-12 Thomas Lidbetter

We study a variant of the thresholding bandit problem (TBP) in the context of outlier detection, where the objective is to identify the outliers whose rewards are above a threshold. Distinct from the traditional TBP, the threshold is…

Machine Learning · Computer Science 2022-03-22 Xiaojin Zhang , Honglei Zhuang , Shengyu Zhang , Yuan Zhou

Recently, Frazier et al. proposed a natural model for crowdsourced exploration of different a priori unknown options: a principal is interested in the long-term welfare of a population of agents who arrive one by one in a multi-armed bandit…

Computer Science and Game Theory · Computer Science 2015-12-29 Li Han , David Kempe , Ruixin Qiang

The orienteering problem (OP) is a combinatorial optimization problem that seeks a path visiting a subset of locations to maximize collected rewards under a limited resource budget. This article presents a systematic PRISMA-based review of…

Optimization and Control · Mathematics 2025-12-19 Songhao Shen , Yufeng Zhou , Qin Lei , Zhibin Wu

Since Grover's seminal work, quantum search has been studied in great detail. In the usual search problem, we have a collection of n items and we would like to find a marked item. We consider a new variant of this problem in which…

Quantum Physics · Physics 2007-05-23 Andris Ambainis

We investigate some versions of the famous 100 prisoner problem for the infinite case, where there are infinitely many prisoners and infinitely many boxes with labels. In this case, many questions can be asked about the admissible steps of…

General Mathematics · Mathematics 2024-09-26 Attila Losonczi

We study a risk-aware robot planning problem where a dispatcher must construct a package delivery plan that maximizes the expected reward for a robot delivering packages across multiple epochs. Each package has an associated reward for…

Optimization and Control · Mathematics 2021-10-20 Blake Wilson , Jeffrey Hudack , Shreyas Sundaram

The Orienteering Problem with Time Window and Delay (\OPTiWinD) is a variant of the online orienteering problem. A series of requests appear in various locations while a vehicle moves within the territory to serve them. Each request has a…

Discrete Mathematics · Computer Science 2022-01-04 Marc Demange , David Ellison , Bertrand Jouve
‹ Prev 1 3 4 5 6 7 10 Next ›