English
Related papers

Related papers: Optimal Exploration of an Exhaustible Resource wit…

200 papers

In this paper, preys with stochastic evasion policies are considered. The stochasticity adds unpredictable changes to the prey's path for avoiding predator's attacks. The prey's cost function is composed of two terms balancing the…

Robotics · Computer Science 2021-10-12 Farhad Farokhi , Magnus Egerstedt

The Exploration-Exploitation tradeoff arises in Reinforcement Learning when one cannot tell if a policy is optimal. Then, there is a constant need to explore new actions instead of exploiting past experience. In practice, it is common to…

Machine Learning · Computer Science 2019-09-10 Lior Shani , Yonathan Efroni , Shie Mannor

We consider a model of nomadic agents exploring and competing for time-varying location-specific resources, arising in crowdsourced transportation services, online communities, and in traditional location based economic activity. This model…

Computer Science and Game Theory · Computer Science 2016-02-23 Pu Yang , Krishnamurthy Iyer , Peter Frazier

We consider a hide-and-seek game between a Hider and a Seeker over a finite set of locations. The Hider chooses one location to conceal a stationary treasure, while the Seeker visits the locations sequentially along a route. As the search…

Systems and Control · Electrical Eng. & Systems 2026-03-31 Prajakta Surve , Shaunak D. Bopardikar , Daigo Shishika , Dipankar Maity , Michael Dorothy

Ensuring sufficient exploration is a central challenge when training meta-reinforcement learning (meta-RL) agents to solve novel environments. Conventional solutions to the exploration-exploitation dilemma inject explicit incentives such as…

Machine Learning · Computer Science 2025-08-05 Micah Rentschler , Jesse Roberts

One of the classic data mining tasks is to discover bursts, time intervals, where events occur at abnormally high rate. In this paper we revisit Kleinberg's seminal work, where bursts are discovered by using exponential distribution with a…

Data Structures and Algorithms · Computer Science 2019-02-06 Nikolaj Tatti

A promising technique for exploration is to maximize the entropy of visited state distribution, i.e., state entropy, by encouraging uniform coverage of visited state space. While it has been effective for an unsupervised setup, it tends to…

Machine Learning · Computer Science 2024-08-12 Dongyoung Kim , Jinwoo Shin , Pieter Abbeel , Younggyo Seo

We develop a probabilistic framework for analysing model-based reinforcement learning in the episodic setting. We then apply it to study finite-time horizon stochastic control problems with linear dynamics but unknown coefficients and…

Machine Learning · Computer Science 2021-12-22 Lukasz Szpruch , Tanut Treetanthiploet , Yufei Zhang

Many online companies sell advertisement space in second-price auctions with reserve. In this paper, we develop a probabilistic method to learn a profitable strategy to set the reserve price. We use historical auction data with features to…

Machine Learning · Statistics 2015-06-25 Maja R. Rudolph , Joseph G. Ellis , David M. Blei

Tourism demand forecasting is methodologically mature, but it typically treats accommodation supply as fixed or exogenous. In platform-mediated short-term rentals, supply is elastic, decision-driven, and co-evolves with demand through…

Applications · Statistics 2026-04-29 Harrison Katz

Demand estimation plays an important role in dynamic pricing where the optimal price can be obtained via maximizing the revenue based on the demand curve. In online hotel booking platform, the demand or occupancy of rooms varies across…

General Economics · Economics 2022-08-12 Fanwei Zhu , Wendong Xiao , Yao Yu , Ziyi Wang , Zulong Chen , Quan Lu , Zemin Liu , Minghui Wu , Shenghua Ni

We consider a discounted infinite horizon optimal stopping problem. If the underlying distribution is known a priori, the solution of this problem is obtained via dynamic programming (DP) and is given by a well known threshold rule. When…

Machine Learning · Computer Science 2021-02-23 Daniel Russo , Assaf Zeevi , Tianyi Zhang

We consider "time-of-use" pricing as a technique for matching supply and demand of temporal resources with the goal of maximizing social welfare. Relevant examples include energy, computing resources on a cloud computing platform, and…

Computer Science and Game Theory · Computer Science 2017-04-11 Shuchi Chawla , Nikhil R. Devanur , Alexander E. Holroyd , Anna Karlin , James Martin , Balasubramanian Sivan

We consider a ubiquitous scenario in the Internet economy when individual decision-makers (henceforth, agents) both produce and consume information as they make strategic choices in an uncertain environment. This creates a three-way…

Computer Science and Game Theory · Computer Science 2021-04-09 Yishay Mansour , Aleksandrs Slivkins , Vasilis Syrgkanis , Zhiwei Steven Wu

The problem of pure exploration in Markov decision processes has been cast as maximizing the entropy over the state distribution induced by the agent's policy, an objective that has been extensively studied. However, little attention has…

Machine Learning · Computer Science 2024-06-19 Riccardo Zamboni , Duilio Cirino , Marcello Restelli , Mirco Mutti

We investigate activities that have different periods of duration. We define the profit intensity as a measure of this economic category. The profit intensity in a repeated trading has a unique property of attaining its maximum at a fixed…

Trading and Market Microstructure · Quantitative Finance 2009-11-13 Edward W. Piotrowski , Jan Sladkowski

This paper examines the objective of optimally harvesting a single species in a stochastic environment. This problem has previously been analyzed in Alvarez (2000) using dynamic programming techniques and, due to the natural payoff…

Optimization and Control · Mathematics 2016-08-02 Richard H. Stockbridge , Chao Zhu

Pricing algorithms have demonstrated the capability to learn tacit collusion that is largely unaddressed by current regulations. Their increasing use in markets, including oligopolistic industries with a history of collusion, calls for…

Computer Science and Game Theory · Computer Science 2025-02-26 Paul Friedrich , Barna Pásztor , Giorgia Ramponi

We consider a stochastic optimal control problem in a market model with temporary and permanent price impact, which is related to an expected utility maximization problem under finite fuel constraint. We establish the initial condition…

Mathematical Finance · Quantitative Finance 2015-10-13 Mourad Lazgham

Chase-escape is a competitive growth process in which red particles spread to adjacent empty sites according to a rate-$\lambda$ Poisson process while being chased and consumed by blue particles according to a rate-$1$ Poisson process.…

Probability · Mathematics 2022-05-24 Emma Bernstein , Clare Hamblen , Matthew Junge , Lily Reeves