English
Related papers

Related papers: Optimal Exploration of an Exhaustible Resource wit…

200 papers

I formulate an entropy-rate maximization problem at the observable level for stochastic processes observed through an information-reducing observation map. For a visible stationary law, the map determines an observational fiber of hidden…

Information Theory · Computer Science 2026-04-14 Oleg Kiriukhin

It can be profitable for vehicle service providers to set service prices based on users' travel demand on different origin-destination pairs. The prior studies on the spatial pricing of vehicle service rely on the assumption that providers…

Computer Science and Game Theory · Computer Science 2020-07-08 Haoran Yu , Ermin Wei , Randall A. Berry

Reinforcement Learning is a powerful tool to model decision-making processes. However, it relies on an exploration-exploitation trade-off that remains an open challenge for many tasks. In this work, we study neighboring state-based,…

Machine Learning · Computer Science 2025-11-04 Yu-Teng Li , Justin Lin , Jeffery Cheng , Pedro Pachuca

This paper presents a new condition for the existence of optimal stationary policies in average-cost continuous-time Markov decision processes with unbounded cost and transition rates, arising from controlled queueing systems. This…

Optimization and Control · Mathematics 2015-04-23 Cao Ping , Xie Jingui

Motivated by applications where impatience is pervasive and evaluation times are uncertain, we study a selection model where options may expire at an unknown point in time and evaluation times are stochastic. Initially, the decision-maker…

Optimization and Control · Mathematics 2026-02-05 Yihua Xu , Rohan Ghuge , Sebastian Perez-Salazar

We present an alternative view for the study of optimal control of partially observed Markov Decision Processes (POMDPs). We first revisit the traditional (and by now standard) separated-design method of reducing the problem to fully…

Optimization and Control · Mathematics 2024-12-20 Serdar Yüksel

Learning optimal policies in sparse rewards settings is difficult as the learning agent has little to no feedback on the quality of its actions. In these situations, a good strategy is to focus on exploration, hopefully leading to the…

Machine Learning · Computer Science 2023-09-28 Giuseppe Paolo , Miranda Coninx , Alban Laflaquière , Stephane Doncieux

A sequential decision-making agent balances between exploring to gain new knowledge about an environment and exploiting current knowledge to maximize immediate reward. For environments studied in the traditional literature, optimal…

Machine Learning · Computer Science 2024-07-23 Dilip Arumugam , Wanqiao Xu , Benjamin Van Roy

This paper studies a search problem where a consumer is initially aware of only a few products. At every point in time, the consumer then decides between searching among alternatives he is already aware of and discovering more products. I…

Theoretical Economics · Economics 2022-02-21 Rafael P. Greminger

We consider excursions for a class of stochastic processes describing a population of discrete individuals experiencing density-limited growth, such that the population has a finite carrying capacity and behaves qualitatively like the…

Populations and Evolution · Quantitative Biology 2017-04-10 Todd L. Parsons

This paper deals with the unconstrained and constrained cases for continuous-time Markov decision processes under the finite-horizon expected total cost criterion. The state space is denumerable and the transition and cost rates are allowed…

Optimization and Control · Mathematics 2014-08-26 Qingda Wei , Xian Chen

When an item goes out of stock, sales transaction data no longer reflect the original customer demand, since some customers leave with no purchase while others substitute alternative products for the one that was out of stock. Here we…

Applications · Statistics 2016-01-15 Benjamin Letham , Lydia M. Letham , Cynthia Rudin

In the maximum state entropy exploration framework, an agent interacts with a reward-free environment to learn a policy that maximizes the entropy of the expected state visitations it is inducing. Hazan et al. (2019) noted that the class of…

Machine Learning · Computer Science 2022-07-11 Mirco Mutti , Riccardo De Santi , Marcello Restelli

Balancing exploration and exploitation remains a central challenge in reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs). Current RLVR methods often overemphasize exploitation, leading to entropy…

Computation and Language · Computer Science 2026-04-14 Liang Chen , Xueting Han , Qizhou Wang , Bo Han , Jing Bai , Hinrich Schutze , Kam-Fai Wong

We study offline dynamic pricing when historical data provide incomplete coverage of the price space such that some candidate prices, including the optimal one, may be entirely unobserved. This setting is common in practice and is…

Machine Learning · Statistics 2026-05-25 Zeyu Bian , Lan Wang , Zhengling Qi

We are interested in risk constraints for infinite horizon discrete time Markov decision processes (MDPs). Starting with average reward MDPs, we show that increasing concave stochastic dominance constraints on the empirical distribution of…

Optimization and Control · Mathematics 2012-06-21 William B. Haskell , Rahul Jain

We revisit the incremental autonomous exploration problem proposed by Lim & Auer (2012). In this setting, the agent aims to learn a set of near-optimal goal-conditioned policies to reach the $L$-controllable states: states that are…

Machine Learning · Computer Science 2022-05-24 Haoyuan Cai , Tengyu Ma , Simon Du

The Random Utility Maximization model is by far the most adopted framework to estimate consumer choice behavior. However, behavioral economics has provided strong empirical evidence of irrational choice behavior, such as halo effects, that…

Econometrics · Economics 2021-09-10 Sanjay Dominik Jena , Andrea Lodi , Claudio Sole

We study decision timing problems on finite horizon with Poissonian information arrivals. In our model, a decision maker wishes to optimally time her action in order to maximize her expected reward. The reward depends on an unobservable…

Optimization and Control · Mathematics 2012-05-07 Michael Ludkovski , Semih Sezer

Any firm whose business strategy has an exposure constraint that limits its potential gain naturally considers expansion, as this can increase its exposure. We model business expansion as an enlargement of the opportunity set for business…

Risk Management · Quantitative Finance 2021-12-14 Ling Wang , Kexin Chen , Mei Choi Chiu , Hoi Ying Wong