English
Related papers

Related papers: Grab It Before It's Gone: Testing Uncertain Reward…

200 papers

We propose a new approach to solving dynamic decision problems with rewards that are unbounded below. The approach involves transforming the Bellman equation in order to convert an unbounded problem into a bounded one. The major advantage…

Theoretical Economics · Economics 2019-12-02 Qingyin Ma , John Stachurski

We study the valuation of an American put option with a random time horizon given by the last exit time of the underlying asset from a fixed level. Since this random time is not a stopping time, the problem falls outside the classical…

Probability · Mathematics 2026-03-31 Zhuoshu Wu , Libo Li

Inference on unknown quantities in dynamical systems via observational data is essential for providing meaningful insight, furnishing accurate predictions, enabling robust control, and establishing appropriate designs for future…

Methodology · Statistics 2018-02-06 M. Chung , M. Binois , R. B. Gramacy , D. J. Moquin , A. P. Smith , A. M. Smith

In this paper we provide a theoretical analysis of Variable Annuities with a focus on the holder's right to an early termination of the contract. We obtain a rigorous pricing formula and the optimal exercise boundary for the surrender…

Mathematical Finance · Quantitative Finance 2024-05-06 Tiziano De Angelis , Alessandro Milazzo , Gabriele Stabile

Recent work has considered natural variations of the multi-armed bandit problem, where the reward distribution of each arm is a special function of the time passed since its last pulling. In this direction, a simple (yet widely applicable)…

Machine Learning · Computer Science 2021-05-25 Alexia Atsidakou , Orestis Papadigenopoulos , Soumya Basu , Constantine Caramanis , Sanjay Shakkottai

Autonomous exploration in mobile robotics often involves a trade-off between two objectives: maximizing environmental coverage and minimizing the total path length. In the widely used information gain paradigm, exploration is guided by the…

Robotics · Computer Science 2025-04-22 Ludvig Ericson , José Pedro , Patric Jensfelt

This paper is concerned with the problem of Model Predictive Control and Rolling Horizon Control of discrete-time systems subject to possibly unbounded random noise inputs, while satisfying hard bounds on the control inputs. We use a…

Optimization and Control · Mathematics 2010-09-08 Peter Hokayem , Debasish Chatterjee , John Lygeros

We study a nonstationary bandit problem where rewards depend on both actions and latent states, the latter governed by unknown linear dynamics. Crucially, the state dynamics also depend on the actions, resulting in tension between…

Machine Learning · Computer Science 2025-10-21 Sunmook Choi , Yahya Sattar , Yassir Jedra , Maryam Fazel , Sarah Dean

How can we perform efficient inference and learning in directed probabilistic models, in the presence of continuous latent variables with intractable posterior distributions, and large datasets? We introduce a stochastic variational…

Machine Learning · Statistics 2022-12-13 Diederik P Kingma , Max Welling

This paper studies finite-time safety and reach-avoid verification for stochastic discrete-time dynamical systems. The aim is to ascertain lower and upper bounds of the probability that, within a predefined finite-time horizon, a system…

Systems and Control · Electrical Eng. & Systems 2025-10-22 Bai Xue

We consider sequential decision problems in which we adaptively choose one of finitely many alternatives and observe a stochastic reward. We offer a new perspective of interpreting Bayesian ranking and selection problems as adaptive…

Machine Learning · Computer Science 2016-06-16 Yingfei Wang , Warren Powell

Most bandit algorithms assume that the reward variances or their upper bounds are known, and that they are the same for all arms. This naturally leads to suboptimal performance and higher regret due to variance overestimation. On the other…

Machine Learning · Computer Science 2023-10-13 Aadirupa Saha , Branislav Kveton

We provide necessary and sufficient conditions for stochastic invariance of finite dimensional submanifolds with boundary in Hilbert spaces for stochastic partial differential equations driven by Wiener processes and Poisson random…

Probability · Mathematics 2014-06-23 Damir Filipovic , Stefan Tappe , Josef Teichmann

We introduce a problem in which a service vehicle seeks to guard a deadline (boundary) from dynamically arriving mobile targets. The environment is a rectangle and the deadline is one of its edges. Targets arrive continuously over time on…

Robotics · Computer Science 2016-11-17 Stephen L. Smith , Shaunak D. Bopardikar , Francesco Bullo

Humanity has been fascinated by the pursuit of fortune since time immemorial, and many successful outcomes benefit from strokes of luck. But success is subject to complexity, uncertainty, and change - and at times becoming increasingly…

General Economics · Economics 2019-04-19 Didier Sornette , Spencer Wheatley , Peter Cauwels

Motivated by applications where impatience is pervasive and evaluation times are uncertain, we study a selection model where options may expire at an unknown point in time and evaluation times are stochastic. Initially, the decision-maker…

Optimization and Control · Mathematics 2026-02-05 Yihua Xu , Rohan Ghuge , Sebastian Perez-Salazar

We study the problem of optimally managing an inventory with unknown demand trend. Our formulation leads to a stochastic control problem under partial observation, in which a Brownian motion with non-observable drift can be singularly…

Optimization and Control · Mathematics 2022-11-28 Salvatore Federico , Giorgio Ferrari , Neofytos Rodosthenous

Markov reward processes (MRPs) are used to model stochastic phenomena arising in operations research, control engineering, robotics, and artificial intelligence, as well as communication and transportation networks. In many of these cases,…

Machine Learning · Statistics 2020-09-17 Ashwin Pananjady , Martin J. Wainwright

We study stochastic linear bandits where, in each round, the learner receives a set of actions (i.e., feature vectors), from which it chooses an element and obtains a stochastic reward. The expected reward is a fixed but unknown linear…

Machine Learning · Computer Science 2024-06-04 Tianyuan Jin , Kyoungseok Jang , Nicolò Cesa-Bianchi

We propose a new framework for imposing monotonicity constraints in a Bayesian nonparametric setting based on numerical solutions of stochastic differential equations. We derive a nonparametric model of monotonic functions that allows for…

Machine Learning · Statistics 2020-02-26 Ivan Ustyuzhaninov , Ieva Kazlauskaite , Carl Henrik Ek , Neill D. F. Campbell