Related papers: Grab It Before It's Gone: Testing Uncertain Reward…
We consider a class of time-inhomogeneous optimal stopping problems and we provide sufficient conditions on the data of the problem that guarantee monotonicity of the optimal stopping boundary. In our setting, time-inhomogeneity stems not…
Efficient exploration remains a challenging problem in reinforcement learning, especially for those tasks where rewards from environments are sparse. A commonly used approach for exploring such environments is to introduce some "intrinsic"…
In reinforcement learning episodes, the rewards and punishments are often non-deterministic, and there are invariably stochastic elements governing the underlying situation. Such stochastic elements are often numerous and cannot be known in…
We study the regularity and well-posedness of physical solutions to the supercooled Stefan problem. Assuming only that the initial temperature is integrable, we prove that the free boundary, known to have jump discontinuities as a function…
A new method for solving stiff boundary value problems is described and compared to other known approaches using the Troesch's problem as a test example. The method is based on the general idea of alternate approximation of either the…
Variational inequalities have gained significant attention in machine learning and optimization research. While stochastic methods for solving these problems typically assume independent data sampling, we investigate an alternative approach…
In this paper, we derive a new handy integral equation for the free-boundary of infinite time horizon, continuous time, stochastic, irreversible investment problems with uncertainty modeled as a one-dimensional, regular diffusion $X$. The…
This paper considers a mortgage contract where the borrower pays a fixed mortgage rate and has the choice of making prepayment. Assume the market interest follows the CIR model, a free boundary problem is formulated. Here we focus on the…
We propose a stochastic approximation method for approximating the efficient frontier of chance-constrained nonlinear programs. Our approach is based on a bi-objective viewpoint of chance-constrained programs that seeks solutions on the…
This paper studies an optimal stochastic impulse control problem in a finite horizon with a decision lag, by which we mean that after an impulse is made, a fixed number units of time has to be elapsed before the next impulse is allowed to…
In this note we introduce and solve a soft classification version of the famous Bayesian sequential testing problem for a Brownian motion's drift. We establish that the value function is the unique non-trivial solution to a free boundary…
We study a controlled version of the Bayesian sequential testing problem for the drift of a Wiener process, in which the observer exercises discretion over the signal intensity. This control incurs a running cost that reflects the resource…
We consider a singular control problem that aims to maximize the expected cumulative rewards, where the instantaneous returns depend on the state of a controlled process. The contributions of this paper are twofold. Firstly, to establish…
We study approximation of non-autonomous linear differential equations with variable delay over infinite intervals. We use piecewise constant argument to obtain a corresponding discrete difference equation. The study of numerical…
Tipping points have been actively studied in various applications as well as from a mathematical viewpoint. A main technique to theoretically understand early-warning signs for tipping points is to use the framework of fast-slow stochastic…
In this paper, motivated by a problem in stochastic impulse control theory, we aim to study solutions to a free boundary problem of obstacle-type. We obtain sharp estimates for the solution using nonlinear tools which are independent of the…
We consider Markov decision processes (MDPs) in which the transition probabilities and rewards belong to an uncertainty set parametrized by a collection of random variables. The probability distributions for these random parameters are…
Online decision-making can be formulated as the popular stochastic multi-armed bandit problem where a learner makes decisions (or takes actions) to maximize cumulative rewards collected from an unknown environment. This paper proposes to…
This note is devoted to continuity results of the time derivative of the solution to the one-dimensional parabolic obstacle problem with variable coefficients. It applies to the smooth fit principle in numerical analysis and in financial…
Boundary estimation in images and videos has been a very active topic of research, and organizing visual information into boundaries and segments is believed to be a corner stone of visual perception. While prior work has focused on…