Related papers: Grab It Before It's Gone: Testing Uncertain Reward…
We consider the Wiener process with drift $$ dX_t=\mu dt +\sigma d W_t $$ with initial value problem $X_0=x_0$, where $x_0 \in R$, $ \mu \in R$ and $\sigma > 0$ are parameters. By use values $(z_k)_{k \in N}$ of corresponding trajectories…
The stochastic generalised linear bandit is a well-understood model for sequential decision-making problems, with many algorithms achieving near-optimal regret guarantees under immediate feedback. However, the stringent requirement for…
The problem of stochastic deadline scheduling is considered. A constrained Markov decision process model is introduced in which jobs arrive randomly at a service center with stochastic job sizes, rewards, and completion deadlines. The…
We consider the Gittins index for a normal distribution with unknown mean $\theta$ and known variance where $\theta$ has a normal prior. In addition to presenting some monotonicity properties of the Gittins index, we derive an approximation…
We study an inverse first-passage-time problem for Wiener process $X(t)$ subject to hold and jump from a boundary $c.$ Let be given a threshold $S>X(0) \ge c,$ and a distribution function $F$ on $[0, + \infty ).$ The problem consists in…
We study optimal stopping of Feller-Markov processes to maximise an undiscounted functional consisting of running and terminal rewards. In a finite-time horizon setting, we extend classical results to unbounded rewards. In infinite horizon,…
In this paper, we introduce a new method for applying the implicit function theorem to find nontrivial solutions to overdetermined problems with a fixed boundary (given) and a free boundary (to be determined). The novelty of this method…
This paper is devoted to studying an infinite time horizon stochastic recursive control problem with jumps, where infinite time horizon stochastic differential equation and backward stochastic differential equation with jumps describe the…
We provide sufficient conditions for the continuity of the free-boundary in a general class of finite-horizon optimal stopping problems arising for instance in finance and economics. The underlying process is a strong solution of one…
The Lipschitz bandit problem extends stochastic bandits to a continuous action set defined over a metric space, where the expected reward function satisfies a Lipschitz condition. In this work, we introduce a new problem of Lipschitz bandit…
In order to understand the impact of random influences at physical boundary on the evolution of multiscale systems, a stochastic partial differential equation model under a fast random dynamical boundary condition is investigated. The…
We study the problem of scheduling periodic real-time tasks so as to meet their individual minimum reward requirements. A task generates jobs that can be given arbitrary service times before their deadlines. A task then obtains rewards…
Control barrier functions are widely used to synthesize safety-critical controls. However, the presence of Gaussian-type noise in dynamical systems can generate unbounded signals and potentially result in severe consequences. Although…
The random greedy algorithm for constructing a large partial Steiner-Triple-System is defined as follows. Begin with a complete graph on $n$ vertices and proceed to remove the edges of triangles one at a time, where each triangle removed is…
In this paper, we address the stochastic contextual linear bandit problem, where a decision maker is provided a context (a random set of actions drawn from a distribution). The expected reward of each action is specified by the inner…
We develop novel empirical Bernstein inequalities for the variance of bounded random variables. Our inequalities hold under constant conditional variance and mean, without further assumptions like independence or identical distribution of…
We focus on a stochastic learning model where the learner observes a finite set of training examples and the output of the learning process is a data-dependent distribution over a space of hypotheses. The learned data-dependent distribution…
The input to the stochastic orienteering problem consists of a budget $B$ and metric $(V,d)$ where each vertex $v$ has a job with deterministic reward and random processing time (drawn from a known distribution). The processing times are…
This paper presents a class of Dynamic Multi-Armed Bandit problems where the reward can be modeled as the noisy output of a time varying linear stochastic dynamic system that satisfies some boundedness constraints. The class allows many…
We study an infinite horizon optimal stopping problem which arises naturally in the optimal timing of a firm/project sale or in the valuation of natural resources: the functional to be maximised is a sum of a discounted running reward and a…