Related papers: On a connection between total positivity and Berno…
Let $X_n, n \ge 0$ be a Markov chain with finite state space $M$. If $x,y \in M$ such that $x$ is transient we have $P^y(X_n = x) \to 0$ for $n \to \infty$, and under mild aperiodicity conditions this convergence is monotone in that for…
This paper deals with control of partially observable discrete-time stochastic systems. It introduces and studies Markov Decision Processes with Incomplete Information and with semi-uniform Feller transition probabilities. The important…
This paper investigates MDPs with intermittent state information. We consider a scenario where the controller perceives the state information of the process via an unreliable communication channel. The transmissions of state information…
Myopic strategy is one of the most important strategies when studying bandit problems. In this paper, we consider the two-armed bandit problem proposed by Feldman. With general distributions and utility functions, we obtain a necessary and…
This paper is devoted to studying the average optimality in continuous-time Markov decision processes with fairly general state and action spaces. The criterion to be maximized is expected average rewards. The transition rates of underlying…
This paper explores continuous-time and state-space optimal stopping problems from a reinforcement learning perspective. We begin by formulating the stopping problem using randomized stopping times, where the decision maker's control is…
We consider an optimal stopping time problem related with many models found in real options problems. The main goal of this work is to bring for the field of real options, different and more realistic pay-off functions, and negative…
We consider an optimal stopping problem with n correlated offers where the goal is to design a (randomized) stopping strategy that maximizes the expected value of the offer in the sequence at which we stop. Instead of assuming to know the…
Causal reversibility blends reversibility and causality for concurrent systems. It indicates that an action can be undone provided that all of its consequences have been undone already, thus making it possible to bring the system back to a…
Given a dynamic ordinal game, we deem a strategy sequentially rational if there exist a Bernoulli utility function and a conditional probability system with respect to which the strategy is a maximizer. We establish a complete class theorem…
We consider the classical last-success problem for sequential Bernoulli trials in the homogeneous setting where $X_1,\ldots,X_n$ are i.i.d. $\mathrm{Bernoulli}(p)$ but the success probability $p\in(0,1)$ is unknown to the decision maker.…
We consider two variations of the classical secretary problem. * A variation of the returning secretary problem where each interviewee may appear a second time with a fixed probability p. The decision-maker observes interviewees…
We study a sequential coin-flipping game in which a player starts with~$n$ coins, each landing heads independently with probability~$p$. In each round the player flips all remaining coins and must set aside at least one coin showing heads;…
Deciding the positivity of a sequence defined by a linear recurrence with polynomial coefficients and initial condition is difficult in general. Even in the case of recurrences with constant coefficients, it is known to be decidable only…
We formulate an optimal stopping problem for a geometric Brownian motion where the probability scale is distorted by a general nonlinear function. The problem is inherently time inconsistent due to the Choquet integration involved. We…
Consider the problem of approximating the optimal policy of a Markov decision process (MDP) by sampling state transitions. In contrast to existing reinforcement learning methods that are based on successive approximations to the nonlinear…
In the best choice problem with random arrivals, an unknown number $n$ of rankable items arrive at times sampled from the uniform distribution. As is well known, a real-time player can ensure stopping at the overall best item with…
A basic question for zero-sum repeated games consists in determining whether the mean payoff per time unit is independent of the initial state. In the special case of "zero-player" games, i.e., of Markov chains equipped with additive…
The problem of making sequential decisions in unknown probabilistic environments is studied. In cycle $t$ action $y_t$ results in perception $x_t$ and reward $r_t$, where all quantities in general may depend on the complete history. The…
Given an initial (resp., terminal) probability measure $\mu$ (resp., $\nu$) on $\mathbb{R}^d$, we characterize those optimal stopping times $\tau$ that maximize or minimize the functional $\mathbb{E} |B_0 - B_\tau|^{\alpha}$, $\alpha > 0$,…