相关论文: Long-Run Average Sustainable Harvesting Policies: …
The aim of this paper is to study the reward based policy exploration problem in a supervised learning approach and enable robots to form complex movement trajectories in challenging reward settings and search spaces. For this, the…
In this paper, we consider a general time-inconsistent optimal control problem for a non homogeneous linear system, in which its state evolves according to a stochastic differential equation with deterministic coefficients, when the noise…
This paper concerns rollout and certainty-equivalent rollout policies for stochastic shortest path problems with absorbing terminal states. The main result provides a direct non-asymptotic performance certificate for a fixed rollout policy:…
In order to approximate the exit time of a one-dimensional diffusion process, we propose an algorithm based on a random walk. Such an algorithm was already introduced in both the Brownian context and in the Ornstein-Uhlenbeck context. Here…
This paper investigates the dynamics and optimal harvesting of age-structured populations governed by McKendrick--von Foerster equations, contrasting two distinct harvesting mechanisms: rate-control and effort-control. For the rate-control…
This article studies typical dynamics and fluctuations for a slow-fast dynamical system perturbed by a small fractional Brownian noise. Based on an ergodic theorem with explicit rates of convergence, which may be of independent interest, we…
We study the maximum likelihood estimator of the drift parameters of a stochastic differential equation, with both drift and diffusion coefficients constant on the positive and negative axis, yet discontinuous at zero. This threshold…
We present an approach for approximately solving discrete-time stochastic optimal-control problems by combining direct trajectory optimization, deterministic sampling, and policy optimization. Our feedback motion-planning algorithm uses a…
In this paper, we consider the classic stochastic (dynamic) knapsack problem, a fundamental mathematical model in revenue management, with general time-varying random demand. Our main goal is to study the optimal policies, which can be…
This work deals with the solution of a non-convex optimization problem to enhance the performance of an energy harvesting device, which involves a nonlinear objective function and a discontinuous constraint. This optimization problem, which…
We establish almost sure invariance principles, a strong form of approximation by Brownian motion, for non-stationary time-series arising as observations on dynamical systems. Our examples include observations on sequential expanding maps,…
When monitoring the dynamics of stochastic systems, such as interacting particles agitated by thermal noise, disentangling deterministic forces from Brownian motion is challenging. Indeed, we show that there is an information-theoretic…
We analyze the effect of additive fractional noise with Hurst parameter $H > \frac{1}{2}$ on fast-slow systems. Our strategy is based on sample paths estimates, similar to the approach by Berglund and Gentz in the Brownian motion case. Yet,…
This paper addresses the problem of managing perishable inventory under multiple sources of uncertainty, including stochastic demand, unreliable supplier fulfillment, and probabilistic product shelf life. We develop a discrete-event…
In the automation of many kinds of processes, the observable outcome can often be described as the combined effect of an entire sequence of actions, or controls, applied throughout its execution. In these cases, strategies to optimise…
In this paper we study a Pontryagin type stochastic maximum principle for the optimal control of a system, where the state dynamics satisfy a stochastic partial differential equation (SPDE) driven by a two-parameter (time-space) Brownian…
We consider slow / fast systems where the slow system is driven by fractional Brownian motion with Hurst parameter $H>{1\over 2}$. We show that unlike in the case $H={1\over 2}$, convergence to the averaged solution takes place in…
The autonomous systems need to decide how to react to the changes at runtime efficiently. The ability to rigorously analyze the environment and the system together is theoretically possible by the model-driven approaches; however, the model…
Temporal abstraction and efficient planning pose significant challenges in offline reinforcement learning, mainly when dealing with domains that involve temporally extended tasks and delayed sparse rewards. Existing methods typically plan…
This paper develops a unified methodology for probabilistic analysis and optimal control design for jump diffusion processes defined by polynomials. For such systems, the evolution of the moments of the state can be described via a system…