English
Related papers

Related papers: A random measure approach to reinforcement learnin…

200 papers

By formulating data samples' formation as a Markov denoising process, diffusion models achieve state-of-the-art performances in a collection of tasks. Recently, many variants of diffusion models have been proposed to enable controlled…

Machine Learning · Computer Science 2023-04-17 Hengtong Zhang , Tingyang Xu

In this paper we propose the notion of continuous-time dynamic spectral risk-measure (DSR). Adopting a Poisson random measure setting, we define this class of dynamic coherent risk-measures in terms of certain backward stochastic…

Probability · Mathematics 2017-04-19 Dilip Madan , Martijn Pistorius , Mitja Stadje

Model-free reinforcement learning attempts to find an optimal control action for an unknown dynamical system by directly searching over the parameter space of controllers. The convergence behavior and statistical properties of these…

Optimization and Control · Mathematics 2021-03-17 Hesameddin Mohammadi , Armin Zare , Mahdi Soltanolkotabi , Mihailo R. Jovanović

Small quantum systems can now be continuously monitored experimentally which allows for the reconstruction of quantum trajectories. A peculiar feature of these trajectories is the emergence of jumps between the eigenstates of the observable…

Mathematical Physics · Physics 2015-06-09 Michel Bauer , Denis Bernard , Antoine Tilloy

We introduce a new probabilistic method for solving a class of impulse control problems based on their representations as Backward Stochastic Differential Equations (BSDEs for short) with constrained jumps. As an example, our method is used…

Computational Finance · Quantitative Finance 2015-03-17 Marie Bernhart , Huyên Pham , Peter Tankov , Xavier Warin

This paper explores continuous-time and state-space optimal stopping problems from a reinforcement learning perspective. We begin by formulating the stopping problem using randomized stopping times, where the decision maker's control is…

Optimization and Control · Mathematics 2026-03-12 Jodi Dianetti , Giorgio Ferrari , Renyuan Xu

We develop an approach for solving time-consistent risk-sensitive stochastic optimization problems using model-free reinforcement learning (RL). Specifically, we assume agents assess the risk of a sequence of random variables using dynamic…

Machine Learning · Computer Science 2022-12-01 Anthony Coache , Sebastian Jaimungal

This paper mainly investigates reflected stochastic recursive control problems governed by jump-diffusion dynamics. The system's state evolution is described by a stochastic differential equation driven by both Brownian motion and Poisson…

Optimization and Control · Mathematics 2025-05-15 Lu Liu , Qingmeng Wei

Engineering problems that are modeled using sophisticated mathematical methods or are characterized by expensive-to-conduct tests or experiments, are encumbered with limited budget or finite computational resources. Moreover, practical…

Machine Learning · Computer Science 2021-12-24 Yonatan Ashenafi , Piyush Pandita , Sayan Ghosh

We propose Deterministic Sequencing of Exploration and Exploitation (DSEE) algorithm with interleaving exploration and exploitation epochs for model-based RL problems that aim to simultaneously learn the system model, i.e., a Markov…

Machine Learning · Computer Science 2022-12-21 Piyush Gupta , Vaibhav Srivastava

In this study, we develop a stochastic optimal control approach with reinforcement learning structure to learn the unknown parameters appeared in the drift and diffusion terms of the stochastic differential equation. By choosing an…

Optimization and Control · Mathematics 2023-08-22 Shuzhen Yang

Reinforcement Learning from Human Feedback (RLHF) is increasingly used to fine-tune diffusion models, but a key challenge arises from the mismatch between stochastic samplers used during training and deterministic samplers used during…

Machine Learning · Computer Science 2025-12-17 Jiayuan Sheng , Hanyang Zhao , Haoxian Chen , David D. Yao , Wenpin Tang

This paper studies a discrete-time mean-variance model based on reinforcement learning. Compared with its continuous-time counterpart in \cite{zhou2020mv}, the discrete-time model makes more general assumptions about the asset's return…

Mathematical Finance · Quantitative Finance 2023-12-27 Xiangyu Cui , Xun Li , Yun Shi , Si Zhao

In reinforcement learning (RL) algorithms, exploratory control inputs are used during learning to acquire knowledge for decision making and control, while the true dynamics of a controlled object is unknown. However, this exploring property…

Machine Learning · Computer Science 2021-03-08 Yoshihiro Okawa , Tomotake Sasaki , Hidenao Iwane

We describe a measurement device principle based on discrete iterations of Bayesian updating of system state probability distributions. Although purely classical by nature, these measurements are accompanied with a progressive collapse of…

Mathematical Physics · Physics 2015-06-11 Michel Bauer , Denis Bernard , Tristan Benoist

In this paper we make a survey on the so called randomization method, a recent methodology to study stochastic optimization problems. It allows to represent the value function of an optimal control problem by a suitable backward stochastic…

Optimization and Control · Mathematics 2025-06-12 Marco Fuhrman

We consider dynamic risk measures induced by Backward Stochastic Differential Equations (BSDEs) in enlargement of filtration setting. On a fixed probability space, we are given a standard Brownian motion and a pair of random variables…

Risk Management · Quantitative Finance 2020-09-25 Alessandro Calvia , Emanuela Rosazza Gianin

Numerical approximation of the long time behavior of a stochastic differential equation (SDE) is considered. Error estimates for time-averaging estimators are obtained and then used to show that the stationary behavior of the numerical…

Probability · Mathematics 2013-11-26 Jonathan C. Mattingly , Andrew M. Stuart , M. V. Tretyakov

The paper introduces an interactive machine learning mechanism to process the measurements of an uncertain, nonlinear dynamic process and hence advise an actuation strategy in real-time. For concept demonstration, a trajectory-following…

Systems and Control · Electrical Eng. & Systems 2023-03-16 Mohammed Abouheaf , Derek Boase , Wail Gueaieb , Davide Spinello , Salah Al-Sharhan

This paper proposes a reinforcement learning (RL) algorithm for infinite horizon $\rm {H_{2}/H_{\infty}}$ problem in a class of stochastic discrete-time systems, rather than using a set of coupled generalized algebraic Riccati equations…

Optimization and Control · Mathematics 2023-11-28 Xiushan Jiang , Li Wang , Dongya Zhao , Ling Shi