English
Related papers

Related papers: A random measure approach to reinforcement learnin…

200 papers

We consider reinforcement learning (RL) in continuous time and study the problem of achieving the best trade-off between exploration of a black box environment and exploitation of current knowledge. We propose an entropy-regularized reward…

Optimization and Control · Mathematics 2019-02-14 Haoran Wang , Thaleia Zariphopoulou , Xunyu Zhou

We establish well-posedness for a class of systems of SDEs with non-Lipschitz coefficients in the diffusion and jump terms and with two sources of interdependence: a monotone function of all the components in the drift of each SDE and the…

Probability · Mathematics 2026-03-24 Ying Jiao , Nikolaos Kolliopoulos

We study the problem of safe learning and exploration in sequential control problems. The goal is to safely collect data samples from operating in an environment, in order to learn to achieve a challenging control goal (e.g., an agile…

Machine Learning · Computer Science 2020-06-30 Anqi Liu , Guanya Shi , Soon-Jo Chung , Anima Anandkumar , Yisong Yue

We consider the problem of designing control laws for stochastic jump linear systems where the disturbances are drawn randomly from a finite sample space according to an unknown distribution, which is estimated from a finite sample of…

Systems and Control · Computer Science 2019-10-31 Mathijs Schuurmans , Pantelis Sopasakis , Panagiotis Patrinos

Intensity control is a class of continuous-time dynamic optimization problems with many important applications in Operations Research including queueing and revenue management. In this study, we propose a practical continuous-time…

Machine Learning · Computer Science 2026-04-14 Huiling Meng , Ningyuan Chen , Xuefeng Gao

We study the learning dynamics of a multi-pass, mini-batch Stochastic Gradient Descent (SGD) procedure for empirical risk minimization in high-dimensional multi-index models with isotropic random data. In an asymptotic regime where the…

Machine Learning · Statistics 2026-02-19 Zhou Fan , Leda Wang

In this paper, we examine a stochastic linear-quadratic control problem characterized by regime switching and Poisson jumps. All the coefficients in the problem are random processes adapted to the filtration generated by Brownian motion and…

Optimization and Control · Mathematics 2024-12-30 Xiaomin Shi , Zuo Quan Xu

IIn this paper, we study a partially observed progressive optimal control problem of forward-backward stochastic differential equations with random jumps, where the control domain is not necessarily convex, and the control variable enter…

Optimization and Control · Mathematics 2022-06-27 Yueyang Zheng , Jingtao Shi

We introduce a novel grid-independent model for learning partial differential equations (PDEs) from noisy and partial observations on irregular spatiotemporal grids. We propose a space-time continuous latent neural PDE model with an…

Machine Learning · Computer Science 2023-10-27 Valerii Iakovlev , Markus Heinonen , Harri Lähdesmäki

In this paper, we study a class of multi-dimensional reflected backward stochastic differential equations when the noise is driven by a Brownian motion and an independent Poisson point process, and when the solution is forced to stay in a…

Probability · Mathematics 2015-01-26 Imade Fakhouri , Youssef Ouknine , Yong Ren

Learned representations in deep reinforcement learning (DRL) have to extract task-relevant information from complex observations, balancing between robustness to distraction and informativeness to the policy. Such stable and rich…

Machine Learning · Computer Science 2021-10-28 Mete Kemertas , Tristan Aumentado-Armstrong

A modified Deep BSDE (backward differential equation) learning method with measurability loss, called Deep BSDE-ML method, is introduced in this paper to solve a kind of linear decoupled forward-backward stochastic differential equations…

Optimization and Control · Mathematics 2022-01-06 Yutian Wang , Yuan-Hua Ni

Preference-based Reinforcement Learning (PbRL) circumvents the need for reward engineering by harnessing human preferences as the reward signal. However, current PbRL methods excessively depend on high-quality feedback from domain experts,…

Machine Learning · Computer Science 2024-10-29 Jie Cheng , Gang Xiong , Xingyuan Dai , Qinghai Miao , Yisheng Lv , Fei-Yue Wang

This article begins with a brief review of random matrix theory, followed by a discussion of how the large-$N$ limit of random matrix models can be realized using operator algebras. I then explain the notion of "Brown measure," which play…

Probability · Mathematics 2022-05-02 Brian C. Hall

This paper considers a portfolio optimization problem in which asset prices are represented by SDEs driven by Brownian motion and a Poisson random measure, with drifts that are functions of an auxiliary diffusion factor process. The…

Portfolio Management · Quantitative Finance 2010-11-16 Mark Davis , Sebastien Lleo

We unify and extend the semigroup and the PDE approaches to stochastic maximal regularity of time-dependent semilinear parabolic problems with noise given by a cylindrical Brownian motion. We treat random coefficients that are only…

Analysis of PDEs · Mathematics 2019-02-12 Pierre Portal , Mark Veraar

The principal aim of the present work is to explore limit theorems for small random perturbations of dynamical systems with periodic impulse effects, in the limit of vanishing noise intensity. We start with a system whose time evolution is…

Probability · Mathematics 2026-03-25 Ashif Khan , Chetan D. Pahlajani

This paper presents a novel approach to reinforcement learning (RL) for control systems that provides probabilistic stability guarantees using finite data. Leveraging Lyapunov's method, we propose a probabilistic stability theorem that…

Machine Learning · Computer Science 2026-03-03 Minghao Han , Lixian Zhang , Chenliang Liu , Zhipeng Zhou , Jun Wang , Wei Pan

We study the discrete-time linear-quadratic (LQ) control model using reinforcement learning (RL). Using entropy to measure the cost of exploration, we prove that the optimal feedback policy for the problem must be Gaussian type. Then, we…

Machine Learning · Statistics 2025-02-05 Lucky Li

Stochastic differential equations (SDEs) provide a flexible framework for modeling temporal dynamics in partially observed systems. A central task is to calibrate such models from data, which requires inferring latent trajectories and…

Machine Learning · Statistics 2026-05-08 Yu Wang , Arnab Ganguly