English
Related papers

Related papers: A random measure approach to reinforcement learnin…

200 papers

We propose a comprehensive framework for policy gradient methods tailored to continuous time reinforcement learning. This is based on the connection between stochastic control problems and randomised problems, enabling applications across…

Optimization and Control · Mathematics 2024-05-01 Robert Denkert , Huyên Pham , Xavier Warin

We propose a sparse grid stochastic collocation method for long-time simulations of stochastic differential equations (SDEs) driven by white noise. The method uses pre-determined sparse quadrature rules for the forcing term and constructs…

Numerical Analysis · Mathematics 2017-06-13 H. Cagan Ozen , Guillaume Bal

Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization signals from label noise. Through an empirical study, we observe that responses with medium…

Machine Learning · Computer Science 2026-04-24 Yongcan Yu , Lingxiao He , Jian Liang , Kuangpu Guo , Meng Wang , Qianlong Xie , Xingxing Wang , Ran He

In this paper, we study one-dimensional backward stochastic differential equation with jump under logarithmic growth assumption in the z-variable (|z|\sqrt{|\ln|z|}|) and an L^p terminal value (for a suitable p>2). We show the existence and…

Probability · Mathematics 2021-03-17 Khalid Oufdil

Reinforcement learning (RL) enables robots to learn skills from interactions with the real world. In practice, the unstructured step-based exploration used in Deep RL -- often very successful in simulation -- leads to jerky motion patterns…

Machine Learning · Computer Science 2022-09-16 Antonin Raffin , Jens Kober , Freek Stulp

Balancing exploration and exploitation is crucial in reinforcement learning (RL). In this paper, we study model-based posterior sampling for reinforcement learning (PSRL) in continuous state-action spaces theoretically and empirically.…

Machine Learning · Computer Science 2021-11-18 Ying Fan , Yifei Ming

This paper bridges reinforcement learning (RL) and risk-sensitive stochastic control by introducing a tractable exploration mechanism for policy search in risk-sensitive portfolio management, with known and unknown model parameters, that…

Portfolio Management · Quantitative Finance 2026-03-03 Sebastien Lleo , Wolfgang Runggaldier

We consider a unifying framework for stochastic control problem including the following features: partial observation, path-dependence (both with respect to the state and the control), and without any non-degeneracy condition on the…

Probability · Mathematics 2016-09-14 Elena Bandini , Andrea Cosso , Marco Fuhrman , Huyên Pham

We study the convergence of $N-$particle systems described by SDEs driven by Brownian motion and Poisson random measure, where the coefficients depend on the empirical measure of the system. Every particle jumps with a jump rate depending…

Probability · Mathematics 2021-03-09 Xavier Erny , Eva Löcherbach , Dasha Loukianova

Model-based reinforcement learning (MBRL) approaches rely on discrete-time state transition models whereas physical systems and the vast majority of control tasks operate in continuous-time. To avoid time-discretization approximation of the…

Machine Learning · Computer Science 2021-06-14 Çağatay Yıldız , Markus Heinonen , Harri Lähdesmäki

Inverse problems in scientific computing often require optimization over infinite-dimensional Hilbert spaces. A commonly used solver in such settings is stochastic gradient descent (SGD), where gradients are approximated using randomly…

Optimization and Control · Mathematics 2026-04-14 Sandra Cerrai , Qin Li , Anjali Nair , Jaeyoung Yoon

Time change is a powerful technique for generating noises and providing flexible models. In the framework of time changed Brownian and Poisson random measures we study the existence and uniqueness of a solution to a general mean-field…

Probability · Mathematics 2016-08-23 Giulia Di Nunno , Hannes Haferkorn

Stochastic policies (also known as relaxed controls) are widely used in continuous-time reinforcement learning algorithms. However, executing a stochastic policy and evaluating its performance in a continuous-time environment remain open…

Machine Learning · Computer Science 2025-10-03 Yanwei Jia , Du Ouyang , Yufei Zhang

Iterative self-training (self-distillation) repeatedly refits a model on pseudo-labels generated by its own predictions. We study this procedure in overparameterized linear regression: an initial estimator is trained on noisy labels, and…

Machine Learning · Statistics 2026-02-17 Mingqi Wu , Archer Y. Yang , Qiang Sun

In this article, we study the dynamics of a nonlinear system governed by an ordinary differential equation under the combined influence of fast periodic sampling with period $\delta$ and small jump noise of size $\varepsilon, 0<…

Probability · Mathematics 2024-11-28 Shivam Singh Dhama

Rewards play an essential role in reinforcement learning. In contrast to rule-based game environments with well-defined reward functions, complex real-world robotic applications, such as contact-rich manipulation, lack explicit and…

Machine Learning · Computer Science 2022-05-30 Yuning Wu , Jieliang Luo , Hui Li

In this paper we construct a framework for doing statistical inference for discretely observed stochastic differential equations (SDEs) where the driving noise has 'memory'. Classical SDE models for inference assume the driving noise to be…

Methodology · Statistics 2013-07-05 Martin Lysy , Natesh S. Pillai

This paper studies the continuous-time reinforcement learning for stochastic singular control with the application to an infinite-horizon irreversible reinsurance problems. The singular control is equivalently characterized as a pair of…

Optimization and Control · Mathematics 2025-12-03 Zongxia Liang , Xiaodong Luo , Xiang Yu

Regression discontinuity designs assess causal effects in settings where treatment is determined by whether an observed running variable crosses a pre-specified threshold. Here we propose a new approach to identification, estimation, and…

Methodology · Statistics 2025-04-01 Dean Eckles , Nikolaos Ignatiadis , Stefan Wager , Han Wu

Reinforcement Learning (RL) has proven effective in solving complex decision-making tasks across various domains, but challenges remain in continuous-time settings, particularly when state dynamics are governed by stochastic differential…

Machine Learning · Computer Science 2025-09-19 Chenyang Jiang , Donggyu Kim , Alejandra Quintos , Yazhen Wang