English
Related papers

Related papers: Reinforcement Learning, Optimal Control, and Bayes…

200 papers

We consider the inverse reinforcement learning problem, that is, the problem of learning from, and then predicting or mimicking a controller based on state/action data. We propose a statistical model for such data, derived from the…

Machine Learning · Statistics 2012-11-27 Sumeetpal S. Singh , Nicolas Chopin , Nick Whiteley

Reinforcement learning is commonly used with function approximation. However, very few positive results are known about the convergence of function approximation based RL control algorithms. In this paper we show that TD(0) and Sarsa(0)…

Machine Learning · Computer Science 2007-05-23 Istvan Szita , Andras Lorincz

Current methods for regularization in machine learning require quite specific model assumptions (e.g. a kernel shape) that are not derived from prior knowledge about the application, but must be imposed merely to make the method work. We…

Machine Learning · Statistics 2022-11-01 Matthias Wieler

We present a formalisation of finite Markov decision processes with rewards in the Isabelle theorem prover. We focus on the foundations required for dynamic programming and the use of reinforcement learning agents over such processes. In…

Logic in Computer Science · Computer Science 2021-12-14 Mark Chevallier , Jacques Fleuriot

Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inherited from pre-training. In this work, we introduce…

Machine Learning · Computer Science 2025-10-21 Jin Peng Zhou , Kaiwen Wang , Jonathan Chang , Zhaolin Gao , Nathan Kallus , Kilian Q. Weinberger , Kianté Brantley , Wen Sun

Sampling from complex target distributions is a challenging task fundamental to Bayesian inference. Parallel tempering (PT) addresses this problem by constructing a Markov chain on the expanded state space of a sequence of distributions…

Computation · Statistics 2023-01-18 Nikola Surjanovic , Saifuddin Syed , Alexandre Bouchard-Côté , Trevor Campbell

This study proposes a materials search method combining a data assimilation technique based on a multivariate Gaussian distribution with Bayesian optimization. The efficiency of the search using this method was demonstrated using a pair of…

In this paper, we present a novel algorithm named synchronous integral Q-learning, which is based on synchronous policy iteration, to solve the continuous-time infinite horizon optimal control problems of input-affine system dynamics. The…

Systems and Control · Electrical Eng. & Systems 2021-05-20 Lei Guo , Han Zhao

A wide range of machine learning algorithms iteratively add data to the training sample. Examples include semi-supervised learning, active learning, multi-armed bandits, and Bayesian optimization. We embed this kind of data addition into…

Machine Learning · Statistics 2024-06-25 Julian Rodemann

This paper employs a policy iteration reinforcement learning (RL) method to study continuous-time linear-quadratic mean-field control problems in infinite horizon. The drift and diffusion terms in the dynamics involve the states, the…

Optimization and Control · Mathematics 2024-11-05 Na Li , Xun Li , Zuo Quan Xu

The long-timescale behavior of complex dynamical systems can be described by linear Markov or Koopman models in a suitable latent space. Recent variational approaches allow the latent space representation and the linear dynamical model to…

Computational Physics · Physics 2019-12-17 Andreas Mardt , Luca Pasquali , Frank Noé , Hao Wu

The control of spatio-temporally chaos is challenging because of high dimensionality and unpredictability. Model-free reinforcement learning (RL) discovers optimal control policies by interacting with the system, typically requiring…

Systems and Control · Electrical Eng. & Systems 2026-01-13 Defne E. Ozan , Andrea Nóvoa , Georgios Rigas , Luca Magri

We address the inverse problem of cosmic large-scale structure reconstruction from a Bayesian perspective. For a linear data model, a number of known and novel reconstruction schemes, which differ in terms of the underlying signal prior,…

Astrophysics · Physics 2009-11-06 F. S. Kitaura , T. A. Ensslin

We consider a task of surveillance-evading path-planning in a continuous setting. An Evader strives to escape from a 2D domain while minimizing the risk of detection (and immediate capture). The probability of detection is path-dependent…

Machine Learning · Computer Science 2023-02-24 Dongping Qi , David Bindel , Alexander Vladimirsky

Sequential Bayesian filters in non-linear dynamic systems require the recursive estimation of the predictive and posterior distributions. This paper introduces a Bayesian filter called the adaptive kernel Kalman filter (AKKF). With this…

Signal Processing · Electrical Eng. & Systems 2023-04-12 Mengwei Sun , Mike E. Davies , Ian K. Proudler , James R. Hopgood

Koopman operator theory is a key tool in data assimilation of complex dynamical systems, with the potential to be applied to multimodal data. We formulate the problem of learning Koopman eigenfunctions from observations at arbitrary,…

Systems and Control · Electrical Eng. & Systems 2026-04-14 Younghwan Cho , Richard Sowers

Reinforcement learning (RL) is frequently employed in fine-tuning large language models (LMs), such as GPT-3, to penalize them for undesirable features of generated sequences, such as offensiveness, social bias, harmfulness or falsehood.…

Machine Learning · Computer Science 2022-10-24 Tomasz Korbak , Ethan Perez , Christopher L Buckley

Normalizing flows can generate complex target distributions and thus show promise in many applications in Bayesian statistics as an alternative or complement to MCMC for sampling posteriors. Since no data set from the target posterior…

Machine Learning · Statistics 2021-07-19 Marylou Gabrié , Grant M. Rotskoff , Eric Vanden-Eijnden

We show that several major algorithms of reinforcement learning (RL) fit into the framework of categorical cybernetics, that is to say, parametrised bidirectional processes. We build on our previous work in which we show that value…

Machine Learning · Computer Science 2025-09-26 Jules Hedges , Riu Rodríguez Sakamoto

Recent progress in reinforcement learning has led to remarkable performance in a range of applications, but its deployment in high-stakes settings remains quite rare. One reason is a limited understanding of the behavior of reinforcement…

Machine Learning · Computer Science 2020-11-04 Feicheng Wang , Lucas Janson
‹ Prev 1 3 4 5 6 7 10 Next ›