中文
相关论文

相关论文: Policy Guided Monte Carlo: Reinforcement Learning …

200 篇论文

Markov chain Monte Carlo (MCMC) is a sampling-based method for estimating features of probability distributions. MCMC methods produce a serially correlated, yet representative, sample from the desired distribution. As such it can be…

统计计算 · 统计学 2019-12-10 Dootika Vats , Nathan Robertson , James M Flegal , Galin L Jones

Despite high reliability, modern power systems with growing renewable penetration face an increasing risk of cascading outages. Real-time cascade mitigation requires fast, complex operational decisions under uncertainty. In this work, we…

物理与社会 · 物理学 2025-06-11 Kai Zhou , Youbiao He , Chong Zhong , Yifu Wu

Sampling the three-dimensional (3D) spin glass -- i.e., generating equilibrium configurations of a 3D lattice with quenched random couplings -- is widely regarded as one of the central and long-standing open problems in statistical physics.…

统计力学 · 物理学 2025-09-30 Tao Chen , Jing Liu , Youjin Deng , Pan Zhang

Monte Carlo simulation is an unbiased numerical tool for studying classical and quantum many-body systems. One of its bottlenecks is the lack of general and efficient update algorithm for large size systems close to phase transition or with…

强关联电子 · 物理学 2017-01-11 Junwei Liu , Yang Qi , Zi Yang Meng , Liang Fu

Model-free Reinforcement Learning (RL) works well when experience can be collected cheaply and model-based RL is effective when system dynamics can be modeled accurately. However, both assumptions can be violated in real world problems such…

机器学习 · 计算机科学 2020-05-07 Mohak Bhardwaj , Ankur Handa , Dieter Fox , Byron Boots

Robotic systems must be able to quickly and robustly make decisions when operating in uncertain and dynamic environments. While Reinforcement Learning (RL) can be used to compute optimal policies with little prior knowledge about the…

机器人学 · 计算机科学 2016-09-13 Yunpeng Pan , Xinyan Yan , Evangelos Theodorou , Byron Boots

Markov chain Monte Carlo (MCMC) algorithms are indispensable when sampling from a complex, high-dimensional distribution by a conventional method is intractable. Even though MCMC is a powerful tool, it is also hard to control and tune in…

图形学 · 计算机科学 2025-10-14 Sascha Holl , Gurprit Singh , Hans-Peter Seidel

We propose a Markov chain Monte Carlo (MCMC) scheme to perform state inference in non-linear non-Gaussian state-space models. Current state-of-the-art methods to address this problem rely on particle MCMC techniques and its variants, such…

统计计算 · 统计学 2019-05-15 Alexander Y. Shestopaloff , Arnaud Doucet

Linear Temporal Logic (LTL) is widely used to specify high-level objectives for system policies, and it is highly desirable for autonomous systems to learn the optimal policy with respect to such specifications. However, learning the…

机器学习 · 计算机科学 2023-10-26 Daqian Shao , Marta Kwiatkowska

We investigate Monte Carlo based algorithms for solving stochastic control problems with probabilistic constraints. Our motivation comes from microgrid management, where the controller tries to optimally dispatch a diesel generator while…

最优化与控制 · 数学 2024-02-06 Alessandro Balata , Michael Ludkovski , Aditya Maheshwari , Jan Palczewski

We propose a method to encourage safety in Model Predictive Control (MPC)-based Reinforcement Learning (RL) via Gaussian Process (GP) regression. This framework consists of 1) a parametric MPC scheme that is employed as model-based…

系统与控制 · 电气工程与系统科学 2024-12-13 Filippo Airaldi , Bart De Schutter , Azita Dabiri

Integration of reinforcement learning and imitation learning is an important problem that has been studied for a long time in the field of intelligent robotics. Reinforcement learning optimizes policies to maximize the cumulative reward,…

机器学习 · 计算机科学 2023-01-18 Akira Kinose , Tadahiro Taniguchi

In this paper, we build and explore supervised learning models of ferromagnetic system behavior, using Monte-Carlo sampling of the spin configuration space generated by the 2D Ising model. Given the enormous size of the space of all…

统计力学 · 物理学 2017-09-06 Nataliya Portman , Isaac Tamblyn

We develop off-lattice simulations of semiflexible polymer chains subjected to applied mechanical forces using Markov Chain Monte Carlo. Our approach models the polymer as a chain of fixed-length bonds, with configurations updated through…

软凝聚态物质 · 物理学 2024-11-26 Lijie Ding , Chi-Huan Tung , Bobby G. Sumpter , Wei-Ren Chen , Changwoo Do

This paper addresses the problem of training a reinforcement learning (RL) policy under partial observability by exploiting a privileged, anytime-feasible planner agent available exclusively during training. We formalize this as a Partially…

机器学习 · 计算机科学 2026-04-10 Mohsen Amiri , Mohsen Amiri , Ali Beikmohammadi , Sindri Magnuśson , Mehdi Hosseinzadeh

Markov chain Monte Carlo (MCMC) methods are ubiquitous tools for simulation-based inference in many fields but designing and identifying good MCMC samplers is still an open question. This paper introduces a novel MCMC algorithm, namely,…

Most successful applications of deep learning involve similar training and test conditions. However, tasks such as biological sequence design involve searching for sequences that improve desirable properties beyond previously known values,…

机器学习 · 计算机科学 2025-05-27 Sophia Hager , Aleem Khan , Andrew Wang , Nicholas Andrews

In this paper, we consider solving discounted Markov Decision Processes (MDPs) under the constraint that the resulting policy is stabilizing. In practice MDPs are solved based on some form of policy approximation. We will leverage recent…

机器学习 · 计算机科学 2021-02-03 Mario Zanon , Sébastien Gros , Michele Palladino

Efficient sampling of complex high-dimensional probability distributions is a central task in computational science. Machine learning methods like autoregressive neural networks, used with Markov chain Monte Carlo sampling, provide good…

统计力学 · 物理学 2021-11-11 Dian Wu , Riccardo Rossi , Giuseppe Carleo

In this article we consider computing expectations w.r.t.~probability laws associated to a certain class of stochastic systems. In order to achieve such a task, one must not only resort to numerical approximation of the expectation, but…

统计计算 · 统计学 2017-10-30 Ajay Jasra , Kengo Kamatani , Kody Law , Yan Zhou