中文
相关论文

相关论文: A survey of random processes with reinforcement

200 篇论文

Deep reinforcement learning has shown remarkable success in the past few years. Highly complex sequential decision making problems from game playing and robotics have been solved with deep model-free methods. Unfortunately, the sample…

机器学习 · 计算机科学 2021-07-20 Aske Plaat , Walter Kosters , Mike Preuss

In this paper, we present a comprehensive, in-depth survey of the literature on reinforcement learning approaches to decision optimization problems in a typical ridesharing system. Papers on the topics of rideshare matching, vehicle…

机器学习 · 计算机科学 2022-10-25 Zhiwei Qin , Hongtu Zhu , Jieping Ye

We introduce a theorem proving algorithm that uses practically no domain heuristics for guiding its connection-style proof search. Instead, it runs many Monte-Carlo simulations guided by reinforcement learning from previous proof attempts.…

人工智能 · 计算机科学 2018-05-22 Cezary Kaliszyk , Josef Urban , Henryk Michalewski , Mirek Olšák

Recent advances at the intersection of reinforcement learning (RL) and visual intelligence have enabled agents that not only perceive complex visual scenes but also reason, generate, and act within them. This survey offers a critical and…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Weijia Wu , Chen Gao , Joya Chen , Kevin Qinghong Lin , Qingwei Meng , Yiming Zhang , Yuke Qiu , Hong Zhou , Mike Zheng Shou

Empirical processes for stationary, causal sequences are considered. We establish empirical central limit theorems for classes of indicators of left half lines, absolutely continuous functions and piecewise differentiable functions. Sample…

统计理论 · 数学 2007-06-13 Wei Biao Wu

Probabilistic model checking is an approach to the formal modelling and analysis of stochastic systems. Over the past twenty five years, the number of different formalisms and techniques developed in this field has grown considerably, as…

计算机科学中的逻辑 · 计算机科学 2025-09-17 Marta Kwiatkowska , Gethin Norman , David Parker

We analyse a preferential urn model with randomness using the replica method. The preferential urn model is a stochastic model based on the concept "the rich get richer." The replica analysis clarifies that the preferential urn model with…

无序系统与神经网络 · 物理学 2009-11-11 Jun Ohkubo , Muneki Yasuda , Kazuyuki Tanaka

This paper develops generalizations of empowerment to continuous states. Empowerment is a recently introduced information-theoretic quantity motivated by hypotheses about the efficiency of the sensorimotor loop in biological organisms, but…

人工智能 · 计算机科学 2012-02-01 Tobias Jung , Daniel Polani , Peter Stone

Reinforcement learning methods have recently been very successful at performing complex sequential tasks like playing Atari games, Go and Poker. These algorithms have outperformed humans in several tasks by learning from scratch, using only…

机器学习 · 计算机科学 2021-09-28 Ajay Subramanian , Sharad Chitlangia , Veeky Baths

Challenging research in various fields has driven a wide range of methodological advances in variable selection for regression models with high-dimensional predictors. In comparison, selection of nonlinear functions in models with additive…

统计方法学 · 统计学 2013-03-05 Fabian Scheipl , Thomas Kneib , Ludwig Fahrmeir

Model-based Reinforcement Learning approaches have the promise of being sample efficient. Much of the progress in learning dynamics models in RL has been made by learning models via supervised learning. But traditional model-based…

机器学习 · 计算机科学 2019-06-12 Shagun Sodhani , Anirudh Goyal , Tristan Deleu , Yoshua Bengio , Sergey Levine , Jian Tang

We prove that the edge-reinforced random walk on the ladder ${\mathbb{Z}\times\{1,2\}}$ with initial weights $a>3/4$ is recurrent. The proof uses a known representation of the edge-reinforced random walk on a finite piece of the ladder as a…

概率论 · 数学 2007-05-23 Franz Merkl , Silke W. W. Rolles

Molecular dynamics simulations hold great promise for providing insight into the microscopic behavior of complex molecular systems. However, their effectiveness is often constrained by long timescales associated with rare events. Enhanced…

计算物理 · 物理学 2026-03-03 Kai Zhu , Enrico Trizio , Jintu Zhang , Renling Hu , Linlong Jiang , Tingjun Hou , Luigi Bonati

Experience replay is a key component in reinforcement learning for stabilizing learning and improving sample efficiency. Its typical implementation samples transitions with replacement from a replay buffer. In contrast, in supervised…

机器学习 · 计算机科学 2025-12-05 Yasuhiro Fujita

The evolution of grammatical systems of syntactic and semantic composition is modeled here with a novel application of reinforcement learning theory. To test the functionalist thesis that speakers' expressive purposes shape their language,…

计算与语言 · 计算机科学 2025-03-04 Stephen Wechsler , James W. Shearer , Katrin Erk

One of the biggest hurdles robotics faces is the facet of sophisticated and hard-to-engineer behaviors. Reinforcement learning offers a set of tools, and a framework to address this problem. In parallel, the misgivings of robotics offer a…

机器人学 · 计算机科学 2022-10-17 Akash Nagaraj , Mukund Sood , Bhagya M Patil

This work describes numerical methods that are useful in many areas: examples include statistical modelling (bioinformatics, computational biology), theoretical physics, and even pure mathematics. The methods are primarily useful for the…

数值分析 · 数学 2025-10-20 U. D. Jentschura , S. V. Aksenov , P. J. Mohr , M. A. Savageau , G. Soff

In a recent paper [2] the author introduced and investigated a random walk model similar to a model introduced in [1]. In these models the increment of the random walk depends on the complete past of the process. In this note I will point…

数据分析、统计与概率 · 物理学 2015-03-12 Rüdiger Kürsten

This report introduces general ideas and some basic methods of the Bayesian probability theory applied to physics measurements. Our aim is to make the reader familiar, through examples rather than rigorous formalism, with concepts such as:…

数据分析、统计与概率 · 物理学 2009-11-10 G. D'Agostini

Recent advances in steady-state analysis of power systems have introduced the equivalent split-circuit approach and corresponding continuation methods that can reliably find the correct physical solution of large-scale power system…

系统与控制 · 计算机科学 2018-04-24 Martin R. Wagner , Amritanshu Pandey , Marko Jereminov , Larry Pileggi