中文
相关论文

相关论文: Algorithms for reachability problems on stochastic…

200 篇论文

Reward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing test-time compute to enhance reward model performance. In…

计算与语言 · 计算机科学 2025-05-21 Jiaxin Guo , Zewen Chi , Li Dong , Qingxiu Dong , Xun Wu , Shaohan Huang , Furu Wei

In this paper, we provide a new algorithm for the problem of prediction in Reinforcement Learning, \emph{i.e.}, estimating the Value Function of a Markov Reward Process (MRP) using the linear function approximation architecture, with memory…

系统与控制 · 计算机科学 2016-09-30 Ajin George Joseph , Shalabh Bhatnagar

Gaussian Mixture Models (GMMs) are one of the most potent parametric density models used extensively in many applications. Flexibly-tied factorization of the covariance matrices in GMMs is a powerful approach for coping with the challenges…

机器学习 · 计算机科学 2023-11-14 Mohammad Pasande , Reshad Hosseini , Babak Nadjar Araabi

Given a unichain Markov reward process (MRP), we provide an explicit expression for the bias values in terms of mean first passage times. This result implies a generalization of known Markov chain perturbation bounds for the stationary…

概率论 · 数学 2024-08-09 Ronald Ortner

We consider the task of estimating a structural model of dynamic decisions by a human agent based upon the observable history of implemented actions and visited states. This problem has an inherent nested structure: in the inner problem, an…

机器学习 · 计算机科学 2024-03-04 Siliang Zeng , Mingyi Hong , Alfredo Garcia

We present a general framework for applying learning algorithms and heuristical guidance to the verification of Markov decision processes (MDPs). The primary goal of our techniques is to improve performance by avoiding an exhaustive…

Constrained decision-making is essential for designing safe policies in real-world control systems, yet simulated environments often fail to capture real-world adversities. We consider the problem of learning a policy that will maximize the…

机器学习 · 计算机科学 2026-02-10 Sourav Ganguly , Kishan Panaganti , Arnob Ghosh , Adam Wierman

We introduce a novel framework to account for sensitivity to rewards uncertainty in sequential decision-making problems. While risk-sensitive formulations for Markov decision processes studied so far focus on the distribution of the…

机器学习 · 计算机科学 2020-09-16 Nelson Vadori , Sumitra Ganesh , Prashant Reddy , Manuela Veloso

Motivated by practical applications where stable long-term performance is critical-such as robotics, operations research, and healthcare-we study the problem of distributionally robust (DR) average-reward reinforcement learning. We propose…

机器学习 · 计算机科学 2026-02-03 Zijun Chen , Shengbo Wang , Nian Si

Markov automata combine non-determinism, probabilistic branching, and exponentially distributed delays. This compositional variant of continuous-time Markov decision processes is used in reliability engineering, performance evaluation and…

计算机科学中的逻辑 · 计算机科学 2017-05-11 Tim Quatmann , Sebastian Junges , Joost-Pieter Katoen

We provide a framework for speeding up algorithms for time-bounded reachability analysis of continuous-time Markov decision processes. The principle is to find a small, but almost equivalent subsystem of the original system and only analyse…

系统与控制 · 计算机科学 2018-07-26 Pranav Ashok , Yuliya Butkova , Holger Hermanns , Jan Křetínský

Motivated by robotic surveillance applications, this paper studies the novel problem of maximizing the return time entropy of a Markov chain, subject to a graph topology with travel times and stationary distribution. The return time entropy…

最优化与控制 · 数学 2018-05-29 Xiaoming Duan , Mishel George , Francesco Bullo

The success of reinforcement learning in typical settings is predicated on Markovian assumptions on the reward signal by which an agent learns optimal policies. In recent years, the use of reward machines has relaxed this assumption by…

机器学习 · 计算机科学 2022-03-29 Taylor Dohmen , Noah Topper , George Atia , Andre Beckus , Ashutosh Trivedi , Alvaro Velasquez

Suppose an online platform wants to compare a treatment and control policy, e.g., two different matching algorithms in a ridesharing system, or two different inventory management algorithms in an online retail site. Standard randomized…

统计方法学 · 统计学 2022-12-27 Peter Glynn , Ramesh Johari , Mohammad Rasouli

Fluctuations in stochastic systems are usually characterized by the full counting statistics, which analyzes the distribution of the number of events taking place in the fixed time interval. In an alternative approach, the distribution of…

统计力学 · 物理学 2018-01-24 Krzysztof Ptaszynski

Markov chains are the de facto finite-state model for stochastic dynamical systems, and Markov decision processes (MDPs) extend Markov chains by incorporating non-deterministic behaviors. Given an MDP and rewards on states, a classical…

计算机科学中的逻辑 · 计算机科学 2024-11-13 Krishnendu Chatterjee , Laurent Doyen

Markov decision processes (MDPs) are widely used in modeling decision making problems in stochastic environments. However, precise specification of the reward functions in MDPs is often very difficult. Recent approaches have focused on…

人工智能 · 计算机科学 2012-02-20 Eunsoo Oh , Kee-Eung Kim

We present a novel method for computing reachability probabilities of parametric discrete-time Markov chains whose transition probabilities are fractions of polynomials over a set of parameters. Our algorithm is based on two key…

We present a new algorithm for the statistical model checking of Markov chains with respect to unbounded temporal properties, such as reachability and full linear temporal logic. The main idea is that we monitor each simulation run on the…

计算机科学中的逻辑 · 计算机科学 2016-03-04 Przemysław Daca , Thomas A. Henzinger , Jan Křetínský , Tatjana Petrov

Fastest arrival events, where the first among many diffusing particles reaches a target, are central in triggering signal initiation in molecular stochastic systems. Classical approaches to simulate such events rely on full trajectory…

概率论 · 数学 2026-05-26 Emmanuel Akame Mfoumou , David Holcman