中文
相关论文

相关论文: Model-based Bootstrap of Controlled Markov Chains

200 篇论文

Predictability of behavior has emerged an an important characteristic in many fields including biology, medicine, and marketing. Behavior can be recorded as a sequence of actions performed by an individual over a given time period. This…

统计方法学 · 统计学 2017-11-13 Brian Vegetabile , Jenny Molet , Tallie Z. Baram , Hal Stern

Background and Objective: Uncertainty in non-linear mixed effect models is often assessed using the Fisher information matrix to derive the standard errors of estimation. The bootstrap is an alternative to the asymptotic method, with…

统计方法学 · 统计学 2026-05-05 Sofia Kaisaridi , Moreno Ursino , Emmanuelle Comets

A robust model predictive control scheme for a class of constrained norm-bounded uncertain discrete-time linear systems is developed under the hypothesis that only partial state measurements are available for feedback. Off-line calculations…

系统与控制 · 计算机科学 2018-07-23 Giuseppe Franzè , Massimiliano Mattei , Luciano Ollio , Valerio Scordamaglia

Optimizing or sampling complex cost functions of combinatorial optimization problems is a longstanding challenge across disciplines and applications. When employing family of conventional algorithms based on Markov Chain Monte Carlo (MCMC)…

机器学习 · 计算机科学 2025-08-15 Dmitrii Dobrynin , Masoud Mohseni , John Paul Strachan

Complex mechanical systems such as vehicle powertrains are inherently subject to multiple nonlinearities and uncertainties arising from parametric variations. Modeling errors are therefore unavoidable, making the transfer of control systems…

系统与控制 · 电气工程与系统科学 2026-02-13 Heisei Yonezawa , Ansei Yonezawa , Itsuro Kajiwara

Model-based offline reinforcement learning is brittle under distribution shift: policy improvement drives rollouts into state--action regions weakly supported by the dataset, where compounding model error yields severe value overestimation.…

机器学习 · 计算机科学 2026-02-04 Zeyu Fang , Zuyuan Zhang , Mahdi Imani , Tian Lan

Recent advancements in offline reinforcement learning (RL) have underscored the capabilities of Conditional Sequence Modeling (CSM), a paradigm that learns the action distribution based on history trajectory and target returns for each…

机器学习 · 计算机科学 2024-05-28 Shengchao Hu , Ziqing Fan , Chaoqin Huang , Li Shen , Ya Zhang , Yanfeng Wang , Dacheng Tao

We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under drifting non-stationarity, i.e., both the reward and state transition distributions are allowed to evolve over time, as long as their respective…

机器学习 · 计算机科学 2020-06-26 Wang Chi Cheung , David Simchi-Levi , Ruihao Zhu

We present a learning model predictive control (MPC) scheme for chance-constrained Markov jump systems with unknown switching probabilities. Using samples of the underlying Markov chain, ambiguity sets of transition probabilities are…

最优化与控制 · 数学 2023-01-06 Mathijs Schuurmans , Panagiotis Patrinos

We study the sequential decision-making problem for automated weaning of mechanical circulatory support (MCS) devices in cardiogenic shock patients. MCS devices are percutaneous micro-axial flow pumps that provide left ventricular unloading…

机器学习 · 计算机科学 2025-11-11 Aysin Tumay , Sophia Sun , Sonia Fereidooni , Aaron Dumas , Elise Jortberg , Rose Yu

Many practical applications of reinforcement learning (RL) constrain the agent to learn from a fixed offline dataset of logged interactions, which has already been gathered, without offering further possibility for data collection. However,…

机器学习 · 计算机科学 2021-07-06 Zizhou Su

Bootstrap is an idea that imposing consistency conditions on a physical system may lead to rigorous and nontrivial statements about its physical observables. In this work, we discuss the bootstrap problem for the invariant measure of the…

高能物理 - 理论 · 物理学 2023-10-24 Minjae Cho , Xin Sun

We introduce a bootstrap procedure for high-frequency statistics of Brownian semistationary processes. More specifically, we focus on a hypothesis test on the roughness of sample paths of Brownian semistationary processes, which uses an…

统计理论 · 数学 2021-01-06 Mikkel Bennedsen , Ulrich Hounyo , Asger Lunde , Mikko S. Pakkanen

We study off-dynamics Reinforcement Learning (RL), where the policy is trained on a source domain and deployed to a distinct target domain. We aim to solve this problem via online distributionally robust Markov decision processes (DRMDPs),…

机器学习 · 计算机科学 2024-02-26 Zhishuai Liu , Pan Xu

This article presents a bootstrap approximation to the Lp_statistics of kernel density estimator in length-biased model. Length-biased data arise in many situations, such as survival analysis, renewal processes and physics. The article…

概率论 · 数学 2017-05-30 Raheleh Zamini

We propose a bootstrap-based calibrated projection procedure to build confidence intervals for single components and for smooth functions of a partially identified parameter vector in moment (in)equality models. The method controls…

统计理论 · 数学 2024-07-03 Hiroaki Kaido , Francesca Molinari , Jörg Stoye

Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift. The learned policy may visit out-of-distribution state-action pairs…

人工智能 · 计算机科学 2026-03-17 Hongqiang Lin , Zhenghui Fu , Weihao Tang , Pengfei Wang , Yiding Sun , Qixian Huang , Dongxu Zhang

In this paper, we develop a general law of large numbers and central limit theorem for cumulative reward processes associated with finite state Markov jump processes with non-stationary transition rates. Such models commonly arise in…

概率论 · 数学 2025-10-15 Monte Fischer , Peter W. Glynn

We introduce a class of Adapted Increasingly Rarely Markov Chain Monte Carlo (AirMCMC) algorithms where the underlying Markov kernel is allowed to be changed based on the whole available chain output but only at specific time points…

统计计算 · 统计学 2018-01-30 Cyril Chimisov , Krzysztof Latuszynski , Gareth Roberts

In this paper, we present a nonlinear robust model predictive control (MPC) framework for general (state and input dependent) disturbances. This approach uses an online constructed tube in order to tighten the nominal (state and input)…

系统与控制 · 电气工程与系统科学 2020-06-05 Johannes Köhler , Raffaele Soloperto , Matthias A. Müller , Frank Allgöwer