中文
相关论文

相关论文: Markovian Interference in Experiments

200 篇论文

Estimation of social influence in networks can be substantially biased in observational studies due to homophily and network correlation in exposure to exogenous events. Randomized experiments, in which the researcher intervenes in the…

社会与信息网络 · 计算机科学 2017-09-28 Sean J. Taylor , Dean Eckles

We present an approach to reduce the communication required between agents in a Multi-Agent learning system by exploiting the inherent robustness of the underlying Markov Decision Process. We compute so-called robustness surrogate functions…

多智能体系统 · 计算机科学 2022-09-08 Daniel Jarne Ornia , Manuel Mazo

We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased…

机器学习 · 计算机科学 2018-10-31 Qiang Liu , Lihong Li , Ziyang Tang , Dengyong Zhou

The understanding of memory effects arising from the interaction between system and environment is a key for engineering quantum thermodynamic devices beyond the standard Markovian limit. We study the performance of measurement-based…

量子物理 · 物理学 2020-01-14 Obinna Abah , Mauro Paternostro

In this paper, we propose feedback designs for manipulating a quantum state to a target state by performing sequential measurements. In light of Belavkin's quantum feedback control theory, for a given set of (projective or non-projective)…

量子物理 · 物理学 2015-06-23 Shuangshuang Fu , Guodong Shi , Alexandre Proutiere , Matthew R. James

Off-policy evaluation (OPE) in reinforcement learning is an important problem in settings where experimentation is limited, such as education and healthcare. But, in these very same settings, observed actions are often confounded by…

机器学习 · 计算机科学 2020-07-29 Andrew Bennett , Nathan Kallus , Lihong Li , Ali Mousavi

This paper discusses the problem of estimating a stochastic signal from nonlinear uncertain observations with time-correlated additive noise described by a first-order Markov process. Random deception attacks are assumed to be launched by…

信号处理 · 电气工程与系统科学 2024-05-09 R. Caballero-Águila , J. Hu , J. Linares-Pérez

We develop methods for estimating how infinitesimal policy changes affect long-term outcomes in dynamic systems. We show that dynamic marginal policy effects (MPEs) can be identified via tractable reduced-form expressions, and can be…

统计方法学 · 统计学 2026-05-26 I-han Lai , Stefan Wager

In the context of modern environmental and societal concerns, there is an increasing demand for methods able to identify management strategies for civil engineering systems, minimizing structural failure risks while optimally planning…

In the maximum state entropy exploration framework, an agent interacts with a reward-free environment to learn a policy that maximizes the entropy of the expected state visitations it is inducing. Hazan et al. (2019) noted that the class of…

机器学习 · 计算机科学 2022-07-11 Mirco Mutti , Riccardo De Santi , Marcello Restelli

Standard estimators of the global average treatment effect can be biased in the presence of interference. This paper proposes regression adjustment estimators for removing bias due to interference in Bernoulli randomized experiments. We use…

统计方法学 · 统计学 2019-03-06 Alex Chin

Our work addresses a fundamental problem in the context of counterfactual inference for Markov Decision Processes (MDPs). Given an MDP path $\tau$, this kind of inference allows us to derive counterfactual paths $\tau'$ describing what-if…

人工智能 · 计算机科学 2025-03-28 Milad Kazemi , Jessica Lally , Ekaterina Tishchenko , Hana Chockler , Nicola Paoletti

Recent studies have greatly improved reinforcement learning, and an increased interest in real-world implementation has emerged. In many cases, the implementation is challenged by time-varying disturbances as it introduces hidden states,…

机器学习 · 计算机科学 2026-03-04 Saki Omi , Hyo-Sang Shin , Namhoon Cho , Antonios Tsourdos

We consider a longitudinal data structure consisting of baseline covariates, time-varying treatment variables, intermediate time-dependent covariates, and a possibly time dependent outcome. Previous studies have shown that estimating the…

统计理论 · 数学 2018-10-09 Linh Tran , Maya Petersen , Joshua Schwab , Mark J van der Laan

Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the historical data…

机器学习 · 计算机科学 2020-04-02 Tengyang Xie , Yifei Ma , Yu-Xiang Wang

Policy gradient methods are very attractive in reinforcement learning due to their model-free nature and convergence guarantees. These methods, however, suffer from high variance in gradient estimation, resulting in poor sample efficiency.…

机器学习 · 计算机科学 2018-11-16 Sergey Pankov

If an experimental treatment is experienced by both treated and control group units, tests of hypotheses about causal effects may be difficult to conceptualize let alone execute. In this paper, we show how counterfactual causal models may…

统计方法学 · 统计学 2012-08-03 Jake Bowers , Mark Fredrickson , Costas Panagopoulos

Quantum memory effects can be qualitatively understood as a consequence of an environment-to-system backflow of information. Here, we analyze and compare how this concept is interpreted and implemented in different approaches to quantum…

量子物理 · 物理学 2022-05-09 Adrián A. Budini

Multiple randomization designs (MRDs) are a class of experimental designs used to handle interference in two-sided marketplaces. We investigate regression adjustment strategies for estimating total, spillover, and direct effects in MRDs. We…

统计方法学 · 统计学 2026-03-23 Timothy Sudijono , Lihua Lei , Lorenzo Masoero , Suhas Vijaykumar , Guido Imbens , James McQueen

Experience replay is a core ingredient of modern deep reinforcement learning, yet its benefits in policy optimization are poorly understood beyond empirical heuristics. This paper develops a novel theoretical framework for experience replay…

机器学习 · 计算机科学 2026-02-04 Hua Zheng , Wei Xie , M. Ben Feng