English

A Reinforcement Learning Approach to Estimating Long-term Treatment Effects

Machine Learning 2022-10-17 v1 Methodology Machine Learning

Abstract

Randomized experiments (a.k.a. A/B tests) are a powerful tool for estimating treatment effects, to inform decisions making in business, healthcare and other applications. In many problems, the treatment has a lasting effect that evolves over time. A limitation with randomized experiments is that they do not easily extend to measure long-term effects, since running long experiments is time-consuming and expensive. In this paper, we take a reinforcement learning (RL) approach that estimates the average reward in a Markov process. Motivated by real-world scenarios where the observed state transition is nonstationary, we develop a new algorithm for a class of nonstationary problems, and demonstrate promising results in two synthetic datasets and one online store dataset.

Keywords

Cite

@article{arxiv.2210.07536,
  title  = {A Reinforcement Learning Approach to Estimating Long-term Treatment Effects},
  author = {Ziyang Tang and Yiheng Duan and Stephanie Zhang and Lihong Li},
  journal= {arXiv preprint arXiv:2210.07536},
  year   = {2022}
}
R2 v1 2026-06-28T03:37:12.570Z