中文
相关论文

相关论文: Finding Counterfactually Optimal Action Sequences …

200 篇论文

Multi-objective Markov decision processes are sequential decision-making problems that involve multiple conflicting reward functions that cannot be optimized simultaneously without a compromise. This type of problems cannot be solved by a…

机器学习 · 计算机科学 2023-08-22 Sherif Abdelfattah , Kathryn Merrick , Jiankun Hu

How do decisions change with the economic environment and with time? This paper studies general nonstationary stopping problems and provides the methodological tools to answer these questions. First, we identify conditions that ensure a…

理论经济学 · 经济学 2024-08-01 Théo Durandard , Matteo Camboni

Consider a subject or unit in a longitudinal biomedical, public health, engineering, economic, or social science study which is being monitored over a possibly random duration. Over time this unit experiences competing recurrent events and…

统计方法学 · 统计学 2024-12-30 Lili Tong , Piaomu Liu , Edsel Pena

Most approaches to visual scene analysis have emphasised parallel processing of the image elements. However, one area in which the sequential nature of vision is apparent, is that of segmenting multiple, potentially similar and partially…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Nikita Araslanov , Constantin Rothkopf , Stefan Roth

Consider a setting in which a policy maker assigns subjects to treatments, observing each outcome before the next subject arrives. Initially, it is unknown which treatment is best, but the sequential nature of the problem permits learning…

计量经济学 · 经济学 2020-08-13 Anders Bredahl Kock , David Preinerstorfer , Bezirgen Veliyev

We show that combinations of optimal (stationary) policies in unichain Markov decision processes are optimal. That is, let M be a unichain Markov decision process with state space S, action space A and policies \pi_j^*: S -> A (1\leq j\leq…

组合数学 · 数学 2007-05-23 Ronald Ortner

Sequential learning systems are used in a wide variety of problems from decision making to optimization, where they provide a 'belief' (opinion) to nature, and then update this belief based on the feedback (result) to minimize (or maximize)…

机器学习 · 计算机科学 2020-09-22 Kaan Gokcesu , Hakan Gokcesu

Conventional methods for query autocompletion aim to predict which completed query a user will select from a list. A shortcoming of this approach is that users often do not know which query will provide the best retrieval performance on the…

信息检索 · 计算机科学 2022-04-26 Adam Block , Rahul Kidambi , Daniel N. Hill , Thorsten Joachims , Inderjit S. Dhillon

Dynamic treatment regimes operationalize the clinical decision process as a sequence of functions, one for each clinical decision, where each function takes as input up-to-date patient information and gives as output a single recommended…

统计方法学 · 统计学 2012-08-08 Eric B. Laber , Daniel J. Lizotte , Bradley Ferguson

The paper focuses on identifying the causes of student performance to provide personalized recommendations for improving pass rates. We introduce the need to move beyond predictive models and instead identify causal relationships. We…

计算机与社会 · 计算机科学 2023-09-26 Bevan I. Smith

Observational longitudinal studies are a common means to study treatment efficacy and safety in chronic mental illness. In many such studies, treatment changes may be initiated by either the patient or by their clinician and can thus vary…

统计方法学 · 统计学 2020-06-12 Zekun Xu , Eric Laber , Ana-Maria Staicu , Emanuel Severus

We study decision timing problems on finite horizon with Poissonian information arrivals. In our model, a decision maker wishes to optimally time her action in order to maximize her expected reward. The reward depends on an unobservable…

最优化与控制 · 数学 2012-05-07 Michael Ludkovski , Semih Sezer

We consider after-study statistical inference for sequentially designed experiments wherein multiple units are assigned treatments for multiple time points using treatment policies that adapt over time. Our goal is to provide inference…

机器学习 · 统计学 2025-06-10 Raaz Dwivedi , Katherine Tian , Sabina Tomkins , Predrag Klasnja , Susan Murphy , Devavrat Shah

A combination of deep reinforcement learning and supervised learning is proposed for the problem of active sequential hypothesis testing in completely unknown environments. We make no assumptions about the prior probability, the action and…

人工智能 · 计算机科学 2023-06-07 George Stamatelis , Nicholas Kalouptsidis

Estimating the counterfactual outcome of treatment is essential for decision-making in public health and clinical science, among others. Often, treatments are administered in a sequential, time-varying manner, leading to an exponentially…

机器学习 · 统计学 2024-07-16 Shenghao Wu , Wenbin Zhou , Minshuo Chen , Shixiang Zhu

This paper introduces a novel causal framework for multi-stage decision-making in natural language action spaces where outcomes are only observed after a sequence of actions. While recent approaches like Proximal Policy Optimization (PPO)…

计算与语言 · 计算机科学 2025-02-26 Bohan Zhang , Yixin Wang , Paramveer S. Dhillon

We consider the batch (off-line) policy learning problem in the infinite horizon Markov Decision Process. Motivated by mobile health applications, we focus on learning a policy that maximizes the long-term average reward. We propose a…

统计理论 · 数学 2022-09-20 Peng Liao , Zhengling Qi , Runzhe Wan , Predrag Klasnja , Susan Murphy

We consider a multidimensional search problem that is motivated by questions in contextual decision-making, such as dynamic pricing and personalized medicine. Nature selects a state from a $d$-dimensional unit ball and then generates a…

数据结构与算法 · 计算机科学 2017-04-27 Ilan Lobel , Renato Paes Leme , Adrian Vladu

Decision-makers are faced with the challenge of estimating what is likely to happen when they take an action. For instance, if I choose not to treat this patient, are they likely to die? Practitioners commonly use supervised learning…

机器学习 · 统计学 2018-02-02 Peter Schulam , Suchi Saria

The problem of making sequential decisions in unknown probabilistic environments is studied. In cycle $t$ action $y_t$ results in perception $x_t$ and reward $r_t$, where all quantities in general may depend on the complete history. The…

人工智能 · 计算机科学 2007-05-23 Marcus Hutter