中文
相关论文

相关论文: Policy Optimization for Personalized Interventions…

200 篇论文

Dynamic treatment regimes are of growing interest across the clinical sciences as these regimes provide one way to operationalize and thus inform sequential personalized clinical decision making. A dynamic treatment regime is a sequence of…

统计方法学 · 统计学 2013-11-27 Eric B. Laber , Min Qian , Dan J. Lizotte , William E. Pelham , Susan A. Murphy

Multi-agent planning in stochastic environments can be framed formally as a decentralized Markov decision problem. Many real-life distributed problems that arise in manufacturing, multi-robot coordination and information gathering scenarios…

人工智能 · 计算机科学 2011-11-02 Claudia V. Goldman , Shlomo Zilberstein

Online artificial intelligence (AI) algorithms are an important component of digital health interventions. These online algorithms are designed to continually learn and improve their performance as streaming data is collected on…

计算机与社会 · 计算机科学 2025-10-29 Susobhan Ghosh , Bhanu T. Gullapalli , Daiqi Gao , Asim Gazi , Anna Trella , Ziping Xu , Kelly Zhang , Susan A. Murphy

Decentralized optimization is widely used in large scale and privacy preserving machine learning and various distributed control and sensing systems. It is assumed that every agent in the network possesses a local objective function, and…

A treatment policy defines when and what treatments are applied to affect some outcome of interest. Data-driven decision-making requires the ability to predict what happens if a policy is changed. Existing methods that predict how the…

机器学习 · 计算机科学 2023-06-21 Çağlar Hızlı , ST John , Anne Juuti , Tuure Saarinen , Kirsi Pietiläinen , Pekka Marttinen

We study the problem of learning personalized decision policies from observational data while accounting for possible unobserved confounding. Previous approaches, which assume unconfoundedness, i.e., that no unobserved confounders affect…

机器学习 · 计算机科学 2019-11-05 Nathan Kallus , Angela Zhou

Sample-based trajectory optimisers are a promising tool for the control of robotics with non-differentiable dynamics and cost functions. Contemporary approaches derive from a restricted subclass of stochastic optimal control where the…

机器人学 · 计算机科学 2021-10-07 Tom Lefebvre , Guillaume Crevecoeur

We address the personalized policy learning problem using longitudinal mobile health application usage data. Personalized policy represents a paradigm shift from developing a single policy that may prescribe personalized decisions by…

统计方法学 · 统计学 2020-01-13 Xinyu Hu , Min Qian , Bin Cheng , Ying Kuen Cheung

Addressing uncertainty is critical for autonomous systems to robustly adapt to the real world. We formulate the problem of model uncertainty as a continuous Bayes-Adaptive Markov Decision Process (BAMDP), where an agent maintains a…

机器人学 · 计算机科学 2019-05-09 Gilwoo Lee , Brian Hou , Aditya Mandalika , Jeongseok Lee , Sanjiban Choudhury , Siddhartha S. Srinivasa

We present a new approach to the problems of evaluating and learning personalized decision policies from observational data of past contexts, decisions, and outcomes. Only the outcome of the enacted decision is available and the historical…

机器学习 · 统计学 2019-06-04 Nathan Kallus

We consider approximate dynamic programming in $\gamma$-discounted Markov decision processes and apply it to approximate planning with linear value-function approximation. Our first contribution is a new variant of Approximate Policy…

机器学习 · 计算机科学 2022-10-31 Gellért Weisz , András György , Tadashi Kozuno , Csaba Szepesvári

Tuberculosis remains a significant global health challenge, with millions of new cases reported annually. Recent studies suggest that expanding the accessibility of TB intervention programs can lead to a substantial decrease in both TB…

Precision medicine has received attention both in and outside the clinic. We focus on the latter, by exploiting the relationship between individuals' social interactions and their mental health to develop a predictive model of one's…

社会与信息网络 · 计算机科学 2019-08-08 Shikang Liu , David Hachen , Omar Lizardo , Christian Poellabauer , Aaron Striegel , Tijana Milenkovic

We present a methodology to automatically compute worst-case performance bounds for a large class of first-order decentralized optimization algorithms. These algorithms aim at minimizing the average of local functions that are distributed…

最优化与控制 · 数学 2023-12-14 Sebastien Colla , Julien M. Hendrickx

Online feedback optimization is a controller design paradigm for optimizing the steady-state behavior of a dynamical system. It employs an optimization algorithm as a dynamic feedback controller and utilizes real-time measurements to bypass…

最优化与控制 · 数学 2024-04-01 Wenbin Wang , Zhiyu He , Giuseppe Belgioioso , Saverio Bolognani , Florian Dörfler

This paper studies a sequential decision-making problem in a two-stage queueing system modeled after operations in CVS MinuteClinics, where nurse practitioners (NPs) oversee patient care throughout the entire visit. All services are…

最优化与控制 · 数学 2026-03-11 Shuwen Lu , Mark E. Lewis , Jamol Pender

This article considers the minimization of the total number of infected individuals over the course of an epidemic in which the rate of infectious contacts can be reduced by time-dependent nonpharmaceutical interventions. The societal and…

最优化与控制 · 数学 2023-03-16 Tom Britton , Lasse Leskelä

On-policy reinforcement learning methods, like Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), often demand extensive data per update, leading to sample inefficiency. This paper introduces Reflective Policy…

机器学习 · 计算机科学 2024-06-07 Yaozhong Gan , Renye Yan , Zhe Wu , Junliang Xing

In this paper we propose an on-line policy iteration (PI) algorithm for finite-state infinite horizon discounted dynamic programming, whereby the policy improvement operation is done on-line, only for the states that are encountered during…

最优化与控制 · 数学 2021-06-03 Dimitri Bertsekas

We study an important yet under-addressed problem of quickly and safely improving policies in online reinforcement learning domains. As its solution, we propose a novel exploration strategy - diverse exploration (DE), which learns and…

机器学习 · 计算机科学 2018-02-26 Andrew Cohen , Lei Yu , Robert Wright
‹ 上一页 1 8 9 10 下一页 ›