中文
相关论文

相关论文: Q-learning with censored data

200 篇论文

We study the adaptive control of an unknown linear system with a quadratic cost function subject to safety constraints on both the states and actions. The challenges of this problem arise from the tension among safety, exploration,…

系统与控制 · 电气工程与系统科学 2021-11-02 Yingying Li , Subhro Das , Jeff Shamma , Na Li

In this paper, we investigate dynamic feature selection within multivariate time-series scenario, a common occurrence in clinical prediction monitoring where each feature corresponds to a bio-test result. Many existing feature selection…

机器学习 · 计算机科学 2024-05-31 Yutong Chen , Jiandong Gao , Ji Wu

This paper develops a quantized Q-learning algorithm for the optimal control of controlled diffusion processes on $\mathbb{R}^d$ under both discounted and ergodic (average) cost criteria. We first establish near-optimality of finite-state…

最优化与控制 · 数学 2026-03-16 Erhan Bayraktar , Ali D. Kara , Somnath Pradhan , Serdar Yuksel

The distribution-free method of conformal prediction (Vovk et al, 2005) has gained considerable attention in computer science, machine learning, and statistics. Candes et al. (2023) extended this method to right-censored survival data,…

统计方法学 · 统计学 2025-06-04 Jing Qin , Jin Piao , Jing Ning , Yu Shen

Q-functions are widely used in discrete-time learning and control to model future costs arising from a given control policy, when the initial state and input are given. Although some of their properties are understood, Q-functions…

最优化与控制 · 数学 2019-02-21 Joseph Warrington

Drawing the quantum phase diagram of a many-body system in the parameter space of its Hamiltonian can be seen as a learning problem, which implies labelling the corresponding ground states according to some classification criterium that…

量子物理 · 物理学 2025-10-17 Mehran Khosrojerdi , Alessandro Cuccoli , Paola Verrucchi , Leonardo Banchi

In this paper we address distributed learning problems over peer-to-peer networks. In particular, we focus on the challenges of quantized communications, asynchrony, and stochastic gradients that arise in this set-up. We first discuss how…

最优化与控制 · 数学 2025-09-04 Nicola Bastianello , Apostolos I. Rikos , Karl H. Johansson

In this paper, a tractable methodology is proposed to approximate stochastic optimal feedback treatment in the context of mixed immuno-chemo therapy of cancer. The method uses a fixed-point value iteration that approximately solves a…

系统与控制 · 计算机科学 2020-10-26 Mazen Alamir

Compared to on-policy counterparts, off-policy model-free deep reinforcement learning can improve data efficiency by repeatedly using the previously gathered data. However, off-policy learning becomes challenging when the discrepancy…

机器学习 · 计算机科学 2023-09-27 Baturay Saglam , Dogan C. Cicek , Furkan B. Mutlu , Suleyman S. Kozat

Most decision-focused learning work has focused on single stage problems whereas many real-world decision problems are more appropriately modelled using multistage optimisation. In multistage problems contextual information is revealed over…

最优化与控制 · 数学 2025-05-29 Egon Peršak , Miguel F. Anjos

To know which operators to apply and in which order, as well as attributing good values to their parameters is a challenge for users of computer vision. This paper proposes a solution to this problem as a multi-agent system modeled…

人工智能 · 计算机科学 2013-11-26 Issam Qaffou , Mohamed Sadgal , Abdelaziz Elfazziki

Individualized treatment rules can lead to better health outcomes when patients have heterogeneous responses to treatment. Very few individualized treatment rule estimation methods are compatible with a multi-treatment observational study…

统计方法学 · 统计学 2019-11-14 Owen E. Leete , Nathan Kallus , Michael G. Hudgens , Sonia Napravnik , Michael R. Kosorok

The conditional survival function of a time-to-event outcome subject to censoring and truncation is a common target of estimation in survival analysis. This parameter may be of scientific interest and also often appears as a nuisance in…

统计方法学 · 统计学 2024-08-20 Charles J. Wolock , Peter B. Gilbert , Noah Simon , Marco Carone

A temporally abstract action, or an option, is specified by a policy and a termination condition: the policy guides option behavior, and the termination condition roughly determines its length. Generally, learning with longer options (like…

人工智能 · 计算机科学 2017-12-05 Anna Harutyunyan , Peter Vrancx , Pierre-Luc Bacon , Doina Precup , Ann Nowe

Reinforcement learning techniques achieved human-level performance in several tasks in the last decade. However, in recent years, the need for interpretability emerged: we want to be able to understand how a system works and the reasons…

机器学习 · 计算机科学 2023-01-13 Leonardo Lucio Custode , Giovanni Iacca

This study presents a novel computer system performance optimization and adaptive workload management scheduling algorithm based on Q-learning. In modern computing environments, characterized by increasing data volumes, task complexity, and…

机器学习 · 计算机科学 2024-11-11 Pochun Li , Yuyang Xiao , Jinghua Yan , Xuan Li , Xiaoye Wang

In this work, we present a methodology that enables an agent to make efficient use of its exploratory actions by autonomously identifying possible objectives in its environment and learning them in parallel. The identification of objectives…

人工智能 · 计算机科学 2019-01-11 Thommen George Karimpanal , Erik Wilhelm

Stochastic dynamic teams and games are rich models for decentralized systems and challenging testing grounds for multi-agent learning. Previous work that guaranteed team optimality assumed stateless dynamics, or an explicit coordination…

最优化与控制 · 数学 2024-03-28 Bora Yongacoglu , Gürdal Arslan , Serdar Yüksel

We develop a neural-network framework for multi-period risk--reward stochastic control problems with constrained two-step feedback policies that may be discontinuous in the state. We allow a broad class of objectives built on a…

计算金融 · 定量金融 2026-03-09 Chang Chen , Duy-Minh Dang

We present a new algorithm for solving linear-quadratic regulator (LQR) problems with linear equality constraints, also known as constrained LQR (CLQR) problems. Our method's sequential runtime is linear in the number of stages and…

最优化与控制 · 数学 2024-08-06 João Sousa-Pinto , Dominique Orban