中文
相关论文

相关论文: Time After Time: Deep-Q Effect Estimation for Inte…

200 篇论文

The Deep Q-Network proposed by Mnih et al. [2015] has become a benchmark and building point for much deep reinforcement learning research. However, replicating results for complex systems is often challenging since original scientific…

机器学习 · 计算机科学 2017-11-22 Melrose Roderick , James MacGlashan , Stefanie Tellex

Quality-Diversity (QD) algorithms have emerged as a powerful optimization paradigm with the aim of generating a set of high-quality and diverse solutions. To achieve such a challenging goal, QD algorithms require maintaining a large archive…

机器学习 · 计算机科学 2024-06-07 Ren-Jian Wang , Ke Xue , Cong Guan , Chao Qian

In this paper, a novel Deep Q-Network (DQN) based scheduling method to optimize delay time and fairness among entanglement requests in quantum repeater networks is proposed. The scheduling of requests determines which pairs of end nodes…

量子物理 · 物理学 2025-05-20 Gongyu Ni , Lester Ho , Holger Claussen

Scarcity of health care resources could result in the unavoidable consequence of rationing. For example, ventilators are often limited in supply, especially during public health emergencies or in resource-constrained health care settings,…

机器学习 · 计算机科学 2024-08-23 Yikuan Li , Chengsheng Mao , Kaixuan Huang , Hanyin Wang , Zheng Yu , Mengdi Wang , Yuan Luo

Engineering system design, viewed as a decision-making process, faces challenges due to complexity and uncertainty. In this paper, we present a framework proposing the use of the Deep Q-learning algorithm to optimize the design of…

机器学习 · 计算机科学 2024-01-01 Ramin Giahi , Cameron A. MacKenzie , Reyhaneh Bijari

We present an enhanced entangled quantum clock protocol that incorporates a quantum phase estimation algorithm to directly estimate proper-time differences as an unknown phase. By employing highly entangled multi-clock states, the…

量子物理 · 物理学 2026-04-09 Won-Young Hwang

We present a comprehensive analysis of deep learning approaches for Electronic Health Record (EHR) time-series imputation, examining how architectural and framework biases combine to influence model performance. Our investigation reveals…

机器学习 · 计算机科学 2025-02-05 Linglong Qian , Tao Wang , Jun Wang , Hugh Logan Ellis , Robin Mitra , Richard Dobson , Zina Ibrahim

We present DPIQN, a deep policy inference Q-network that targets multi-agent systems composed of controllable agents, collaborators, and opponents that interact with each other. We focus on one challenging issue in such systems---modeling…

人工智能 · 计算机科学 2018-04-10 Zhang-Wei Hong , Shih-Yang Su , Tzu-Yun Shann , Yi-Hsiang Chang , Chun-Yi Lee

In statistical dialogue management, the dialogue manager learns a policy that maps a belief state to an action for the system to perform. Efficient exploration is key to successful policy optimisation. Current deep reinforcement learning…

机器学习 · 统计学 2017-12-04 Christopher Tegho , Paweł Budzianowski , Milica Gašić

Despite advancements in state-of-the-art models and information retrieval techniques, current systems still struggle to handle temporal information and to correctly answer detailed questions about past events. In this paper, we investigate…

计算与语言 · 计算机科学 2025-03-10 Mehmet Kardan , Bhawna Piryani , Adam Jatowt

One of the major challenges in Deep Reinforcement Learning for control is the need for extensive training to learn the policy. Motivated by this, we present the design of the Control-Tutored Deep Q-Networks (CT-DQN) algorithm, a Deep…

机器学习 · 计算机科学 2022-12-05 Francesco De Lellis , Marco Coraggio , Giovanni Russo , Mirco Musolesi , Mario di Bernardo

Early recognition of abnormal rhythms in ECG signals is crucial for monitoring and diagnosing patients' cardiac conditions, increasing the success rate of the treatment. Classifying abnormal rhythms into exact categories is very challenging…

机器学习 · 计算机科学 2019-12-18 Jing Zhang , Jing Tian , Yang Cao , Yuxiang Yang , Xiaobin Xu

Emphatic Temporal Difference (ETD) learning has recently been proposed as a convergent off-policy learning method. ETD was proposed mainly to address convergence issues of conventional Temporal Difference (TD) learning under off-policy…

人工智能 · 计算机科学 2019-03-04 Xiang Gu , Sina Ghiassian , Richard S. Sutton

An operationally well-defined delayed-choice quantum-eraser experiment is proposed, realizing a genuine delayed choice within presently available quantum-optical technology. A multimode quantum memory supplies a controlled and verifiable…

量子物理 · 物理学 2025-12-23 Taku Ohwada

Deep Q-Learning (DQL), a family of temporal difference algorithms for control, employs three techniques collectively known as the `deadly triad' in reinforcement learning: bootstrapping, off-policy learning, and function approximation.…

机器学习 · 计算机科学 2019-03-22 Joshua Achiam , Ethan Knight , Pieter Abbeel

In reinforcement learning, it is often difficult to automate high-dimensional, rapid decision-making in dynamic environments, especially when domains require real-time online interaction and adaptive strategies such as web-based games. This…

机器学习 · 计算机科学 2024-05-30 Prabhath Reddy Gujavarthy

``Distribution shift'' is the main obstacle to the success of offline reinforcement learning. A learning policy may take actions beyond the behavior policy's knowledge, referred to as Out-of-Distribution (OOD) actions. The Q-values for…

机器学习 · 计算机科学 2025-01-14 Jing Zhang , Linjiajie Fang , Kexin Shi , Wenjia Wang , Bing-Yi Jing

Optimal execution is an important problem faced by any trader. Most solutions are based on the assumption of constant market impact, while liquidity is known to be dynamic. Moreover, models with time-varying liquidity typically assume that…

交易与市场微观结构 · 定量金融 2024-02-21 Andrea Macrì , Fabrizio Lillo

In this study, we present a novel Survival Analysis algorithm designed to efficiently handle large-scale longitudinal data. Our approach draws inspiration from Reinforcement Learning principles, particularly the Deep Q-Network paradigm,…

机器学习 · 计算机科学 2024-10-10 Mariana Vargas Vieyra , Pascal Frossard

Drawing upon recent advances in language model alignment, we formulate offline Reinforcement Learning as a two-stage optimization problem: First pretraining expressive generative policies on reward-free behavior datasets, then fine-tuning…

机器学习 · 计算机科学 2024-10-31 Huayu Chen , Kaiwen Zheng , Hang Su , Jun Zhu