English
Related papers

Related papers: Augmented Outcome-weighted Learning for Optimal Tr…

200 papers

Reinforcement learning (RL) effectively optimizes Large Language Model (LLM)-based recommenders by contrasting positive and negative items. Empirically, training with beam-search negatives consistently outperforms random negatives, yet the…

Information Retrieval · Computer Science 2026-04-27 Wentao Shi , Qifan Wang , Chen Chen , Fei Liu , Dongfang Liu , Xu Liu , Wanli Ma , Junfeng Pan , Linhong Zhu , Fuli Feng

In recent years, reinforcement learning (RL) has acquired a prominent position in health-related sequential decision-making problems, gaining traction as a valuable tool for delivering adaptive interventions (AIs). However, in part due to a…

Machine Learning · Statistics 2024-07-16 Nina Deliu , Joseph Jay Williams , Bibhas Chakraborty

This paper adds to the growing literature of reinforcement learning (RL) for healthcare by proposing a novel paradigm: augmenting any predictor with Rule-based RL Layer (RRLL) that corrects the model's physiologically impossible…

Machine Learning · Computer Science 2025-02-03 Lingwei Zhu , Zheng Chen , Yukie Nagai , Jimeng Sun

Reliable causal effect estimation from observational data requires adjustment for confounding and sufficient overlap in covariate distributions between treatment groups. However, in high-dimensional settings, lack of overlap often inflates…

Methodology · Statistics 2025-03-21 Linying Yang , Robin J. Evans

Individualized treatment regimes (ITRs) aim to improve clinical outcomes by assigning treatment based on patient-specific characteristics. However, existing methods often struggle with high-dimensional covariates, limiting accuracy,…

Machine Learning · Statistics 2026-01-13 Sungtaek Son , Eardi Lila , Kwun Chuen Gary Chan

Quantile optimal treatment regimes (OTRs) aim to assign treatments that maximize a specified quantile of patients' outcomes. Compared to treatment regimes that target the mean outcomes, quantile OTRs offer fairer regimes when a lower…

Methodology · Statistics 2026-01-07 Junwen Xia , Jingxiao Zhang , Dehan Kong

Reinforcement Learning (RL) has emerged as a central paradigm for advancing Large Language Models (LLMs), where pre-training and RL post-training share the same log-likelihood formulation. In contrast, recent RL approaches for diffusion…

Machine Learning · Computer Science 2025-09-30 Shuchen Xue , Chongjian Ge , Shilong Zhang , Yichen Li , Zhi-Ming Ma

Post-training with Reinforcement Learning (RL) has substantially improved reasoning in Large Language Models (LLMs) via test-time scaling. However, extending this paradigm to Multimodal LLMs (MLLMs) through verbose rationales yields limited…

Computation and Language · Computer Science 2026-02-16 Bangzheng Li , Jianmo Ni , Chen Qu , Ian Miao , Liu Yang , Xingyu Fu , Muhao Chen , Derek Zhiyuan Cheng

Reinforcement learning (RL) has been a promising essence in future 5G-beyond and 6G systems. Its main advantage lies in its robust model-free decision-making in complex and large-dimension wireless environments. However, most existing RL…

Robotics · Computer Science 2025-02-04 Eslam Eldeeb , Hirley Alves

Offline reinforcement learning (RL) optimizes the policy on a previously collected dataset without any interactions with the environment, yet usually suffers from the distributional shift problem. To mitigate this issue, a typical solution…

Machine Learning · Computer Science 2023-09-06 Qisen Yang , Shenzhi Wang , Qihang Zhang , Gao Huang , Shiji Song

Automated machine learning (AutoML) methods improve upon existing models by optimizing various aspects of their design. While present methods focus on hyperparameters and neural network topologies, other aspects of neural network design can…

Machine Learning · Computer Science 2023-04-10 Garrett Bingham

Optimal Order Execution is a well-established problem in finance that pertains to the flawless execution of a trade (buy or sell) for a given volume within a specified time frame. This problem revolves around optimizing returns while…

Computational Finance · Quantitative Finance 2026-01-13 Khabbab Zakaria , Jayapaulraj Jerinsh , Andreas Maier , Patrick Krauss , Stefano Pasquali , Dhagash Mehta

Class imbalance remains a fundamental challenge in machine learning, where standard classifiers exhibit severe performance degradation in minority classes. Although existing approaches address imbalance through resampling or cost-sensitive…

Machine Learning · Computer Science 2026-02-10 Zahir Alsulaimawi

Large-scale constrained optimization is pivotal in modern scientific, engineering, and industrial computation, often involving complex systems with numerous variables and constraints. This paper provides a unified and comprehensive…

Optimization and Control · Mathematics 2025-10-21 Kangkang Deng , Rui Wang , Zhenyuan Zhu , Junyu Zhang , Zaiwen Wen

Medical image analysis relies on accurate segmentation, and benefits from controllable synthesis (of new training images). Yet both tasks of the cyclical pipeline face spatial imbalance: lesions occupy small regions against vast…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Anugunj Naman , Ayushman Singh , Gaibo Zhang , Yaguang Zhang

A dynamic treatment regime (DTR) is an approach to delivering precision medicine that uses patient characteristics to guide treatment decisions for optimal health outcomes. Numerous methods have been proposed for DTR estimation, including…

Methodology · Statistics 2025-02-03 Adel Ahmadi Nadi , Michael Wallace

Multimodal learning often encounters the under-optimized problem and may perform worse than unimodal learning. Existing approaches attribute this issue to imbalanced learning across modalities and tend to address it through gradient…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Shicai Wei , Chunbo Luo , Yang Luo

Covariate balance is crucial in obtaining unbiased estimates of treatment effects in observational studies. Methods based on inverse probability weights have been widely used to estimate treatment effects with observational data. Machine…

Methodology · Statistics 2021-04-08 Michele Santacatterina

In this work, we address the problem of determining reliable policies in reinforcement learning (RL), with a focus on optimization under uncertainty and the need for performance guarantees. While classical RL algorithms aim at maximizing…

Machine Learning · Computer Science 2025-10-22 Nadir Farhi

Online reinforcement learning (RL) methods are often data-inefficient or unreliable, making them difficult to train on real robotic hardware, especially quadruped robots. Learning robotic tasks from pre-collected data is a promising…

Robotics · Computer Science 2024-10-28 Hongyin Zhang , Shuyu Yang , Donglin Wang