English
Related papers

Related papers: Tilted Quantile Gradient Updates for Quantile-Cons…

200 papers

The growing literature of Federated Learning (FL) has recently inspired Federated Reinforcement Learning (FRL) to encourage multiple agents to federatively build a better decision-making policy without sharing raw trajectories. Despite its…

Machine Learning · Computer Science 2022-11-04 Flint Xiaofeng Fan , Yining Ma , Zhongxiang Dai , Wei Jing , Cheston Tan , Bryan Kian Hsiang Low

The promise of fault-tolerant quantum computing is challenged by environmental drift that relentlessly degrades the quality of quantum operations. The contemporary solution, halting the entire quantum computation for recalibration, is…

Quantum Physics · Physics 2026-03-10 Volodymyr Sivak , Alexis Morvan , Michael Broughton , Rodrigo G. Cortiñas , Johannes Bausch , Andrew W. Senior , Matthew Neeley , Alec Eickbusch , Noah Shutty , Laleh Aghababaie Beni , James S. Spencer , Francisco J. H Heras , Thomas Edlich , Dmitry Abanin , Amira Abbas , Rajeev Acharya , Georg Aigeldinger , Ross Alcaraz , Sayra Alcaraz , Trond I. Andersen , Markus Ansmann , Frank Arute , Kunal Arya , Walt Askew , Nikita Astrakhantsev , Juan Atalaya , Brian Ballard , Joseph C. Bardin , Hector Bates , Andreas Bengtsson , Majid Bigdeli Karimi , Alexander Bilmes , Simon Bilodeau , Felix Borjans , Alexandre Bourassa , Jenna Bovaird , Dylan Bowers , Leon Brill , Peter Brooks , David A. Browne , Brett Buchea , Bob B. Buckley , Tim Burger , Brian Burkett , Nicholas Bushnell , Jamal Busnaina , Anthony Cabrera , Juan Campero , Hung-Shen Chang , Silas Chen , Ben Chiaro , Liang-Ying Chih , Agnetta Y. Cleland , Bryan Cochrane , Matt Cockrell , Josh Cogan , Roberto Collins , Paul Conner , Harold Cook , William Courtney , Alexander L. Crook , Ben Curtin , Martin Damyanov , Sayan Das , Dripto M. Debroy , Sean Demura , Paul Donohoe , Ilya Drozdov , Andrew Dunsworth , Valerie Ehimhen , Aviv Moshe Elbag , Lior Ella , Mahmoud Elzouka , David Enriquez , Catherine Erickson , Vinicius S. Ferreira , Marcos Flores , Leslie Flores Burgos , Ebrahim Forati , Jeremiah Ford , Austin G. Fowler , Brooks Foxen , Masaya Fukami , Alan Wing Lun Fung , Lenny Fuste , Suhas Ganjam , Gonzalo Garcia , Christopher Garrick , Robert Gasca , Helge Gehring , Robert Geiger , Élie Genois , William Giang , Dar Gilboa , James E. Goeders , Edward C. Gonzales , Raja Gosula , Stijn J. de Graaf , Alejandro Grajales Dau , Dietrich Graumann , Joel Grebel , Alex Greene , Jonathan A. Gross , Jose Guerrero , Loïck Le Guevel , Tan Ha , Steve Habegger , Tanner Hadick , Ali Hadjikhani , Michael C. Hamilton , Matthew P. Harrigan , Sean D. Harrington , Jeanne Hartshorn , Stephen Heslin , Paula Heu , Oscar Higgott , Reno Hiltermann , Hsin-Yuan Huang , Mike Hucka , Christopher Hudspeth , Ashley Huff , William J. Huggins , Evan Jeffrey , Shaun Jevons , Zhang Jiang , Xiaoxuan Jin , Chaitali Joshi , Pavol Juhas , Andreas Kabel , Dvir Kafri , Hui Kang , Kiseo Kang , Amir H. Karamlou , Ryan Kaufman , Kostyantyn Kechedzhi , Tanuj Khattar , Mostafa Khezri , Seon Kim , Can M. Knaut , Bryce Kobrin , Fedor Kostritsa , John Mark Kreikebaum , Ryuho Kudo , Ben Kueffler , Arun Kumar , Vladislav D. Kurilovich , Vitali Kutsko , Nathan Lacroix , David Landhuis , Tiano Lange-Dei , Brandon W. Langley , Pavel Laptev , Kim-Ming Lau , Justin Ledford , Joy Lee , Kenny Lee , Brian J. Lester , Wendy Leung , Lily Li , Wing Yan Li , Ming Li , Alexander T. Lill , William P. Livingston , Matthew T. Lloyd , Aditya Locharla , Laura De Lorenzo , Daniel Lundahl , Aaron Lunt , Sid Madhuk , Aniket Maiti , Ashley Maloney , Salvatore Mandrà , Leigh S. Martin , Orion Martin , Eric Mascot , Paul Masih Das , Dmitri Maslov , Melvin Mathews , Cameron Maxfield , Jarrod R. McClean , Matt McEwen , Seneca Meeks , Kevin C. Miao , Zlatko K. Minev , Reza Molavi , Sebastian Molina , Shirin Montazeri , Charles Neill , Michael Newman , Anthony Nguyen , Murray Nguyen , Chia-Hung Ni , Murphy Yuezhen Niu , Logan Oas , Raymond Orosco , Kristoffer Ottosson , Alice Pagano , Agustin Di Paolo , Sherman Peek , David Peterson , Alex Pizzuto , Elias Portoles , Rebecca Potter , Orion Pritchard , Michael Qian , Chris Quintana , Arpit Ranadive , Matthew J. Reagor , Rachel Resnick , David M. Rhodes , Daniel Riley , Gabrielle Roberts , Roberto Rodriguez , Emma Ropes , Lucia B. De Rose , Eliott Rosenberg , Emma Rosenfeld , Dario Rosenstock , Elizabeth Rossi , Pedram Roushan , David A. Rower , Robert Salazar , Kannan Sankaragomathi , Murat Can Sarihan , Kevin J. Satzinger , Max Schaefer , Sebastian Schroeder , Henry F. Schurkus , Aria Shahingohar , Michael J. Shearn , Aaron Shorter , Vladimir Shvarts , Spencer Small , W. Clarke Smith , David A. Sobel , Barrett Spells , Sofia Springer , George Sterling , Jordan Suchard , Aaron Szasz , Alexander Sztein , Madeline Taylor , Jothi Priyanka Thiruraman , Douglas Thor , Dogan Timucin , Eifu Tomita , Alfredo Torres , M. Mert Torunbalci , Hao Tran , Abeer Vaishnav , Justin Vargas , Sergey Vdovichev , Guifre Vidal , Catherine Vollgraff Heidweiller , Meghan Voorhees , Steven Waltman , Jonathan Waltz , Shannon X. Wang , Brayden Ware , James D. Watson , Yonghua Wei , Travis Weidel , Theodore White , Kristi Wong , Bryan W. K. Woo , Christopher J. Wood , Maddy Woodson , Cheng Xing , Z. Jamie Yao , Ping Yeh , Bicheng Ying , Juhwan Yoo , Noureldin Yosri , Elliot Young , Grayson Young , Adam Zalcman , Ran Zhang , Yaxing Zhang , Ningfeng Zhu , Nicholas Zobrist , Zhenjie Zou , Ryan Babbush , Dave Bacon , Sergio Boixo , Yu Chen , Zijun Chen , Michel Devoret , Monica Hansen , Jeremy Hilton , Cody Jones , Julian Kelly , Alexander N. Korotkov , Erik Lucero , Anthony Megrant , Hartmut Neven , William D. Oliver , Ganesh Ramachandran , Vadim Smelyanskiy , Paul V. Klimov

Risk-sensitive reinforcement learning (RL) aims to optimize policies that balance the expected reward and risk. In this paper, we present a novel risk-sensitive RL framework that employs an Iterated Conditional Value-at-Risk (CVaR)…

Machine Learning · Computer Science 2023-12-05 Yu Chen , Yihan Du , Pihe Hu , Siwei Wang , Desheng Wu , Longbo Huang

Modern approaches to autonomous driving rely heavily on learned components trained with large amounts of human driving data via imitation learning. However, these methods require large amounts of expensive data collection and even then face…

Reinforcement learning (RL) agents with pre-specified reward functions cannot provide guaranteed safety across variety of circumstances that an uncertain system might encounter. To guarantee performance while assuring satisfaction of safety…

Artificial Intelligence · Computer Science 2021-04-20 Aquib Mustafa , Majid Mazouchi , Subramanya Nageshrao , Hamidreza Modares

Real-world reinforcement learning (RL) problems often demand that agents behave safely by obeying a set of designed constraints. We address the challenge of safe RL by coupling a safety guide based on model predictive control (MPC) with a…

Machine Learning · Computer Science 2022-03-30 Samuel Pfrommer , Tanmay Gautam , Alec Zhou , Somayeh Sojoudi

Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose Sampling-Based Safe Reinforcement Learning (SBSRL), a model-based RL algorithm that…

Machine Learning · Computer Science 2026-05-20 Luca Vignola , Bruce D. Lee , Manish Prajapat , Manuel Wendl , Melanie Zeilinger , Andreas Krause , Yarden As

Safe reinforcement learning (RL) is a standard paradigm for safety-critical decision making. However, real-world safety constraints can be complex, subjective, and even hard to explicitly specify. Existing works on constraint inference rely…

Machine Learning · Computer Science 2026-05-25 Chenglin Li , Grant Ruan , Hua Geng

Given a set of trajectories demonstrating the execution of a task safely in a constrained MDP with observable rewards but with unknown constraints and non-observable costs, we aim to find a policy that maximizes the likelihood of…

Machine Learning · Computer Science 2026-03-02 George Papadopoulos , George A. Vouros

In this paper, we propose the Quantile Option Architecture (QUOTA) for exploration based on recent advances in distributional reinforcement learning (RL). In QUOTA, decision making is based on quantiles of a value distribution, not only the…

Machine Learning · Computer Science 2018-11-09 Shangtong Zhang , Borislav Mavrin , Linglong Kong , Bo Liu , Hengshuai Yao

Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing methods often rely on soft expected-cost objectives or iterative generative inference, which can be…

Machine Learning · Computer Science 2026-03-17 Mumuksh Tayal , Manan Tayal , Ravi Prakash

Reinforcement learning (RL) enables agents to learn optimal behaviors through interaction with their environment and has been increasingly deployed in safety-critical applications, including autonomous driving. Despite its promise, RL is…

Continuous-time reinforcement learning (CTRL) provides a natural framework for sequential decision-making in dynamic environments where interactions evolve continuously over time. While CTRL has shown growing empirical success, its ability…

Machine Learning · Computer Science 2025-12-04 Runze Zhao , Yue Yu , Ruhan Wang , Chunfeng Huang , Dongruo Zhou

Several works have addressed the problem of incorporating constraints in the reinforcement learning (RL) framework, however majority of them can only guarantee the satisfaction of soft constraints. In this work, we address the problem of…

Machine Learning · Computer Science 2020-06-16 Kwangyeon Kim , Akshita Gupta , Hong-Cheol Choi , Inseok Hwang

Ensuring that reinforcement learning (RL) controllers satisfy safety and reliability constraints in real-world settings remains challenging: state-avoidance and constrained Markov decision processes often fail to capture trajectory-level…

Machine Learning · Computer Science 2026-04-06 Alper Kamil Bozkurt , Calin Belta , Ming C. Lin

We introduce Coarse Q-learning (CQL), a reinforcement-learning model for bandit problems with stochastically varying menus. Alternatives are exogenously partitioned into similarity classes, and feedback from sampled alternatives is pooled…

Theoretical Economics · Economics 2026-05-13 Philippe Jehiel , Aviman Satpathy

This paper proposes a novel formulation for reinforcement learning (RL) with large language models, explaining why and under what conditions the true sequence-level reward can be optimized via a surrogate token-level objective in policy…

Machine Learning · Computer Science 2025-12-04 Chujie Zheng , Kai Dang , Bowen Yu , Mingze Li , Huiqiang Jiang , Junrong Lin , Yuqiong Liu , Hao Lin , Chencan Wu , Feng Hu , An Yang , Jingren Zhou , Junyang Lin

The use of Reinforcement Learning (RL) is still restricted to simulation or to enhance human-operated systems through recommendations. Real-world environments (e.g. industrial robots or power grids) are generally designed with safety…

Machine Learning · Computer Science 2020-08-14 Mathieu Seurin , Philippe Preux , Olivier Pietquin

Quantum reinforcement learning (QRL) is a promising paradigm for near-term quantum devices. While existing QRL methods have shown success in discrete action spaces, extending these techniques to continuous domains is challenging due to the…

Quantum Physics · Physics 2025-03-19 Shaojun Wu , Shan Jin , Dingding Wen , Donghong Han , Xiaoting Wang

In many RL applications, ensuring an agent's actions adhere to constraints is crucial for safety. Most previous methods in Action-Constrained Reinforcement Learning (ACRL) employ a projection layer after the policy network to correct the…

Machine Learning · Computer Science 2025-02-18 Janaka Chathuranga Brahmanage , Jiajing Ling , Akshat Kumar