English
Related papers

Related papers: Quantum Computing Provides Exponential Regret Impr…

200 papers

Stochastic contextual bandits are fundamental for sequential decision-making but pose significant challenges for existing neural network-based algorithms, particularly when scaling to quantum neural networks (QNNs) due to issues such as…

Machine Learning · Computer Science 2026-01-07 Yuqi Huang , Vincent Y. F Tan , Sharu Theresa Jose

Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.…

Quantum Physics · Physics 2025-07-25 Gilberto Cunha , Alexandra Ramôa , André Sequeira , Michael de Oliveira , Luís Barbosa

The theory of reinforcement learning currently suffers from a mismatch between its empirical performance and the theoretical characterization of its performance, with consequences for, e.g., the understanding of sample efficiency, safety,…

Machine Learning · Computer Science 2022-02-14 Feicheng Wang , Lucas Janson

We study online learning in episodic finite-horizon Markov decision processes (MDPs) with convex objective functions, known as the concave utility reinforcement learning (CURL) problem. This setting generalizes RL from linear to convex…

Machine Learning · Computer Science 2025-05-13 Bianca Marin Moreno , Khaled Eldowa , Pierre Gaillard , Margaux Brégère , Nadia Oudjane

In this work, we address the open problem of finding low-complexity near-optimal multi-armed bandit algorithms for sequential decision making problems. Existing bandit algorithms are either sub-optimal and computationally simple (e.g.,…

Machine Learning · Computer Science 2018-04-18 Fang Liu , Sinong Wang , Swapna Buccapatnam , Ness Shroff

In the past decade, the field of quantum machine learning has drawn significant attention due to the prospect of bringing genuine computational advantages to now widespread algorithmic methods. However, not all domains of machine learning…

Many applications require optimizing an unknown, noisy function that is expensive to evaluate. We formalize this task as a multi-armed bandit problem, where the payoff function is either sampled from a Gaussian process (GP) or has low RKHS…

Machine Learning · Computer Science 2015-03-13 Niranjan Srinivas , Andreas Krause , Sham M. Kakade , Matthias Seeger

In this paper, we consider the multi-armed bandit problem with high-dimensional features. First, we prove a minimax lower bound, $\mathcal{O}\big((\log d)^{\frac{\alpha+1}{2}}T^{\frac{1-\alpha}{2}}+\log T\big)$, for the cumulative regret,…

Machine Learning · Computer Science 2021-09-27 Ke Li , Yun Yang , Naveen N. Narisetty

We present the first regret bound for classical online Q-learning in infinite-horizon discounted Markov decision processes (MDPs), without relying on optimism or bonus terms. We first analyze Boltzmann Q-learning with decaying temperature…

Machine Learning · Computer Science 2026-05-18 Rahul Singh , Siddharth Chandak , Eric Moulines , Vivek S. Borkar , Nicholas Bambos

Autoregressive processes naturally arise in a large variety of real-world scenarios, including stock markets, sales forecasting, weather prediction, advertising, and pricing. When facing a sequential decision-making problem in such a…

The Adversarial Markov Decision Process (AMDP) is a learning framework that deals with unknown and varying tasks in decision-making applications like robotics and recommendation systems. A major limitation of the AMDP formalism, however, is…

Machine Learning · Statistics 2024-05-06 Sang Bin Moon , Abolfazl Hashemi

We consider the regret minimization problem in reinforcement learning (RL) in the episodic setting. In many real-world RL environments, the state and action spaces are continuous or very large. Existing approaches establish regret…

Machine Learning · Computer Science 2022-06-29 Sayak Ray Chowdhury , Rafael Oliveira

Conservative Contextual Bandits (CCBs) address safety in sequential decision making by requiring that an agent's policy, along with minimizing regret, also satisfies a safety constraint: the performance is not worse than a baseline policy…

Machine Learning · Computer Science 2024-12-10 Rohan Deb , Mohammad Ghavamzadeh , Arindam Banerjee

Machine learning is widely believed to be one of the most promising practical applications of quantum computing. Existing quantum machine learning schemes typically employ a quantum-classical hybrid approach that relies crucially on…

Quantum Physics · Physics 2025-02-11 Qi Ye , Shuangyue Geng , Zizhao Han , Weikang Li , L. -M. Duan , Dong-Ling Deng

Bayesian Optimization (BO) is widely used for optimising black-box functions but requires us to specify the length scale hyperparameter, which defines the smoothness of the functions the optimizer will consider. Most current BO algorithms…

Machine Learning · Statistics 2024-11-26 Juliusz Ziomek , Masaki Adachi , Michael A. Osborne

The promise of fault-tolerant quantum computing is challenged by environmental drift that relentlessly degrades the quality of quantum operations. The contemporary solution, halting the entire quantum computation for recalibration, is…

Quantum Physics · Physics 2026-03-10 Volodymyr Sivak , Alexis Morvan , Michael Broughton , Rodrigo G. Cortiñas , Johannes Bausch , Andrew W. Senior , Matthew Neeley , Alec Eickbusch , Noah Shutty , Laleh Aghababaie Beni , James S. Spencer , Francisco J. H Heras , Thomas Edlich , Dmitry Abanin , Amira Abbas , Rajeev Acharya , Georg Aigeldinger , Ross Alcaraz , Sayra Alcaraz , Trond I. Andersen , Markus Ansmann , Frank Arute , Kunal Arya , Walt Askew , Nikita Astrakhantsev , Juan Atalaya , Brian Ballard , Joseph C. Bardin , Hector Bates , Andreas Bengtsson , Majid Bigdeli Karimi , Alexander Bilmes , Simon Bilodeau , Felix Borjans , Alexandre Bourassa , Jenna Bovaird , Dylan Bowers , Leon Brill , Peter Brooks , David A. Browne , Brett Buchea , Bob B. Buckley , Tim Burger , Brian Burkett , Nicholas Bushnell , Jamal Busnaina , Anthony Cabrera , Juan Campero , Hung-Shen Chang , Silas Chen , Ben Chiaro , Liang-Ying Chih , Agnetta Y. Cleland , Bryan Cochrane , Matt Cockrell , Josh Cogan , Roberto Collins , Paul Conner , Harold Cook , William Courtney , Alexander L. Crook , Ben Curtin , Martin Damyanov , Sayan Das , Dripto M. Debroy , Sean Demura , Paul Donohoe , Ilya Drozdov , Andrew Dunsworth , Valerie Ehimhen , Aviv Moshe Elbag , Lior Ella , Mahmoud Elzouka , David Enriquez , Catherine Erickson , Vinicius S. Ferreira , Marcos Flores , Leslie Flores Burgos , Ebrahim Forati , Jeremiah Ford , Austin G. Fowler , Brooks Foxen , Masaya Fukami , Alan Wing Lun Fung , Lenny Fuste , Suhas Ganjam , Gonzalo Garcia , Christopher Garrick , Robert Gasca , Helge Gehring , Robert Geiger , Élie Genois , William Giang , Dar Gilboa , James E. Goeders , Edward C. Gonzales , Raja Gosula , Stijn J. de Graaf , Alejandro Grajales Dau , Dietrich Graumann , Joel Grebel , Alex Greene , Jonathan A. Gross , Jose Guerrero , Loïck Le Guevel , Tan Ha , Steve Habegger , Tanner Hadick , Ali Hadjikhani , Michael C. Hamilton , Matthew P. Harrigan , Sean D. Harrington , Jeanne Hartshorn , Stephen Heslin , Paula Heu , Oscar Higgott , Reno Hiltermann , Hsin-Yuan Huang , Mike Hucka , Christopher Hudspeth , Ashley Huff , William J. Huggins , Evan Jeffrey , Shaun Jevons , Zhang Jiang , Xiaoxuan Jin , Chaitali Joshi , Pavol Juhas , Andreas Kabel , Dvir Kafri , Hui Kang , Kiseo Kang , Amir H. Karamlou , Ryan Kaufman , Kostyantyn Kechedzhi , Tanuj Khattar , Mostafa Khezri , Seon Kim , Can M. Knaut , Bryce Kobrin , Fedor Kostritsa , John Mark Kreikebaum , Ryuho Kudo , Ben Kueffler , Arun Kumar , Vladislav D. Kurilovich , Vitali Kutsko , Nathan Lacroix , David Landhuis , Tiano Lange-Dei , Brandon W. Langley , Pavel Laptev , Kim-Ming Lau , Justin Ledford , Joy Lee , Kenny Lee , Brian J. Lester , Wendy Leung , Lily Li , Wing Yan Li , Ming Li , Alexander T. Lill , William P. Livingston , Matthew T. Lloyd , Aditya Locharla , Laura De Lorenzo , Daniel Lundahl , Aaron Lunt , Sid Madhuk , Aniket Maiti , Ashley Maloney , Salvatore Mandrà , Leigh S. Martin , Orion Martin , Eric Mascot , Paul Masih Das , Dmitri Maslov , Melvin Mathews , Cameron Maxfield , Jarrod R. McClean , Matt McEwen , Seneca Meeks , Kevin C. Miao , Zlatko K. Minev , Reza Molavi , Sebastian Molina , Shirin Montazeri , Charles Neill , Michael Newman , Anthony Nguyen , Murray Nguyen , Chia-Hung Ni , Murphy Yuezhen Niu , Logan Oas , Raymond Orosco , Kristoffer Ottosson , Alice Pagano , Agustin Di Paolo , Sherman Peek , David Peterson , Alex Pizzuto , Elias Portoles , Rebecca Potter , Orion Pritchard , Michael Qian , Chris Quintana , Arpit Ranadive , Matthew J. Reagor , Rachel Resnick , David M. Rhodes , Daniel Riley , Gabrielle Roberts , Roberto Rodriguez , Emma Ropes , Lucia B. De Rose , Eliott Rosenberg , Emma Rosenfeld , Dario Rosenstock , Elizabeth Rossi , Pedram Roushan , David A. Rower , Robert Salazar , Kannan Sankaragomathi , Murat Can Sarihan , Kevin J. Satzinger , Max Schaefer , Sebastian Schroeder , Henry F. Schurkus , Aria Shahingohar , Michael J. Shearn , Aaron Shorter , Vladimir Shvarts , Spencer Small , W. Clarke Smith , David A. Sobel , Barrett Spells , Sofia Springer , George Sterling , Jordan Suchard , Aaron Szasz , Alexander Sztein , Madeline Taylor , Jothi Priyanka Thiruraman , Douglas Thor , Dogan Timucin , Eifu Tomita , Alfredo Torres , M. Mert Torunbalci , Hao Tran , Abeer Vaishnav , Justin Vargas , Sergey Vdovichev , Guifre Vidal , Catherine Vollgraff Heidweiller , Meghan Voorhees , Steven Waltman , Jonathan Waltz , Shannon X. Wang , Brayden Ware , James D. Watson , Yonghua Wei , Travis Weidel , Theodore White , Kristi Wong , Bryan W. K. Woo , Christopher J. Wood , Maddy Woodson , Cheng Xing , Z. Jamie Yao , Ping Yeh , Bicheng Ying , Juhwan Yoo , Noureldin Yosri , Elliot Young , Grayson Young , Adam Zalcman , Ran Zhang , Yaxing Zhang , Ningfeng Zhu , Nicholas Zobrist , Zhenjie Zou , Ryan Babbush , Dave Bacon , Sergio Boixo , Yu Chen , Zijun Chen , Michel Devoret , Monica Hansen , Jeremy Hilton , Cody Jones , Julian Kelly , Alexander N. Korotkov , Erik Lucero , Anthony Megrant , Hartmut Neven , William D. Oliver , Ganesh Ramachandran , Vadim Smelyanskiy , Paul V. Klimov

We take initial steps in studying PAC-MDP algorithms with limited adaptivity, that is, algorithms that change its exploration policy as infrequently as possible during regret minimization. This is motivated by the difficulty of running…

Machine Learning · Computer Science 2020-02-11 Yu Bai , Tengyang Xie , Nan Jiang , Yu-Xiang Wang

This work studies the problem of learning episodic Markov Decision Processes with known transition and bandit feedback. We develop the first algorithm with a ``best-of-both-worlds'' guarantee: it achieves $\mathcal{O}(log T)$ regret when…

Machine Learning · Computer Science 2020-11-03 Tiancheng Jin , Haipeng Luo

We study the online resource allocation problem in which at each round, a budget $B$ must be allocated across $K$ arms under censored feedback. An arm yields a reward if and only if two conditions are satisfied: (i) the arm is activated…

Machine Learning · Computer Science 2026-02-09 Giovanni Montanari , Côme Fiegel , Corentin Pla , Aadirupa Saha , Vianney Perchet

We study the regret guarantee for risk-sensitive reinforcement learning (RSRL) via distributional reinforcement learning (DRL) methods. In particular, we consider finite episodic Markov decision processes whose objective is the entropic…

Machine Learning · Computer Science 2024-01-26 Hao Liang , Zhi-Quan Luo
‹ Prev 1 8 9 10 Next ›