English
Related papers

Related papers: eBandit: Kernel-Driven Reinforcement Learning for …

200 papers

Sparse Mixture of Experts (SMoE) has become a preferred architecture for scaling Transformer capacity without increasing computational cost, as it activates only a small subset of experts for each input. However, deploying such an approach…

Machine Learning · Computer Science 2026-01-26 Ziyi Han , Xutong Liu , Ruiting Zhou , Xiangxiang Dai , John C. S. Lui

The rapid development of multimedia and communication technology has resulted in an urgent need for high-quality video streaming. However, robust video streaming under fluctuating network conditions and heterogeneous client computing…

Systems and Control · Electrical Eng. & Systems 2022-01-21 Junyan Yang , Yang Jiang , Shuoyao Wang

We propose online algorithms for sequential learning in the contextual multi-armed bandit setting. Our approach is to partition the context space and then optimally combine all of the possible mappings between the partition regions and the…

Machine Learning · Computer Science 2017-12-11 Mohammadreza Mohaghegh Neyshabouri , Kaan Gokcesu , Huseyin Ozkan , Suleyman S. Kozat

In Reinforcement Learning (RL), multi-armed Bandit (MAB) problems have found applications across diverse domains such as recommender systems, healthcare, and finance. Traditional MAB algorithms typically assume stationary reward…

Artificial Intelligence · Computer Science 2024-10-10 Gustavo de Freitas Fonseca , Lucas Coelho e Silva , Paulo André Lima de Castro

Recent advances in wireless radio frequency (RF) energy harvesting allows sensor nodes to increase their lifespan by remotely charging their batteries. The amount of energy harvested by the nodes varies depending on their ambient…

Machine Learning · Computer Science 2020-06-17 Debamita Ghosh , Arun Verma , Manjesh K. Hanawal

Nowadays Dynamic Adaptive Streaming over HTTP (DASH) is the most prevalent solution on the Internet for multimedia streaming and responsible for the majority of global traffic. DASH uses adaptive bit rate (ABR) algorithms, which select the…

Networking and Internet Architecture · Computer Science 2018-09-28 Pablo Gil Pereira , Andreas Schmidt , Thorsten Herfet

Online experimentation with interference is a common challenge in modern applications such as e-commerce and adaptive clinical trials in medicine. For example, in online marketplaces, the revenue of a good depends on discounts applied to…

Machine Learning · Computer Science 2024-05-30 Abhineet Agarwal , Anish Agarwal , Lorenzo Masoero , Justin Whitehouse

We formulate a new problem at the intersectionof semi-supervised learning and contextual bandits,motivated by several applications including clini-cal trials and ad recommendations. We demonstratehow Graph Convolutional Network (GCN), a…

Machine Learning · Computer Science 2020-10-26 Sohini Upadhyay , Mikhail Yurochkin , Mayank Agarwal , Yasaman Khazaeni , DjallelBouneffouf

Immersive media streaming, especially virtual reality (VR)/360-degree video streaming which is very bandwidth demanding, has become more and more popular due to the rapid growth of the multimedia and networking deployments. To better…

Multimedia · Computer Science 2018-03-22 Wei Huang , Lianghui Ding , Hung-Yu Wei , Jenq-Neng Hwang , Yiling Xu , Wenjun Zhang

The network performance is usually assessed by drive tests, where teams of people with specially equipped vehicles physically drive out to test various locations throughout a radio network. However, intelligent and autonomous…

Networking and Internet Architecture · Computer Science 2023-05-01 Hakan Gokcesu , Ozgur Ercetin , Gokhan Kalem , Salih Ergut

Contextual bandits (CB) are online sequential decision-making problems under partial feedback that underpin many adaptive services. There is a growing demand to deploy CB agents directly on-device, under strict constraints on memory,…

Machine Learning · Computer Science 2026-05-14 Marco Angioli , Kevin Johansson , Antonello Rosato , Amy Loutfi , Denis Kleyko

In this paper, we consider a risk-averse multi-armed bandit (MAB) problem where the goal is to learn a policy that minimizes the risk of low expected return, as opposed to maximizing the expected return itself, which is the objective in the…

Machine Learning · Computer Science 2022-09-12 Yi Shen , Jessilyn Dunn , Michael M. Zavlanos

We describe the results of a randomized controlled trial of video-streaming algorithms for bitrate selection and network prediction. Over the last eight months, we have streamed 14.2 years of video to 56,000 users across the Internet.…

Networking and Internet Architecture · Computer Science 2020-07-14 Francis Y. Yan , Hudson Ayers , Chenzhi Zhu , Sadjad Fouladi , James Hong , Keyi Zhang , Philip Levis , Keith Winstein

This paper describes a high-performance, low-latency video surveillance system designed for resource-constrained environments. We have proposed a formal entropy-based adaptive frame buffering algorithm and integrated that with MobileNetV2…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Poojashree Chandrashekar Pankaj M Sajjanar

Reinforcement learning generalizes multi-armed bandit problems with additional difficulties of a longer planning horizon and unknown transition kernel. We explore a black-box reduction from discounted infinite-horizon tabular reinforcement…

Machine Learning · Computer Science 2024-03-12 Ian A. Kash , Lev Reyzin , Zishun Yu

Coordination among multiple access points (APs) is integral to IEEE 802.11bn (Wi-Fi 8) for managing contention in dense networks. This letter explores the benefits of Coordinated Spatial Reuse (C-SR) and proposes the use of reinforcement…

Networking and Internet Architecture · Computer Science 2025-04-01 Maksymilian Wojnar , Wojciech Ciezobka , Katarzyna Kosek-Szott , Krzysztof Rusek , Szymon Szott , David Nunez , Boris Bellalta

We study the Inverse Contextual Bandit (ICB) problem, in which a learner seeks to optimize a policy while an observer, who cannot access the learner's rewards and only observes actions, aims to recover the underlying problem parameters.…

Machine Learning · Computer Science 2026-03-05 Yuqi Kong , Xiao Zhang , Weiran Shen

We consider a multi-armed bandit problem in a setting where each arm produces a noisy reward realization which depends on an observable random covariate. As opposed to the traditional static multi-armed bandit problem, this setting allows…

Statistics Theory · Mathematics 2013-05-27 Vianney Perchet , Philippe Rigollet

We study a stochastic bandit problem with a general unknown reward function and a general unknown constraint function. Both functions can be non-linear (even non-convex) and are assumed to lie in a reproducing kernel Hilbert space (RKHS)…

Machine Learning · Computer Science 2022-03-30 Xingyu Zhou , Bo Ji

Many past attempts at modeling repeated Cournot games assume that demand is stationary. This does not align with real-world scenarios in which market demands can evolve over a product's lifetime for a myriad of reasons. In this paper, we…

Machine Learning · Computer Science 2022-01-04 Kshitija Taywade , Brent Harrison , Judy Goldsmith