English

Cost-Efficient Online Decision Making: A Combinatorial Multi-Armed Bandit Approach

Machine Learning 2025-01-30 v3

Abstract

Online decision making plays a crucial role in numerous real-world applications. In many scenarios, the decision is made based on performing a sequence of tests on the incoming data points. However, performing all tests can be expensive and is not always possible. In this paper, we provide a novel formulation of the online decision making problem based on combinatorial multi-armed bandits and take the (possibly stochastic) cost of performing tests into account. Based on this formulation, we provide a new framework for cost-efficient online decision making which can utilize posterior sampling or BayesUCB for exploration. We provide a theoretical analysis of Thompson Sampling for cost-efficient online decision making, and present various experimental results that demonstrate the applicability of our framework to real-world problems.

Keywords

Cite

@article{arxiv.2308.10699,
  title  = {Cost-Efficient Online Decision Making: A Combinatorial Multi-Armed Bandit Approach},
  author = {Arman Rahbar and Niklas Åkerblom and Morteza Haghir Chehreghani},
  journal= {arXiv preprint arXiv:2308.10699},
  year   = {2025}
}

Comments

Accepted in Transactions on Machine Learning Research (01/2025)

R2 v1 2026-06-28T12:00:24.953Z