Stochastic Low-Rank Bandits
Abstract
Many problems in computer vision and recommender systems involve low-rank matrices. In this work, we study the problem of finding the maximum entry of a stochastic low-rank matrix from sequential observations. At each step, a learning agent chooses pairs of row and column arms, and receives the noisy product of their latent values as a reward. The main challenge is that the latent values are unobserved. We identify a class of non-negative matrices whose maximum entry can be found statistically efficiently and propose an algorithm for finding them, which we call LowRankElim. We derive a upper bound on its -step regret, where is the number of rows, is the number of columns, is the rank of the matrix, and is the minimum gap. The bound depends on other problem-specific constants that clearly do not depend . To the best of our knowledge, this is the first such result in the literature.
Cite
@article{arxiv.1712.04644,
title = {Stochastic Low-Rank Bandits},
author = {Branislav Kveton and Csaba Szepesvari and Anup Rao and Zheng Wen and Yasin Abbasi-Yadkori and S. Muthukrishnan},
journal= {arXiv preprint arXiv:1712.04644},
year = {2017}
}