中文

Learning to Sparsify Stochastic Linear Bandits

机器学习 2026-05-12 v1 系统与控制 系统与控制 最优化与控制

摘要

This paper addresses the problem of learning to sparsify stochastic linear bandits, where a decision-maker sequentially selects actions from a high-dimensional space subject to a sparsity constraint on the number of nonzero elements in the action vector. The key challenge lies in minimizing cumulative regret while tackling the potential NP-hardness of finding optimal sparse actions due to the inherent combinatorial structure of the problem. We propose an adaptively phased exploration and exploitation algorithmic framework, utilizing ordinary least squares for parameter learning and specialized subroutines for sparse action selection. When the action set is a Euclidean ball, optimal sparse actions can be efficiently computed, enabling us to establish a O~(dT)\tilde{\mathcal{O}}(d\sqrt{T}) regret, where dd is the dimension of the action vector and TT is the time horizon length. For general convex and compact action sets where finding optimal sparse actions is intractable, we employ a greedy subroutine. For general strongly convex action sets, we derive a O~(dT)\tilde{\mathcal{O}}(d \sqrt{T}) α\alpha-regret; for general compact sets lacking strong convexity, we establish a O~(dT2/3)\tilde{\mathcal{O}}(d T^{2/3}) α\alpha-regret, where α\alpha pertains to the approximation ratio of the greedy algorithm. Finally, we validate the performance of our algorithms using extensive experiments including an application to recommendation system.

关键词

引用

@article{arxiv.2605.10151,
  title  = {Learning to Sparsify Stochastic Linear Bandits},
  author = {Zhengmiao Wang and Ming Chi and Zhi-Wei Liu and Lintao Ye and Carla Fabiana Chiasserini},
  journal= {arXiv preprint arXiv:2605.10151},
  year   = {2026}
}

备注

Include all the omitted details and proofs from the conference paper accepted to IJCAI 2026