English

The Sample-Communication Complexity Trade-off in Federated Q-Learning

Machine Learning 2024-10-31 v2 Optimization and Control Machine Learning

Abstract

We consider the problem of federated Q-learning, where MM agents aim to collaboratively learn the optimal Q-function of an unknown infinite-horizon Markov decision process with finite state and action spaces. We investigate the trade-off between sample and communication complexities for the widely used class of intermittent communication algorithms. We first establish the converse result, where it is shown that a federated Q-learning algorithm that offers any speedup with respect to the number of agents in the per-agent sample complexity needs to incur a communication cost of at least an order of 11γ\frac{1}{1-\gamma} up to logarithmic factors, where γ\gamma is the discount factor. We also propose a new algorithm, called Fed-DVR-Q, which is the first federated Q-learning algorithm to simultaneously achieve order-optimal sample and communication complexities. Thus, together these results provide a complete characterization of the sample-communication complexity trade-off in federated Q-learning.

Keywords

Cite

@article{arxiv.2408.16981,
  title  = {The Sample-Communication Complexity Trade-off in Federated Q-Learning},
  author = {Sudeep Salgia and Yuejie Chi},
  journal= {arXiv preprint arXiv:2408.16981},
  year   = {2024}
}

Comments

Accepted to NeurIPS 2024

R2 v1 2026-06-28T18:28:21.952Z