English

Robust Asynchronous Q-Learning under Reward and State Corruption via Batching

Machine Learning 2026-07-23 v1 Systems and Control

Abstract

Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback. Specifically, at each time-step, an adversary can perturb both the reward and state observations of the learner following the Huber contamination model. To defend against such data corruption, we propose {{\texttt{BR-Async-Q}}}: a novel, epoch-based, robust QQ-learning algorithm built upon two key ideas: (i) partitioning the online data stream into batches to reduce variance, and (ii) constructing robust estimates of the Bellman optimality operator using such batched data. We prove a high-probability \ell_\infty error bound for {{\texttt{BR-Async-Q}}} that matches that for vanilla QQ-learning, up to a small additive term that scales with the fraction of corrupted samples. To our knowledge, this provides the first robustness guarantee for asynchronous QQ-learning subject to both reward and state corruption. Furthermore, when only rewards are corrupted, the dependence of our algorithm's bound on the corruption fraction is minimax optimal.

Cite

@article{arxiv.2607.20822,
  title  = {Robust Asynchronous Q-Learning under Reward and State Corruption via Batching},
  author = {Sreejeet Maity and Aritra Mitra},
  journal= {arXiv preprint arXiv:2607.20822},
  year   = {2026}
}

Comments

To appear at the 65th IEEE Conference on Decision and Control (CDC) 2026