English

Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks

Machine Learning 2026-05-21 v3 Artificial Intelligence

Abstract

Meta reinforcement learning aims to develop policies that generalize to unseen tasks sampled from a task distribution. While context-based meta-RL methods improve task representation using task latents, they often struggle with out-of-distribution (OOD) tasks. To address this, we propose Task-Aware Virtual Training (TAVT), a novel algorithm that accurately captures task characteristics for both training and OOD scenarios using metric-based representation learning. Our method successfully preserves task characteristics in virtual tasks and employs a state regularization technique to mitigate overestimation errors in state-varying environments. Numerical results demonstrate that TAVT significantly enhances generalization to OOD tasks across various MuJoCo and MetaWorld environments. Our code is available at https://github.com/JM-Kim-94/tavt.git.

Keywords

Cite

@article{arxiv.2502.02834,
  title  = {Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks},
  author = {Jeongmo Kim and Yisak Park and Minung Kim and Seungyul Han},
  journal= {arXiv preprint arXiv:2502.02834},
  year   = {2026}
}

Comments

9 pages main paper, 20 pages appendices with reference. Accepted to ICML 2025

R2 v1 2026-06-28T21:32:55.081Z