English

AREAL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language Models

Machine Learning 2026-02-03 v1

Abstract

Reinforcement learning (RL) based post-training for large language models (LLMs) is computationally expensive, as it generates many rollout sequences that could frequently share long token prefixes. Existing RL frameworks usually process these sequences independently, repeatedly recomputing identical prefixes during forward and backward passes during policy model training, leading to substantial inefficiencies in computation and memory usage. Although prefix sharing naturally induces a tree structure over rollouts, prior tree-attention-based solutions rely on fully materialized attention masks and scale poorly in RL settings. In this paper, we introduce AREAL-DTA to efficiently exploit prefix sharing in RL training. AREAL-DTA employs a depth-first-search (DFS)-based execution strategy that dynamically traverses the rollout prefix tree during both forward and backward computation, materializing only a single root-to-leaf path at a time. To further improve scalability, AREAL-DTA incorporates a load-balanced distributed batching mechanism that dynamically constructs and processes prefix trees across multiple GPUs. Across the popular RL post-training workload, AREAL-DTA achieves up to 8.31×8.31\times in τ2\tau^2-bench higher training throughput.

Keywords

Cite

@article{arxiv.2602.00482,
  title  = {AREAL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language Models},
  author = {Jiarui Zhang and Yuchen Yang and Ran Yan and Zhiyu Mei and Liyuan Zhang and Daifeng Li and Wei Fu and Jiaxuan Gao and Shusheng Xu and Yi Wu and Binhang Yuan},
  journal= {arXiv preprint arXiv:2602.00482},
  year   = {2026}
}
R2 v1 2026-07-01T09:29:00.675Z