We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. Because this source state is already computed during ordinary decoding, LRT adds a cross-layer recurrent latent pathway across positions without inserting pause tokens or extra depth loops, and the standard attention mechanism and KV-cache interface are preserved. To pretrain this recurrence at scale without sequentially unrolling the transformer, we introduce interleaved parallel training: a single full-sequence initialization forward pass builds a shared buffer; then disjoint position subsets are refined in parallel and written back, so that all tokens receive recurrent-memory-aware supervision at roughly 2 times baseline compute. Across nanochat style backbones and a wide range of tokens-per-parameter budgets, LRT improves both language-modeling loss and in-context learning under matched effective compute while adding as little as 0.3% parameters.
@article{arxiv.2605.26797,
title = {Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior},
author = {Zeyi Huang and Xuehai He and LiLiang Ren and Yiping Wang and Baolin Peng and Hao Cheng and Shuohang Wang and Pengcheng He and Jianfeng Gao and Yong Jae Lee and Yelong Shen},
journal= {arXiv preprint arXiv:2605.26797},
year = {2026}
}