English

Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds

Machine Learning 2026-05-13 v1

Abstract

We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kullback--Leibler (KL) divergence. Our goal is to characterize how the sequence horizon HH affects both approximation and estimation errors in this joint-distribution, sequence-level regime. By establishing matching upper and lower bounds, we provide, to our knowledge, the first complete characterization of long-horizon error behavior under the natural joint KL objective, with improved rates and optimality justification relative to existing work. On the approximation side, we show that joint KL admits a horizon-free approximation factor, in sharp contrast to Hellinger-based analyses that exhibit an Ω(H)\Omega(H) dependence for computationally efficient methods; this isolates the choice of divergence as the source of approximation amplification. On the estimation side, we prove a fundamental information-theoretic lower bound of order Ω(H)\Omega(H) that holds for both decomposable policy classes and fully shared policies, matching the O~(H)\widetilde O(H) upper bounds achieved by computationally efficient algorithms. Our analysis clarifies the landscape of recent autoregressive learning results by aligning the log-loss training objective, the sequence-level evaluation metric, and the approximation metric {\color{black}through a sharp joint-KL oracle theory}. We further show that these joint-KL guarantees imply policy learning regret bounds at rates matching prior imitation learning literature.

Keywords

Cite

@article{arxiv.2605.12316,
  title  = {Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds},
  author = {Yunbei Xu and Yuzhe Yuan and Ruohan Zhan},
  journal= {arXiv preprint arXiv:2605.12316},
  year   = {2026}
}
R2 v1 2026-07-22T07:08:02.010Z