Scaling laws for large language models (LLMs) have provided useful guidance in training ever larger models for predictable performance gains. Time series forecasting shares a similar sequential structure to language, and is amenable to large-scale transformer architectures. Here we show that foundational decoder-only time series transformer models exhibit analogous scaling-behavior to LLMs, with architectural details (aspect ratio and number of heads) having a minimal effect over broad ranges. We assemble a large corpus of heterogenous time series data on which to train, and establish for the first time power-law scaling with parameter count, dataset size, and training compute, spanning five orders of magnitude.
@article{arxiv.2405.13867,
title = {Scaling-laws for Large Time-series Models},
author = {Thomas D. P. Edwards and James Alvey and Justin Alsing and Nam H. Nguyen and Benjamin D. Wandelt},
journal= {arXiv preprint arXiv:2405.13867},
year = {2025}
}
Comments
4 main pages (16 total), 4 figures; Accepted for oral presentation in Time Series in the Age of Large Models (TSALM) Workshop at Neurips 2024