English

Synthetic Series-Symbol Data Generation for Time Series Foundation Models

Machine Learning 2025-10-21 v3 Artificial Intelligence

Abstract

Foundation models for time series analysis (TSA) have attracted significant attention. However, challenges such as training data scarcity and imbalance continue to hinder their development. Inspired by complex dynamic system theories, we design a series-symbol data generation mechanism, enabling the unrestricted creation of high-quality time series data paired with corresponding symbolic expressions. To leverage series-symbol data pairs with strong correlations, we develop SymTime, a pre-trained foundation model for enhancing time series representation using symbolic information. SymTime demonstrates competitive performance across five major TSA tasks when fine-tunes with downstream tasks, rivaling foundation models pre-trained on real-world datasets. This approach underscores the potential of series-symbol data generation and pretraining mechanisms in overcoming data scarcity and enhancing task performance. The code is available at https://github.com/wwhenxuan/SymTime.

Cite

@article{arxiv.2510.08445,
  title  = {Synthetic Series-Symbol Data Generation for Time Series Foundation Models},
  author = {Wenxuan Wang and Kai Wu and Yujian Betterest Li and Dan Wang and Xiaoyu Zhang},
  journal= {arXiv preprint arXiv:2510.08445},
  year   = {2025}
}

Comments

64 pages, 25 figures, 35 tables, NeurIPS 2025 accepted

R2 v1 2026-07-01T06:27:18.778Z