English

A Systematic Evaluation of Generated Time Series and Their Effects in Self-Supervised Pretraining

Machine Learning 2024-08-16 v1

Abstract

Self-supervised Pretrained Models (PTMs) have demonstrated remarkable performance in computer vision and natural language processing tasks. These successes have prompted researchers to design PTMs for time series data. In our experiments, most self-supervised time series PTMs were surpassed by simple supervised models. We hypothesize this undesired phenomenon may be caused by data scarcity. In response, we test six time series generation methods, use the generated data in pretraining in lieu of the real data, and examine the effects on classification performance. Our results indicate that replacing a real-data pretraining set with a greater volume of only generated samples produces noticeable improvement.

Keywords

Cite

@article{arxiv.2408.07869,
  title  = {A Systematic Evaluation of Generated Time Series and Their Effects in Self-Supervised Pretraining},
  author = {Audrey Der and Chin-Chia Michael Yeh and Xin Dai and Huiyuan Chen and Yan Zheng and Yujie Fan and Zhongfang Zhuang and Vivian Lai and Junpeng Wang and Liang Wang and Wei Zhang and Eamonn Keogh},
  journal= {arXiv preprint arXiv:2408.07869},
  year   = {2024}
}

Comments

To appear in CIKM 2024 as a short paper; the version here is the self-contained version that includes the non-mandatory supplementary material available on the paper's companion website