English

Building a Multivariate Time Series Benchmarking Datasets Inspired by Natural Language Processing (NLP)

Computation and Language 2024-10-15 v1 Artificial Intelligence

Abstract

Time series analysis has become increasingly important in various domains, and developing effective models relies heavily on high-quality benchmark datasets. Inspired by the success of Natural Language Processing (NLP) benchmark datasets in advancing pre-trained models, we propose a new approach to create a comprehensive benchmark dataset for time series analysis. This paper explores the methodologies used in NLP benchmark dataset creation and adapts them to the unique challenges of time series data. We discuss the process of curating diverse, representative, and challenging time series datasets, highlighting the importance of domain relevance and data complexity. Additionally, we investigate multi-task learning strategies that leverage the benchmark dataset to enhance the performance of time series models. This research contributes to the broader goal of advancing the state-of-the-art in time series modeling by adopting successful strategies from the NLP domain.

Keywords

Cite

@article{arxiv.2410.10687,
  title  = {Building a Multivariate Time Series Benchmarking Datasets Inspired by Natural Language Processing (NLP)},
  author = {Mohammad Asif Ibna Mustafa and Ferdinand Heinrich},
  journal= {arXiv preprint arXiv:2410.10687},
  year   = {2024}
}
R2 v1 2026-06-28T19:20:54.111Z