English
Related papers

Related papers: Entropy-Assisted Quality Pattern Identification in…

200 papers

Time series analysis has achieved great success in cyber security such as intrusion detection and device identification. Learning similarities among multiple time series is a crucial problem since it serves as the foundation for downstream…

Machine Learning · Computer Science 2025-06-23 Shaoyu Dou , Kai Yang , Yang Jiao , Chengbo Qiu , Kui Ren

Many organisations manage service quality and monitor a large set devices and servers where each entity is associated with telemetry or physical sensor data series. Recently, various methods have been proposed to detect behavioural…

Social and Information Networks · Computer Science 2023-05-10 Len Feremans , Boris Cule , Bart Goethals

Massive volumes of high-dimensional data that evolves over time is continuously collected by contemporary information processing systems, which brings up the problem of organizing this data into clusters, i.e. achieve the purpose of…

Machine Learning · Computer Science 2019-10-22 Di Xu , Tianhang Long , Junbin Gao

We consider the viability of a modularised mechanistic online machine learning framework to learn signals in low-frequency financial time series data. The framework is proved on daily sampled closing time-series data from JSE equity…

Statistical Finance · Quantitative Finance 2021-01-11 Joel da Costa , Tim Gebbie

Clustering is an important data mining technique that groups similar data records, recently categorical transaction clustering is received more attention. In this research, we study the problem of categorical data clustering for…

Databases · Computer Science 2017-05-03 Mahmoud Mahdi , Samir Abdelrahman , Reem Bahgat , Ismail Ismail

From a machine learning point of view, identifying a subset of relevant features from a real data set can be useful to improve the results achieved by classification methods and to reduce their time and space complexity. To achieve this…

Machine Learning · Computer Science 2017-05-23 Pietro Cassara , Alessandro Rozza , Mirco Nanni

Finding interdependency relations between (possibly multivariate) time series provides valuable knowledge about the processes that generate the signals. Information theory sets a natural framework for non-parametric measures of several…

Information Theory · Computer Science 2016-02-09 German Gomez-Herrero , Wei Wu , Kalle Rutanen , Miguel C. Soriano , Gordon Pipa , Raul Vicente

We present a simple model that uses time series momentum in order to construct strategies that systematically outperform their benchmark. The simplicity of our model is elegant: We only require a benchmark time series and several related…

Portfolio Management · Quantitative Finance 2020-02-12 Marc Rohloff , Alexander Vogt

Entropy estimation is of practical importance in information theory and statistical science. Many existing entropy estimators suffer from fast growing estimation bias with respect to dimensionality, rendering them unsuitable for…

Information Theory · Computer Science 2023-08-22 Ziqiao Ao , Jinglai Li

As access to high-quality, domain-specific data grows increasingly scarce, multi-epoch training has become a practical strategy for adapting large language models (LLMs). However, autoregressive models often suffer from performance…

Computation and Language · Computer Science 2025-12-30 Jiapeng Wang , Yiwen Hu , Yanzipeng Gao , Haoyu Wang , Shuo Wang , Hongyu Lu , Jiaxin Mao , Wayne Xin Zhao , Junyi Li , Xiao Zhang

Accurate and timely detection of cyber threats is critical to keeping our online economy and data safe. A key technique in early detection is the classification of unusual patterns of network behaviour, often hidden as low-frequency events…

Cryptography and Security · Computer Science 2024-05-01 Anthony Kenyon , Lipika Deka , David Elizondo

Permutation entropy and its associated frameworks are remarkable examples of physics-inspired techniques adept at processing complex and extensive datasets. Despite substantial progress in developing and applying these tools, their use has…

Data Analysis, Statistics and Probability · Physics 2024-08-14 Leonardo G. J. M. Voltarelli , Arthur A. B. Pessa , Luciano Zunino , Rafael S. Zola , Ervin K. Lenzi , Matjaz Perc , Haroldo V. Ribeiro

Existing clustering algorithms such as K-means often need to preset parameters such as the number of categories K, and such parameters may lead to the failure to output objective and consistent clustering results. This paper introduces a…

Machine Learning · Computer Science 2022-09-15 Shaodong Deng , Long Sheng , Jiayi Nie , Fuyi Deng

We investigate the relative information efficiency of financial markets by measuring the entropy of the time series of high frequency data. Our tool to measure efficiency is the Shannon entropy, applied to 2-symbol and 3-symbol…

Statistical Finance · Quantitative Finance 2016-09-15 Lucio Maria Calcagnile , Fulvio Corsi , Stefano Marmi

Predictive modeling in archaeology is essential for the understanding of people's behavior in the past and for guiding heritage conservation. However, spatial sampling bias caused by uneven research effort can severely limit model…

Applications · Statistics 2025-08-05 Mehmet Sıddık Çadırcı , Golnaz Shahtahmassebi

This paper presents a new clustering algorithm for space-time data based on the concepts of topological data analysis and in particular, persistent homology. Employing persistent homology - a flexible mathematical tool from algebraic…

Machine Learning · Statistics 2019-10-28 Umar Islambekov , Yulia Gel

Preserving biodiversity and ecosystem stability is a challenge that can be pursued through modern statistical mechanics modeling. Here we introduce a variational maximum entropy-based algorithm to evaluate the entropy in a minimal ecosystem…

Biological Physics · Physics 2018-10-17 Mattia Miotto , Lorenzo Monacelli

Data-driven modelling and computational predictions based on maximum entropy principle (MaxEnt-principle) aim at finding as-simple-as-possible - but not simpler then necessary - models that allow to avoid the data overfitting problem. We…

Computation · Statistics 2020-11-25 Horenko Illia , Marchenko Ganna , Gagliardini Patrick

We propose a three-stage framework for forecasting high-dimensional time-series data. Our method first estimates parameters for each univariate time series. Next, we use these parameters to cluster the time series. These clusters can be…

Machine Learning · Computer Science 2021-10-28 Reese Pathak , Rajat Sen , Nikhil Rao , N. Benjamin Erichson , Michael I. Jordan , Inderjit S. Dhillon

The study seeks to develop an effective strategy based on the novel framework of statistical arbitrage based on graph clustering algorithms. Amalgamation of quantitative and machine learning methods, including the Kelly criterion, and an…

Portfolio Management · Quantitative Finance 2024-06-18 Adam Korniejczuk , Robert Ślepaczuk