Related papers: Clustering Financial Time Series: How Long is Enou…
Citation maturity time varies for different articles. However, the impact of all articles is measured in a fixed window. Clustering their citation trajectories helps understand the knowledge diffusion process and reveals that not all…
The community structure of complex networks reveals both their organization and hidden relationships among their constituents. Most community detection methods currently available are not deterministic, and their results typically depend on…
A sample of hundreds of simulated galaxy clusters is used to study the statistical properties of galaxy cluster formation. Individual assembly histories are discussed, the degree of virialization is demonstrated and various commonly used…
Number of connected devices is steadily increasing and these devices continuously generate data streams. Real-time processing of data streams is arousing interest despite many challenges. Clustering is one of the most suitable methods for…
Star clusters are ideal tracers of star formation activity in systems outside the volume that can be studied using individual, resolved stars. These unresolved clusters span orders of magnitude in brightness and mass, and their formation is…
Clustering methods must be tailored to the dataset it operates on, as there is no objective or universal definition of ``cluster,'' but nevertheless arbitrariness in the clustering method must be minimized. This paper develops a…
We study the theoretical and practical runtime limits of k-means and k-median clustering on large datasets. Since effectively all clustering methods are slower than the time it takes to read the dataset, the fastest approach is to quickly…
Time-series clustering serves as a powerful data mining technique for time-series data in the absence of prior knowledge about clusters. A large amount of time-series data with large size has been acquired and used in various research…
Using data from world stock exchange indices prior to and during periods of global financial crises, clusters and networks of indices are built for different thresholds and diverse periods of time, so that it is then possible to analyze how…
The study of pulsar glitch phenomena serves as a valuable probe into the dynamic properties of matter under extreme high-density conditions, offering insights into the physics within neutron stars. Providing theoretical explanations for the…
With rapidly increasing data, clustering algorithms are important tools for data analytics in modern research. They have been successfully applied to a wide range of domains; for instance, bioinformatics, speech recognition, and financial…
Improving the future of healthcare starts by better understanding the current actual practices in hospital settings. This motivates the objective of discovering typical care pathways from patient data. Revealing typical care pathways can be…
We present a new algorithm for clustering longitudinal data. Data of this type can be conceptualized as consisting of individuals and, for each such individual, observations of a time-dependent variable made at various times. Generically,…
The correlation function of galaxy clusters has often been used as a test of cosmological models. A number of assumptions are implicit in the comparison of theoretical expectations to data. Here we use an ensemble of ten large N-body…
Clustered standard errors and approximate randomization tests are popular inference methods that allow for dependence within observations. However, they require researchers to know the cluster structure ex ante. We propose a procedure to…
Unsupervised clustering of temporal data is both challenging and crucial in machine learning. In this paper, we show that neither traditional clustering methods, time series specific or even deep learning-based alternatives generalise well…
In this paper, we propose a technique for time series clustering using community detection in complex networks. Firstly, we present a method to transform a set of time series into a network using different distance functions, where each…
The measured correlations of financial time series in subsequent epochs change considerably as a function of time. When studying the whole correlation matrices, quasi-stationary patterns, referred to as market states, are seen by applying…
Time series are ubiquitous, and a measure to assess their similarity is a core part of many computational systems. In particular, the similarity measure is the most essential ingredient of time series clustering and classification systems.…
Human behavior modeling deals with learning and understanding behavior patterns inherent in humans' daily routines. Existing pattern mining techniques either assume human dynamics is strictly periodic, or require the number of modes as…