English
Related papers

Related papers: Detecting British Columbia Coastal Rainfall Patter…

200 papers

Analyzing large datasets with distributed dataflow systems requires the use of clusters. Public cloud providers offer a large variety and quantity of resources that can be used for such clusters. However, picking the appropriate resources…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-04-28 Jonathan Will , Jonathan Bader , Lauritz Thamsen

Archetype and archetypoid analysis can be extended to functional data. Each function is represented as a mixture of actual observations (functional archetypoids) or functional archetypes, which are a mixture of observations in the data set.…

Methodology · Statistics 2016-09-02 Irene Epifanio

We review clustering as an analysis tool and the underlying concepts from an introductory perspective. What is clustering and how can clusterings be realised programmatically? How can data be represented and prepared for a clustering task?…

Machine Learning · Computer Science 2022-12-05 Jan-Oliver Felix Kapp-Joswig , Bettina G. Keller

Spatial statistics is concerned with the analysis of data that have spatial locations associated with them, and those locations are used to model statistical dependence between the data. The spatial data are treated as a single realisation…

Methodology · Statistics 2022-02-09 Noel Cressie , Matthew Sainsbury-Dale , Andrew Zammit-Mangion

Classically, Bayesian clustering interprets each component of a mixture model as a cluster. The inferred clustering posterior is highly sensitive to any inaccuracies in the kernel within each component. As this kernel is made more flexible,…

Methodology · Statistics 2025-12-12 David Buch , Miheer Dewaskar , David B. Dunson

We present a new method to detect meteor showers using the Density-Based Spatial Clustering of Applications with Noise algorithm (DBSCAN; Ester et al. 1996). DBSCAN is a modern cluster detection algorithm that is well suited to the problem…

Earth and Planetary Astrophysics · Physics 2018-03-14 Glenn Sugar , Althea Moorhead , Peter Brown , Bill Cooke

Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. For most applications, applying clustering is only appropriate when cluster structure is present. As such, the study of clusterability,…

Machine Learning · Statistics 2018-10-30 A. Adolfsson , M. Ackerman , N. C. Brownstein

A recurrent question in climate risk analysis is determining how climate change will affect heavy precipitation patterns. Dividing the globe into homogeneous sub-regions should improve the modelling of heavy precipitation by inferring…

Methodology · Statistics 2021-11-02 Philomène Le Gall , Anne-Catherine Favre , Philippe Naveau , Alexandre Tuel

Hydroclimatic processes are characterized by heterogeneous spatiotemporal correlation structures and marginal distributions that can be continuous, mixed-type, discrete or even binary. Simulating exactly such processes can greatly improve…

Methodology · Statistics 2017-07-24 Simon Michael Papalexiou

Accurately predicting long-term rainfall is challenging. Global climate indices, such as the El Ni\~no-Southern Oscillation, are standard input features for machine learning. However, a significant gap persists regarding local-scale indices…

Machine Learning · Computer Science 2026-02-20 Kiattikun Chobtham

Clustering is one of the most common unsupervised learning tasks in machine learning and data mining. Clustering algorithms have been used in a plethora of applications across several scientific fields. However, there has been limited…

Machine Learning · Computer Science 2017-02-09 Quang N. Tran , Ba-Ngu Vo , Dinh Phung , Ba-Tuong Vo

The mixture of Gaussian distributions, a soft version of k-means , is considered a state-of-the-art clustering algorithm. It is widely used in computer vision for selecting classes, e.g., color, texture, and shapes. In this algorithm, each…

Machine Learning · Statistics 2016-12-30 Mahajabin Rahman , Davi Geiger

Clustering of event stream data is of great importance in many application scenarios, including but not limited to, e-commerce, electronic health, online testing, mobile music service, etc. Existing clustering algorithms fail to take…

Methodology · Statistics 2024-05-29 Yuecheng Zhang , Guanhua Fang , Wen Yu

The growing field of large-scale time domain astronomy requires methods for probabilistic data analysis that are computationally tractable, even with large datasets. Gaussian Processes are a popular class of models used for this purpose…

Instrumentation and Methods for Astrophysics · Physics 2017-11-15 Daniel Foreman-Mackey , Eric Agol , Sivaram Ambikasaran , Ruth Angus

In many areas of science one aims to estimate latent sub-population mean curves based only on observations of aggregated population curves. By aggregated curves we mean linear combination of functional data that cannot be observed…

Methodology · Statistics 2011-02-15 Ronaldo Dias , Nancy L. Garcia , Alexandra M. Schmidt

Time series clustering is a central machine learning task with applications in many fields. While the majority of the methods focus on real-valued time series, very few works consider series with discrete response. In this paper, the…

Machine Learning · Statistics 2023-04-25 Ángel López Oriona , Christian Weiss , José Antonio Vilar

Functional data describe a wide range of processes, such as growth curves and spectral absorption. In this study, we analyze air pollution data from the In-service Aircraft for a Global Observing System, focusing on the spatial interactions…

Methodology · Statistics 2024-11-14 Rita Fici , Gianluca Sottile , Luigi Augugliaro , Ernst-Jan Camiel Wit

Biclustering is an unsupervised data mining technique that aims to unveil patterns (biclusters) from gene expression data matrices. In the framework of this thesis, we propose new biclustering algorithms for microarray data. The latter is…

Machine Learning · Computer Science 2018-11-26 Amina Houari

Clustering algorithms are often used to find subpopulations in exploratory data analysis workflows. Not only the clusters themselves, but also their shape can represent meaningful subpopulations. In this paper, we present FLASC, an…

Machine Learning · Computer Science 2025-04-23 D. M. Bot , J. Peeters , J. Liesenborgs , J. Aerts

Distributed dataflow systems enable the use of clusters for scalable data analytics. However, selecting appropriate cluster resources for a processing job is often not straightforward. Performance models trained on historical executions of…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-10-19 Dominik Scheinert , Lauritz Thamsen , Houkun Zhu , Jonathan Will , Alexander Acker , Thorsten Wittkopp , Odej Kao