相关论文: Asynchronously parallelised percolation on distrib…
We present simulations of a 3-d percolation model studied recently by K.J. Schrenk et al. [Phys. Rev. Lett. 116, 055701 (2016)], obtained with a new and more efficient algorithm. They confirm most of their results in spite of larger systems…
The kernel-based multi-scale method has been proven to be a powerful approximation method for scattered data approximation problems which is computationally superior to conventional kernel-based interpolation techniques. The multi-scale…
Many applications of interest involve data that can be analyzed as unit vectors on a d-dimensional sphere. Specific examples include text mining, in particular clustering of documents, biology, astronomy and medicine among others. Previous…
Motivated by a computer science algorithm known as `linear probing with hashing' we study a new type of percolation model whose basic features include a sequential `dropping' of particles on a substrate followed by their transport via a…
As the data size in Machine Learning fields grows exponentially, it is inevitable to accelerate the computation by utilizing the ever-growing large number of available cores provided by high-performance computing hardware. However, existing…
Markov Chain Monte Carlo methods are algorithms used to sample probability distributions, commonly used to sample the Boltzmann distribution of physical/chemical models (e.g., protein folding, Ising model, etc.). This allows us to study…
This article presents new algorithms for massively parallel granular dynamics simulations on distributed memory architectures using a domain partitioning approach. Collisions are modelled with hard contacts in order to hide their…
This paper introduces a novel K-means clustering algorithm, an advancement on the conventional Big-means methodology. The proposed method efficiently integrates parallel processing, stochastic sampling, and competitive optimization to…
We investigate the problem of growing clusters, which is modeled by two dimensional disks and three dimensional droplets. In this model we place a number of seeds on random locations on a lattice with an initial occupation probability, $p$.…
We consider the problem of clustering a sample of probability distributions from a random distribution on $\mathbb R^p$. Our proposed partitioning method makes use of a symmetric, positive-definite kernel $k$ and its associated reproducing…
The simulation of the metabolism in mammalian cells becomes a severe problem if spatial distributions must be taken into account. Especially the cytoplasm has a very complex geometric structure which cannot be handled by standard…
For massive data stored at multiple machines, we propose a distributed subsampling procedure for the composite quantile regression. By establishing the consistency and asymptotic normality of the composite quantile regression estimator from…
Clustering algorithms start with a fixed divergence, which captures the possibly asymmetric distance between a sample and a centroid. In the mixture model setting, the sample distribution plays the same role. When all attributes have the…
In this paper we propose a new approach for Big Data mining and analysis. This new approach works well on distributed datasets and deals with data clustering task of the analysis. The approach consists of two main phases, the first phase…
Recent results from high-resolution solar granulation observations indicate the existence of a population of small granular cells that are smaller than 600 km in diameter. These small convective cells strongly contribute to the total area…
One of the main challenges in distributed computing is building interfaces and APIs that allow programmers with limited background in distributed systems to write scalable, performant, and fault-tolerant applications on large clusters. In…
Motivation: With the development of droplet based systems, massive single cell transcriptome data has become available, which enables analysis of cellular and molecular processes at single cell resolution and is instrumental to…
The statistical analysis of cosmic large-scale structure is most often based on simple two-point summary statistics, like the power spectrum or the two-point correlation function of a sample of galaxies or other types of tracers. In…
We present a study of connectivity percolation in suspensions of hard spherocylinders by means of Monte Carlo simulation and connectedness percolation theory. We focus attention on polydispersity in the length, the diameter and the…
In this paper we address the problem of estimating the posterior distribution of the static parameters of a continuous time state space model with discrete time observations by an algorithm that combines the Kalman filter and a particle…