Related papers: Classifying the typefaces of the Gutenberg 42-line…
Document clustering is generally the first step for topic identification. Since many clustering methods operate on the similarities between documents, it is important to build representations of these documents which keep their semantics as…
Biclustering is a class of techniques that simultaneously clusters the rows and columns of a matrix to sort heterogeneous data into homogeneous blocks. Although many algorithms have been proposed to find biclusters, existing methods suffer…
We present a novel approach for finding and evaluating structural models of small metallic nanoparticles. Rather than fitting a single model with many degrees of freedom, the approach algorithmically builds libraries of nanoparticle…
We have reanalyzed a data set of 99 low redshift ($ z < 0.1 $) Abell clusters and determined their shapes. For this, three different measures are used. We use Monte-Carlo simulations to investigate the errors in the methods. The corrected…
Galaxy diversification proceeds by transforming events like accretion, interaction or mergers. These explain the formation and evolution of galaxies that can now be described with many observables. Multivariate analyses are the obvious…
Power-law type distributions are extensively found when studying the behaviour of many complex systems. However, due to limitations in data acquisition, empirical datasets often only cover a narrow range of observation, making it difficult…
We begin by reviewing some probabilistic results about the Dirichlet Process and its close relatives, focussing on their implications for statistical modelling and analysis. We then introduce a class of simple mixture models in which…
What are the best methods of capturing thematic similarity between literary texts? Knowing the answer to this question would be useful for automatic clustering of book genres, or any other thematic grouping. This paper compares a variety of…
The clustering of a data set is one of the core tasks in data analytics. Many clustering algorithms exhibit a strong contrast between a favorable performance in practice and bad theoretical worst-cases. Prime examples are least-squares…
I compare the mass values obtained with data taken from the Arcminute Microkelvin Imager (AMI) radio interferometer system and from the Planck satellite. The former of these uses a Bayesian analysis pipeline that parameterises a cluster in…
The fundamental problem of similarity studies, in the frame of data-mining, is to examine and detect similar items in articles, papers, books, with huge sizes. In this paper, we are interested in the probabilistic, and the statistical and…
The different regimes of gravitational lensing constitutes an interesting tool in order to map the mass distribution in galaxy clusters on different scales. In this proceedings article, I review some work I have performed on this topic.…
We present line-strengths and kinematics from the central regions of 32 galaxies with Hubble types ranging from E to Sbc. Spectral indices, based on the Lick system, are measured in the optical and near infra-red (NIR). The 24 indices…
Aims. We use the spectra of more than 30,000 red giant branch (RGB) stars in 25 globular clusters (GC), obtained within the MUSE survey of Galactic globular clusters, to calibrate the Ca II triplet (CaT) metallicity relation and derive…
We study the geometry and topology of the large-scale structure traced by galaxy clusters in numerical simulations of a box of side 320 $h^{-1}$ Mpc, and compare them with available data on real clusters. The simulations we use are…
A recent evaluation of three-loop nonplanar Feynman integrals contributing to Higgs plus jet production has established their dependence on two novel symbol letters. We show that the resulting alphabet is described by a $G_2$ cluster…
As a kind of basic machine learning method, clustering algorithms group data points into different categories based on their similarity or distribution. We present a clustering algorithm by finding hyper-planes to distinguish the data…
Nine popular clustering methods are applied to 42 real data sets. The aim is to give a detailed characterisation of the methods by means of several cluster validation indexes that measure various individual aspects of the resulting clusters…
We use high-precision photometry of red-giant-branch (RGB) stars in 57 Galactic globular clusters (GCs), mostly from the `Hubble Space Telescope (HST) UV Legacy Survey of Galactic globular clusters', to identify and characterize their…
Genome wide comparisons between enteric bacteria yield large sets of conserved putative regulatory sites on a gene by gene basis that need to be clustered into regulons. Using the assumption that regulatory sites can be represented as…