Related papers: Classifying the typefaces of the Gutenberg 42-line…
We first study a new family of graded quiver varieties together with a new $t$-deformation of the associated Grothendieck rings. This provides the geometric foundations for a joint paper by Yoshiyuki Kimura and the author. We further…
Cluster mass profiles are tests of models of structure formation. Only two current observational methods of determining the mass profile, gravitational lensing and the caustic technique, are independent of the assumption of dynamical…
Clustering is one of the main tasks in exploratory data analysis and descriptive statistics where the main objective is partitioning observations in groups. Clustering has a broad range of application in varied domains like climate,…
In this paper we cluster 330 classical music pieces collected from MusicNet database based on their musical note sequence. We use shingling and chord trajectory matrices to create signature for each music piece and performed spectral…
We show that margin-based bitext mining in a multilingual sentence space can be applied to monolingual corpora of billions of sentences. We are using ten snapshots of a curated common crawl corpus (Wenzek et al., 2019) totalling 32.7…
Graph-based clustering methods have demonstrated the effectiveness in various applications. Generally, existing graph-based clustering methods first construct a graph to represent the input data and then partition it to generate the…
Graph clustering is widely used in analysis of biological networks, social networks and etc. For over a decade many graph clustering algorithms have been published, however a comprehensive and consistent performance comparison is not…
Filaments of galaxies are known to stretch between galaxy clusters at all redshifts in a complex manner. In this Letter, we present an analysis of the frequency and distribution of inter-cluster galaxy filaments selected from the 2dF Galaxy…
A keyword search on constrained clustering on Web-of-Science returned just under 3,000 documents. We ran automatic analyses of those, and compiled our own bibliography of 183 papers which we analysed in more detail based on their topic and…
Ongoing and future spectroscopic surveys will measure numerous galaxy redshifts within tens of thousands of galaxy clusters. However, the sampling within these clusters will be low, 15 < N < 50 per cluster. With such data, it will be…
We study the cluster category of a canonical algebra A in terms of the hereditary category of coherent sheaves over the corresponding weighted projective line X. As an application we determine the automorphism group of the cluster category…
One basic requirement of many studies is the necessity of classifying data. Clustering is a proposed method for summarizing networks. Clustering methods can be divided into two categories named model-based approaches and algorithmic…
There is a growing need for unbiased clustering methods, ideally automated. We have developed a topology-based analysis tool called Two-Tier Mapper (TTMap) to detect subgroups in global gene expression datasets and identify their…
A clustering analysis is performed on two samples of $\sim 600$ faint galaxies each, in two widely separated regions of the sky, including the Hubble Deep Field. One of the survey regions is configured so that some galaxy pairs span angular…
This paper is a chapter in the forthcoming Handbook of Cluster Analysis, Hennig et al. (2015). For definitions of basic clustering methods and some further methodology, other chapters of the Handbook are referred to. To read this version of…
Galaxy cluster number counts are an important probe to constrain cosmological parameters. One of the main ingredients of the analysis, along with accurate estimates of the clusters' masses, is the selection function, and in particular the…
Recent progress in generative language models has enabled machines to generate astonishingly realistic texts. While there are many legitimate applications of such models, there is also a rising need to distinguish machine-generated texts…
Cluster analysis is a popular unsupervised learning tool used in many disciplines to identify heterogeneous sub-populations within a sample. However, validating cluster analysis results and determining the number of clusters in a data set…
We use machine learning to study the moduli space of genus two curves, specifically focusing on detecting whether a genus two curve has $(n, n)$-split Jacobian. Based on such techniques, we observe that there are very few rational moduli…
We find by applying MacMahon's partition analysis that all magic labellings of the cube are of eight types, each generated by six basis elements. A combinatorial proof of this fact is given. The number of magic labellings of the cube is…