Related papers: Validity of Clusters Produced By kernel-$k$-means …
We consider the problem of causal structure learning in the setting of heterogeneous populations, i.e., populations in which a single causal structure does not adequately represent all population members, as is common in biological and…
Enhancing classical machine learning (ML) algorithms through quantum kernels is a rapidly growing research topic in quantum machine learning (QML). A key challenge in using kernels -- both classical and quantum -- is that ML workflows…
For the unitary ensembles of $N\times N$ Hermitian matrices associated with a weight function $w$ there is a kernel, expressible in terms of the polynomials orthogonal with respect to the weight function, which plays an important role. For…
We design a new algorithm for the Euclidean $k$-means problem that operates in the local model of differential privacy. Unlike in the non-private literature, differentially private algorithms for the $k$-means objective incur both additive…
Clustering is an unsupervised learning method that constitutes a cornerstone of an intelligent data analysis process. It is used for the exploration of inter-relationships among a collection of patterns, by organizing them into homogeneous…
Reproducibility is essential in machine learning because it ensures that a model or experiment yields the same scientific conclusion. For specific algorithms repeatability with bitwise identical results is also a key for scientific…
Quantum replacer codes are codes that can be protected from errors induced by a given set of quantum replacer channels, an important class of quantum channels that includes the erasures of subsets of qubits that arise in quantum error…
Spectral clustering has become one of the most widely used clustering techniques when the structure of the individual clusters is non-convex or highly anisotropic. Yet, despite its immense popularity, there exists fairly little theory about…
In this paper we extend an earlier result within Dempster-Shafer theory ["Fast Dempster-Shafer Clustering Using a Neural Network Structure," in Proc. Seventh Int. Conf. Information Processing and Management of Uncertainty in Knowledge-Based…
We prove two theorems of Paley and Wiener in the slice regular setting. As an application, we can compute the reproducing kernel for the slice regular Paley-Wiener space, and obtain a related sampling theorem.
This paper focuses on the problem of kernelizing an existing supervised Mahalanobis distance learner. The following features are included in the paper. Firstly, three popular learners, namely, "neighborhood component analysis", "large…
In this paper, the easier methods of my thesis are applied to give a simple proof of a theorem of Goussarov. The theorem relates two possible notions of finite type equivalence of knots, links or string links, showing that the resulting…
We deal with kernel theorems for modulation spaces. We completely characterize the continuity of a linear operator on the modulation spaces $M^p$ for every $1\leq p\leq\infty$, by the membership of its kernel to (mixed) modulation spaces.…
A kernel density is an aggregate of kernel functions, which are itself densities and could be kernel densities. This is used to decompose a kernel into its constituent parts. Pearson's test for equality of proportions is applied to…
K-means is one of the most widely used algorithms for clustering in Data Mining applications, which attempts to minimize the sum of the square of the Euclidean distance of the points in the clusters from the respective means of the…
We study the following distribution clustering problem: Given a hidden partition of $k$ distributions into two groups, such that the distributions within each group are the same, and the two distributions associated with the two clusters…
This paper introduces an approach for detecting differences in the first-order structures of spatial point patterns. The proposed approach leverages the kernel mean embedding in a novel way by introducing its approximate version tailored to…
For each partition of a data set into a given number of parts there is a partition such that every part is as much as possible a good model (an "algorithmic sufficient statistic") for the data in that part. Since this can be done for every…
Clustering is a widely used technique with a long and rich history in a variety of areas. However, most existing algorithms do not scale well to large datasets, or are missing theoretical guarantees of convergence. This paper introduces a…
Neural tangent kernels (NTKs) have been proposed to study the behavior of trained neural networks from the perspective of Gaussian processes. An important result in this body of work is the theorem of equivalence between a trained neural…