Related papers: Better Jet Clustering Algorithms
Discovering and clustering subspaces in high-dimensional data is a fundamental problem of machine learning with a wide range of applications in data mining, computer vision, and pattern recognition. Earlier methods divided the problem into…
This work proposes a clusterization algorithm called k-Morphological Sets (k-MS), based on morphological reconstruction and heuristics. k-MS is faster than the CPU-parallel k-Means in worst case scenarios and produces enhanced…
Jets are an important probe to identify the hard interaction of interest at the LHC. They are routinely used in Standard Model precision measurements as well as in searches for new heavy particles, including jet substructure methods. In…
It has always been a great challenge for clustering algorithms to automatically determine the cluster numbers according to the distribution of datasets. Several approaches have been proposed to address this issue, including the recent…
We consider the problem of subspace clustering: given points that lie on or near the union of many low-dimensional linear subspaces, recover the subspaces. To this end, one first identifies sets of points close to the same subspace and uses…
A robust clustering method for probabilities in Wasserstein space is introduced. This new "trimmed $k$-barycenters" approach relies on recent results on barycenters in Wasserstein space that allow intensive computation, as required by…
We propose a simple, projection-based algorithm for clustering mixtures of discrete (Bernoulli) distributions. Unlike previous approaches that rely on coordinate-specific ``combinatorial projections,'' our algorithm is rotationally…
For many observables, the most difficult part of a single logarithmic resummation is the analytical treatment of the observable's dependence on multiple emissions. We present a general numerical method, which allows the resummation…
We show that the next-to-leading order perturbative prediction, matched with the next-to-leading logarithmic approximation for predicting both two-, three- and four-jet rates using the Durham jet-clustering algorithm, in the 0.001 < ycut <…
Jet angularities are a class of jet substructure observables where a continuous parameter is introduced in order to interpolate between different classic observables such as the jet mass and jet broadening. We consider jet angularities…
Clustering, a fundamental activity in unsupervised learning, is notoriously difficult when the feature space is high-dimensional. Fortunately, in many realistic scenarios, only a handful of features are relevant in distinguishing clusters.…
We consider the classic Correlation Clustering problem: Given a complete graph where edges are labelled either $+$ or $-$, the goal is to find a partition of the vertices that minimizes the number of the \pedges across parts plus the number…
Clustering is a fundamental problem in data analysis. In differentially private clustering, the goal is to identify $k$ cluster centers without disclosing information on individual data points. Despite significant research progress, the…
We present the next-to-leading order (O(alpha_s^3)) perturbative QCD predictions for e^+e^- annihilation into four jets. A previous calculation omitted the O(alpha_s^3) terms suppressed by one or more powers of 1/N_c^2, where N_c is the…
The origin of the breaking of conventional linear k_\perp-factorization for hard processes in a nuclear environment is by now well established. The realization of the nonlinear nuclear k_\perp-factorization which emerges instead was found…
We extend the theoretical analysis of a recently proposed single subspace learning algorithm, called Dual Principal Component Pursuit (DPCP), to the case where the data are drawn from of a union of hyperplanes. To gain insight into the…
We compute the leading clustering (abelian non-global) logarithms, which arise in the distribution of non-global QCD observables when final-state partons are clustered using the $k_t$ jet algorithm, up to six loops in perturbation theory.…
Motivated by the fact that distances between data points in many real-world clustering instances are often based on heuristic measures, Bilu and Linial~\cite{BL} proposed analyzing objective based clustering problems under the assumption…
Clustering methods are often used in physics education research (PER) to identify subgroups of individuals within a population who share similar response patterns or characteristics. K-means (or k-modes, for categorical data) is one of the…
Pruning filters is an effective method for accelerating deep neural networks (DNNs), but most existing approaches prune filters on a pre-trained network directly which limits in acceleration. Although each filter has its own effect in DNNs,…