相关论文: Better Jet Clustering Algorithms
The linear k_\perp-factorization is part and parcel of the pQCD description of high energy hard processes off free nucleons. In the case of heavy nuclear targets the very concept of nuclear parton density becomes ill-defined as exemplified…
This paper proposes a Mixed-Integer Linear Programming approach for the Soft Graph Clustering Problem. This is the first method that simultaneously allocates membership proportion for vertices that lie in multiple clusters, and that…
Scaling clustering algorithms to massive data sets is a challenging task. Recently, several successful approaches based on data summarization methods, such as coresets and sketches, were proposed. While these techniques provide provably…
Clustering categorical data is an integral part of data mining and has attracted much attention recently. In this paper, we present k-histogram, a new efficient algorithm for clustering categorical data. The k-histogram algorithm extends…
Data clustering is a process of arranging similar data into groups. A clustering algorithm partitions a data set into several groups such that the similarity within a group is better than among groups. In this paper a hybrid clustering…
The incremental K-means clustering algorithm has already been proposed and analysed in paper [Chakraborty and Nagwani, 2011]. It is a very innovative approach which is applicable in periodically incremental environment and dealing with a…
Reduced k-means clustering is a method for clustering objects in a low-dimensional subspace. The advantage of this method is that both clustering of objects and low-dimensional subspace reflecting the cluster structure are simultaneously…
Clustering is a fundamental problem in machine learning where distance-based approaches have dominated the field for many decades. This set of problems is often tackled by partitioning the data into K clusters where the number of clusters…
We present a new method to expose the dead cone effect at colliders using iterative declustering techniques. Iterative declustering allows to unwind the jet clustering and to access the subjets or branches at different depths of the jet…
A new algorithm for jet finding in hadronic collisions is presented. The algorithm, based on a Gaussian filter in $(\eta,\phi)$, is specifically intended for use in heavy ion collisions and/or for detectors with limited acceptance. The…
We show that a general purpose clusterization algorithm, Deterministic Annealing, can be adapted to the problem of jet identification in particle production by high energy collisions. In particular we consider the problem of jet searching…
We consider stochastic settings for clustering, and develop provably-good approximation algorithms for a number of these notions. These algorithms yield better approximation ratios compared to the usual deterministic clustering setting.…
Jet finding is a type of optimization problem, where hadrons from a high-energy collision event are grouped into jets based on a clustering criterion. As three interesting examples, one can form a jet cluster that (1) optimizes the overall…
Progress in the theoretical understanding of parton branching dynamics within an expanding Quark Gluon Plasma relies on detailed and fair comparisons with experimental data for reconstructed jets. Such comparisons are only meaningful when…
We show how one can phrase the cut improvement problem for graphs as a sparse recovery problem, whence one can use algorithms originally developed for use in compressive sensing (such as SubspacePursuit or CoSaMP) to solve it. We show that…
We devise coresets for kernel $k$-Means with a general kernel, and use them to obtain new, more efficient, algorithms. Kernel $k$-Means has superior clustering capability compared to classical $k$-Means, particularly when clusters are…
This paper considers a canonical clustering problem where one receives unlabeled samples drawn from a balanced mixture of two elliptical distributions and aims for a classifier to estimate the labels. Many popular methods including PCA and…
We present a study on how to effectively reduce the dimensions of the $k$-means clustering problem, so that provably accurate approximations are obtained. Four algorithms are presented, two \textit{feature selection} and two \textit{feature…
We propose a novel method to accelerate Lloyd's algorithm for K-Means clustering. Unlike previous acceleration approaches that reduce computational cost per iterations or improve initialization, our approach is focused on reducing the…
Jet classification is an important ingredient in measurements and searches for new physics at particle coliders, and secondary vertex reconstruction is a key intermediate step in building powerful jet classifiers. We use a neural network to…