Related papers: Clustering logarithms up to six loops
k-means is a widely used clustering algorithm, but for $k$ clusters and a dataset size of $N$, each iteration of Lloyd's algorithm costs $O(kN)$ time. Although there are existing techniques to accelerate single Lloyd iterations, none of…
We describe how unintegrated parton distributions can be calculated from conventional integrated distributions. We extend and improve the 'last-step' evolution approach, and explain why doubly-unintegrated parton distributions are…
Clustering is one of the main tasks in exploratory data analysis and descriptive statistics where the main objective is partitioning observations in groups. Clustering has a broad range of application in varied domains like climate,…
Jet substructure is typically studied using clustering algorithms, such as kT, which arrange the jets' constituents into trees. Instead of considering a single tree per jet, we propose that multiple trees should be considered, weighted by…
We derive the leading non-global logarithms (NGLs) of ratios of jet masses m_{1,2} and a jet energy veto \Lambda due to soft gluons splitting into regions in and out of jets. Such NGLs appear in any exclusive jet cross section with multiple…
Analog quantum computation offers a route to machine learning using controllable physical dynamics as a computational resource. However, many existing approaches rely on task-specific protocols or observables that are difficult to access…
Clustering is a fundamental problem in unsupervised machine learning with many applications in data analysis. Popular clustering algorithms such as Lloyd's algorithm and $k$-means++ can take $\Omega(ndk)$ time when clustering $n$ points in…
We present a new general algorithm for calculating arbitrary jet cross sections in arbitrary scattering processes to next-to-leading accuracy in perturbative QCD. The algorithm is based on the subtraction method. The key ingredients are new…
We generalise the factorization of abelian gauge theory amplitudes to next-to-leading power (NLP) in a soft scale expansion, following a recent generalisation for Yukawa theory. From an all-order power counting analysis of leading and…
We study $k$-means clustering in a semi-supervised setting. Given an oracle that returns whether two given points belong to the same cluster in a fixed optimal clustering, we investigate the following question: how many oracle queries are…
We derive a novel factorization theorem for $N$-jettiness at hadron colliders, which incorporates coherence-violating effects induced by Glauber gluons and several new momentum modes. Their interplay generates coherence-violating logarithms…
We analyze online and mini-batch k-means variants. Both scale up the widely used Lloyd 's algorithm via stochastic approximation, and have become popular for large-scale clustering and unsupervised feature learning. We show, for the first…
We perform a two-loop calculation in Light Cone Perturbation Theory (LCPT) to evaluate the next-to-leading order nonsinglet splitting function. Our calculation demonstrates the methodology and feasibility of performing higher order…
Clustering is one of the most fundamental problems in data analysis and it has been studied extensively in the literature. Though many clustering algorithms have been proposed, clustering theories that justify the use of these clustering…
We consider the problem of fast time-series data clustering. Building on previous work modeling the correlation-based Hamiltonian of spin variables we present an updated fast non-expensive Agglomerative Likelihood Clustering algorithm…
Distributed data mining techniques and mainly distributed clustering are widely used in the last decade because they deal with very large and heterogeneous datasets which cannot be gathered centrally. Current distributed clustering…
Most convex and nonconvex clustering algorithms come with one crucial parameter: the $k$ in $k$-means. To this day, there is not one generally accepted way to accurately determine this parameter. Popular methods are simple yet theoretically…
Non-global logarithms arise from the sensitivity of collider observables to soft radiation in limited angular regions of phase space. Their resummation to next-to-leading logarithmic (NLL) order has been a long standing problem and its…
Theoretical predictions used in experimental analysis of LHC data have an inherent theoretical uncertainty associated to the relevant order in the perturbative expansion that the observable is computed. In this talk we briefly introduce and…
Clustering is a fundamental task in unsupervised learning. Previous research has focused on learning-augmented $k$-means in Euclidean metrics, limiting its applicability to complex data representations. In this paper, we generalize…