Related papers: Limit theorems for unbounded cluster functionals o…
Convex clustering is a well-regarded clustering method, resembling the similar centroid-based approach of Lloyd's $k$-means, without requiring a predefined cluster count. It starts with each data point as its centroid and iteratively merges…
We develop a model in which interactions between nodes of a dynamic network are counted by non homogeneous Poisson processes. In a block modelling perspective, nodes belong to hidden clusters (whose number is unknown) and the intensity…
Block coordinate methods have been extensively studied for minimization problems, where they come with significant complexity improvements whenever the considered problems are compatible with block decomposition and, moreover, block…
In [1], the authors consider a random walk $(Z_{n,1},\ldots,Z_{n,K+1})\in \mathbb{Z}^{K+1}$ with the constraint that each coordinate of the walk is at distance one from the following one. A functional central limit theorem for the first…
We prove functional limit theorems for dynamical systems in the presence of clusters of large values which, when summed and suitably normalised, get collapsed in a jump of the limiting process observed at the same time point. To keep track…
I introduce a generic method for inference about a scalar parameter in research designs with a finite number of heterogeneous clusters where only a single cluster received treatment. This situation is commonplace in…
Any limiting point process for the time normalized exceedances of high levels by a stationary sequence is necessarily compound Poisson under appropriate long range dependence conditions. Typically exceedances appear in clusters. The…
We study a Markovian model for the random fragmentation of an object. At each time, the state consists of a collection of blocks. Each block waits an exponential amount of time with parameter given by its size to some power $\alpha$,…
Many cluster similarity indices are used to evaluate clustering algorithms, and choosing the best one for a particular task remains an open problem. We demonstrate that this problem is crucial: there are many disagreements among the…
We derive theorems which outline explicit mechanisms by which anomalous scaling for the probability density function of the sum of many correlated random variables asymptotically prevails. The results characterize general anomalous scaling…
Clustering provides a common means of identifying structure in complex data, and there is renewed interest in clustering as a tool for the analysis of large data sets in many fields. A natural question is how many clusters are appropriate…
We address the problem of giving robust performance bounds based on the study of the asymptotic behavior of the insensitive load balancing schemes when the number of servers and the load scales jointly. These schemes have the desirable…
Functional concurrent, or varying-coefficient, regression models are commonly used in biomedical and clinical settings to investigate how the relation between an outcome and observed covariate varies as a function of another covariate. In…
In this paper we propagate a large deviations approach for proving limit theory for (generally) multivariate time series with heavy tails. We make this notion precise by introducing regularly varying time series. We provide general large…
We define the local empirical process, based on $n$ i.i.d. random vectors in dimension $d$, in the neighborhood of the boundary of a fixed set. Under natural conditions on the shrinking neighborhood, we show that, for these local empirical…
We re-consider Leadbetter's extremal index for stationary sequences. It has interpretation as reciprocal of the expected size of an extremal cluster above high thresholds. We focus on heavy-tailed time series, in particular on regularly…
We study the problem of explainability-first clustering where explainability becomes a first-class citizen for clustering. Previous clustering approaches use decision trees for explanation, but only after the clustering is completed. In…
Finding "true" clusters in a data set is a challenging problem. Clustering solutions obtained using different models and algorithms do not necessarily provide compact and well-separated clusters or the optimal number of clusters. Cluster…
Random walks in random scenery are processes defined by $Z_n:=\sum_{k=1}^n\xi_{X_1+...+X_k}$, where $(X_k,k\ge 1)$ and $(\xi_y,y\in{\mathbb Z}^d)$ are two independent sequences of i.i.d. random variables with values in ${\mathbb Z}^d$ and…
In the analysis of binary longitudinal data, it is of interest to model a dynamic relationship between a response and covariates as a function of time, while also investigating similar patterns of time-dependent interactions. We present a…