Related papers: Splitting Localization and Prediction Numbers
Though mostly used as a clustering algorithm, k-means are originally designed as a quantization algorithm. Namely, it aims at providing a compression of a probability distribution with k points. Building upon [21, 33], we try to investigate…
An exponential-time exact algorithm is provided for the task of clustering n items of data into k clusters. Instead of seeking one partition, posterior probabilities are computed for summary statistics: the number of clusters, and pairwise…
We give a simple and elementary proof of the identity $$\sum_{r=1}^n\sum_{k_1,...,k_r\ge 1: \sum_{i=1}^r k_i= n} \frac {n!} {k_1!k_2!...k_r!}k_1^{k_2}...k_{r-1}^{k_r}=(n+1)^{n-1}$$ where $n\in \mathbb N$. A first application of this formula…
Calibration is a conditional property that depends on the information retained by a predictor. We develop decomposition identities for arbitrary proper losses that make this dependence explicit. At any information level $\mathcal A$, the…
The famous Lovasz Local Lemma [EL75] is a powerful tool to non-constructively prove the existence of combinatorial objects meeting a prescribed collection of criteria. Kratochvil et al. applied this technique to prove that a k-CNF in which…
This article deals with the spatio-temporal sensors deployment in order to maximize detection probability of an intelligent and randomly moving target in an area under surveillance. Our work is based on the rare events simulation framework.…
In the present paper we are interested in properties of forcing notions which measure in a sense the distance between the ground model reals and the reals in the extension. We look at the ways the ``new'' reals can be aproximated by ``old''…
The problem of estimating the multiplicity of the zero of a polynomial when restricted to the trajectory of a non-singular polynomial vector field, at one or several points, has been considered by authors in several different fields. The…
A robust clustering method for probabilities in Wasserstein space is introduced. This new "trimmed $k$-barycenters" approach relies on recent results on barycenters in Wasserstein space that allow intensive computation, as required by…
We find that a wide variety of families of partition statistics stabilize in a fashion similar to $p_k(n)$, the number of partitions of n with k parts, which satisfies $p_k(n) = p_{k+1}(n + 1), k \geq n/2$. We bound the regions of…
Spatial confounding between the spatial random effects and fixed effects covariates has been recently discovered and showed that it may bring misleading interpretation to the model results. Solutions to alleviate this problem are based on…
A permutation can be locally classified according to the four local types: peaks, valleys, double rises and double falls. The corresponding classification of binary increasing trees uses four different types of nodes. Flajolet demonstrated…
Let $L$ be the Dunkl Laplacian on the Euclidean space $\mathbb{R}^N$ associated with a normalized root system $R$ and a multiplicity function $k(\nu)\geq 0$, $\nu\in R$. We establish a Leibniz-type rule for the fractional powers of $L$ on…
Loss tomography has received considerable attention in recent years and a number of estimators based on maximum likelihood (ML) or Bayesian principles have been proposed. Almost all of the estimators are devoted to the tree topology despite…
We consider Bayesian linear regression with sparsity-inducing prior and design efficient sampling algorithms leveraging posterior contraction properties. A quasi-likelihood with Gaussian spike-and-slab (that is favorable both statistically…
We propose a new approach to safe variable preselection in high-dimensional penalized regression, such as the lasso. Preselection - to start with a manageable set of covariates - has often been implemented without clear appreciation of its…
This paper describes techniques for growing classification and regression trees designed to induce visually interpretable trees. This is achieved by penalizing splits that extend the subset of features used in a particular branch of the…
Eliciting preferences from human judgements is inherently imprecise, yet most decision analysis methods force a single priority vector from pairwise comparisons, discarding the information embedded in inconsistencies. We instead leverage…
In probabilistic program analysis, quantitative analysis aims at deriving tight numerical bounds for probabilistic properties such as expectation and assertion probability. Most previous works consider numerical bounds over the whole…
In these expository notes, we describe some features of the multiplicative coalescent and its connection with random graphs and minimum spanning trees. We use Pitman's proof of Cayley's formula, which proceeds via a calculation of the…