Related papers: Large deviations analysis for random combinatorial…
Cluster analysis requires many decisions: the clustering method and the implied reference model, the number of clusters and, often, several hyper-parameters and algorithms' tunings. In practice, one produces several partitions, and a final…
We study spatial permutations with cycle weights that are bounded or slowly diverging. We show that a phase transition occurs at an explicit critical density. The long cycles are macroscopic and their cycle lengths satisfy a…
We derive a parallel sampling algorithm for computational inverse problems that present an unknown linear forcing term and a vector of nonlinear parameters to be recovered. It is assumed that the data is noisy and that the linear part of…
We introduce two notions of discrepancy between binary vectors, which are not metric functions in general but nonetheless capture the mathematical structure of the binary asymmetric channel. In turn, these lead to two new fundamental…
Contrastive loss is a powerful approach for representation learning, where larger batch sizes enhance performance by providing more negative samples to better distinguish between similar and dissimilar data. However, scaling batch sizes is…
Model-based clustering of moderate or large dimensional data is notoriously difficult. We propose a model for simultaneous dimensionality reduction and clustering by assuming a mixture model for a set of latent scores, which are then linked…
In the regularly varying time series setting, a cluster of exceedances is a short period for which the supremum norm exceeds a high threshold. We propose to study a generalization of this notion considering short periods, or blocks, with…
For Laplacian models in dimension $(1+1)$ we derive sample path large deviations for the profile height function, that is, we study scaling limits of Gaussian integrated random walks and Gaussian integrated random walk bridges perturbed by…
Distributed Lag Models (DLMs) and similar regression approaches such as MIDAS have been used for many decades in econometrics and more recently to investigate how poor air quality adversely affects human health. In this paper we describe…
Consider a queueing system fed by traffic from $N$ independent and identically distributed marked point processes. We establish several novel sample path large deviations results in the scaled uniform topology for such a system with a small…
Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep epistemic divergence. Using two major reasoning benchmarks -…
Determinantal Point Processes (DPPs) are a family of probabilistic models that have a repulsive behavior, and lend themselves naturally to many tasks in machine learning where returning a diverse set of objects is important. While there are…
Recently we proposed a microscopic approach to the description of the phase behaviour and critical phenomena in binary fluid mixtures. It was based on the method of collective variables (CV) with a reference system. The approach allowed us…
Subsampling algorithms for various parametric regression models with massive data have been extensively investigated in recent years. However, all existing studies on subsampling heavily rely on clean massive data. In practical…
In a recent paper published in this journal [J. Phys. A: Math. Theor. 42 (2009) 495004] we studied a one-dimensional particles system where nearest particles attract with a force inversely proportional to a power \alpha of their distance…
Consider the problem where a statistician in a two-node system receives rate-limited information from a transmitter about marginal observations of a memoryless process generated from two possible distributions. Using its own observations,…
We consider a system of classical particles confined in a box $\Lambda\subset\mathbb{R}^d$ with zero boundary conditions interacting via a stable and regular pair potential. Based on the validity of the cluster expansion for the canonical…
In long-range percolation on $\mathbb{Z}^d$, points $x$ and $y$ are connected by an edge with probability $1-\exp(-\beta\|x-y\|^{-d-\alpha})$, where $\alpha>0$ is fixed and $\beta \geq 0$ is a parameter. As $d$ and $\alpha$ vary, the model…
Large-deviations theory deals with tails of probability distributions and the rare events of random processes, for example spreading packets of particles. Mathematically, it concerns the exponential fall-of of the density of thin-tailed…
Deep Learning (DL) methods show very good performance when trained on large, balanced data sets. However, many practical problems involve imbalanced data sets, or/and classes with a small number of training samples. The performance of DL…