Related papers: Concentration inequalities and cut-off phenomena f…
Artificial intelligence models trained through loss minimization have demonstrated significant success, grounded in principles from fields like information theory and statistical physics. This work explores these established connections…
Discovery of an accurate causal Bayesian network structure from observational data can be useful in many areas of science. Often the discoveries are made under uncertainty, which can be expressed as probabilities. To guide the use of such…
I am presenting a first-ever scientific collection of short sayings on probability and statistics expressed by most various men of science, many classics included, from antiquity to Kepler to our time. Quite understandably, the reader will…
Meta-analysis is an important tool for combining results from multiple studies and has been widely used in evidence-based medicine for several decades. This paper reports, for the first time, an interesting and valuable paradox in…
The aim of this paper is twofold. First, three theoretical principles are formalized: randomization, overrepresentation and restriction. We develop these principles and give a rationale for their use in choosing the sampling design in a…
The Pythagorean Expected Wins Percentage Model was developed by Bill James to estimate a baseball team expected wins percentage over the course of a season. As such, the model can be used to assess how lucky or unfortunate a team was over…
Aggregated predictors are obtained by making a set of basic predictors vote according to some weights, that is, to some probability distribution. Randomized predictors are obtained by sampling in a set of basic predictors, according to some…
For any distribution $\pi$ with support equal to $[n] = \{1, 2,..., n \}$, we study the set $\mathcal{A}_{\pi}$ of tridiagonal stochastic matrices $K$ satisfying $\pi(i) K[i,j] = \pi(j) K[j,i]$ for all $i, j \in [n]$. These matrices…
Random partition models are widely used in Bayesian methods for various clustering tasks, such as mixture models, topic models, and community detection problems. While the number of clusters induced by random partition models has been…
We investigate a quadratic dynamical system known as nonlinear recombinations. This system models the evolution of a probability measure over the Boolean cube, converging to the stationary state obtained as the product of the initial…
Random matrices now play a role in many parts of computational mathematics. To advance these applications, it is desirable to have tools that are flexible, easy to use, and powerful. Over the last 25 years, researchers have developed a…
In this paper we give optimal constants in Talagrand's concentration inequalities for maxima of empirical processes associated to independent and eventually nonidentically distributed random variables. Our approach is based on the entropy…
Random matrices acting on structured sets play a fundamental role in high-dimensional geometry, compressed sensing, and randomized algorithms. Existing results primarily focus on subgaussian models, when random matrices act as…
We consider a random walker whose motion is tethered around a focal point. We use two models that exhibit the same spatial dependence in the steady state but widely different dynamics. In one case, the walker is subject to a deterministic…
The problem of studying rare events is central to many areas of computer simulations. In a recent paper [Kang, P., et al., Nat. Comput. Sci. 4, 451-460, 2024], we have shown that a powerful way of solving this problem passes through the…
Ranking and comparing items is crucial for collecting information about preferences in many areas, from marketing to politics. The Mallows rank model is among the most successful approaches to analyse rank data, but its computational…
Chance-constrained programming (CCP) is one of the most difficult classes of optimization problems that has attracted the attention of researchers since the 1950s. In this survey, we focus on cases when only a limited information on the…
Concentration of measure is a phenomenon in which a random variable that depends in a smooth way on a large number of independent random variables is essentially constant. The random variable will "concentrate" around its median or…
Clustered observations are ubiquitous in controlled and observational studies and arise naturally in multi-centre trials or longitudinal surveys. We present a novel model for the analysis of clustered observations where the marginal…
The paper that is commented by Touchette contains a computational study which opens the door to a desirable generalization of the standard large deviation theory (applicable to a set of $N$ nearly independent random variables) to systems…