Related papers: Large deviations analysis for random combinatorial…
Classical peaks over threshold analysis is widely used for statistical modeling of sample extremes, and can be supplemented by a model for the sizes of clusters of exceedances. Under mild conditions a compound Poisson process model allows…
This article carries out a large dimensional analysis of standard regularized discriminant analysis classifiers designed on the assumption that data arise from a Gaussian mixture model with different means and covariances. The analysis…
In a recent study, the finite-time ($t$) and -population size ($N_c$) scalings in the evaluation of a large deviation function (LDF) estimator were analyzed by means of the cloning algorithm. These scalings provide valuable information…
We present large deviations estimates in the supremum norm for a system of independent random walks superposed with a birth-and-death dynamics evolving on the discrete torus with $N$ sites. The scaling limit considered is the so-called…
The Kullback-Leibler (KL) divergence is frequently used in data science. For discrete distributions on large state spaces, approximations of probability vectors may result in a few small negative entries, rendering the KL divergence…
The peak-background split argument is commonly used to relate the abundance of dark matter halos to their spatial clustering. Testing this argument requires an accurate determination of the halo mass function. We present a Maximum…
We study multifractality in a broad class of disordered systems which includes, e.g., the diluted x-y model. Using renormalized field theory we analyze the scaling behavior of cumulant averaged dynamical variables (in case of the x-y model…
In all local low-dimensional models, scaling at critical points deviates from mean field behavior -- with one possible exception. This exceptional model with ``ordinary" behavior is an inherently non-equilibrium model studied some time ago…
We propose a method for describing a phase behavior of a system consisting of particles of two sorts. The interaction of each species is described by interaction potentials containing the repulsive and attractive components. Asymmetry is…
The disorder-driven phase transition of the RFIM is observed using exact ground-state computer simulations for hyper cubic lattices in d=5,6,7 dimensions. Finite-size scaling analyses are used to calculate the critical point and the…
The ability to compute the exact divergence between two high-dimensional distributions is useful in many applications but doing so naively is intractable. Computing the alpha-beta divergence -- a family of divergences that includes the…
Large deviation functions contain information on the stability and response of systems driven into nonequilibrium steady states, and in such a way are similar to free energies for systems at equilibrium. As with equilibrium free energies,…
In the classical contamination models, such as the gross-error (Huber and Tukey contamination model or Case-wise Contamination), observations are considered as the units to be identified as outliers or not. This model is very useful when…
For Markov processes evolving on multiple time-scales a combination of large component scalings and averaging of rapid fluctuations can lead to useful limits for model approximation. A general approach to proving a law of large numbers to a…
The event of large losses plays an important role in credit risk. As these large losses are typically rare, and portfolios usually consist of a large number of positions, large deviation theory is the natural tool to analyze the tail…
The influence of size differences, shape, mass and persistent motion on phase separation in binary mixtures has been intensively studied. Here we focus on the exclusive role of diffusivity differences in binary mixtures of equal-sized…
We prove a large deviations principle for the empirical measure of the one dimensional symmetric simple exclusion process in contact with reservoirs. The dynamics of the reservoirs is slowed down with respect to the dynamics of the system,…
Symbolic regression that aims to detect underlying data-driven models has become increasingly important for industrial data analysis. For most existing algorithms such as genetic programming (GP), the convergence speed might be too slow for…
We discuss a Bayesian model selection approach to high dimensional data in the deep under sampling regime. The data is based on a representation of the possible discrete states $s$, as defined by the observer, and it consists of $M$…
Recent deep learning models are difficult to train using a large batch size, because commodity machines may not have enough memory to accommodate both the model and a large data batch size. The batch size is one of the hyper-parameters used…