Related papers: A Jarque-Bera test for skew normal data
Class distribution skews in imbalanced datasets may lead to models with prediction bias towards majority classes, making fair assessment of classifiers a challenging task. Metrics such as Balanced Accuracy are commonly used to evaluate a…
Zipf's law is just one out of many universal laws proposed to describe statistical regularities in language. Here we review and critically discuss how these laws can be statistically interpreted, fitted, and tested (falsified). The modern…
The notion of maximal-spacing in several dimensions was introduced and studied by Deheuvels (1983) for data uniformly distributed on the unit cube. Later on, Janson (1987) extended the results to data uniformly distributed on any bounded…
We introduce a new class of algorithms, Stochastic Generalized Method of Moments (SGMM), for estimation and inference on (overidentified) moment restriction models. Our SGMM is a novel stochastic approximation alternative to the popular…
Fully test-time adaptation aims at adapting a pre-trained model to the test stream during real-time inference, which is urgently required when the test distribution differs from the training distribution. Several efforts have been devoted…
Individual gamma ray bursts (GRBs) have very diverse time behavior - from a single pulse to a long complex sequence of chaotic pulses of different timescales. I studied light curves of GRBs using data from the CGRO's BATSE experiment and…
We propose and study a general method for construction of consistent statistical tests on the basis of possibly indirect, corrupted, or partially available observations. The class of tests devised in the paper contains Neyman's smooth…
The log-normal distribution is one of the most common distributions used for modeling skewed and positive data. It frequently arises in many disciplines of science, specially in the biological and medical sciences. The statistical analysis…
We study the error of the number of points of a unimodular lattice that fall in a strictly convex and analytic set having the origin and that is dilated by a factor $t$. The aim is to generalize the result of a previous article. We first…
Laplace distribution is popular in the field of economics and finance. Still, data sets often show a lack of symmetry and a tendency of being bounded from either side of their support. In view of this, we introduce a new family of skew…
Under reasonable algebraic assumptions and under an infinite second order moment assumption, we show that the logarithm of the norm (log-norm) of a product of random i.i.d. matrices with entries in $\mathbb{R}$ or in any other local field…
Recent advances in Natural Language Processing have demonstrated the effectiveness of pretrained language models like BERT for a variety of downstream tasks. We present GiusBERTo, the first BERT-based model specialized for anonymizing…
Anomaly detection is a fundamental yet challenging problem in machine learning due to the lack of label information. In this work, we propose a novel and powerful framework, dubbed as SLA$^2$P, for unsupervised anomaly detection. After…
Let y=A\beta+\epsilon, where y is an N\times1 vector of observations, \beta is a p\times1 vector of unknown regression coefficients, A is an N\times p design matrix and \epsilon is a spherically symmetric error term with unknown scale…
We propose skewed stable random projections for approximating the pth frequency moments of dynamic data streams (0<p<=2), which has been frequently studied in theoretical computer science and database communities. Our method significantly…
This note corrects a technical error in Guardiola (2020, Journal of Statistical Distributions and Applications), presents updated derivations, and offers an extended discussion of the properties of the spherical Dirichlet distribution.…
The continuous random energy model (CREM) is a toy model of disordered systems introduced by Bovier and Kurkova in 2004 based on previous work by Derrida and Spohn in the 80s. In a recent paper by Addario-Berry and Maillard, they raised the…
We propose a general maximum likelihood empirical Bayes (GMLEB) method for the estimation of a mean vector based on observations with i.i.d. normal errors. We prove that under mild moment conditions on the unknown means, the average mean…
There are many information and divergence measures exist in the literature on information theory and statistics. The most famous among them are Kullback-Leiber's (1951)relative information and Jeffreys (1946) J-divergence, Information…
The Gutenberg-Richter power law distribution of earthquake sizes is one of the most famous example illustrating self-similarity. It is well-known that the Gutenberg-Richter distribution has to be modified for large seismic moments, due to…