English
Related papers

Related papers: Correction algorithm for finite sample statistics

200 papers

Message importance measure (MIM) is an important index to describe the message importance in the scenario of big data. Similar to the Shannon Entropy and Renyi Entropy, MIM is required to characterize the uncertainty of a random process and…

Information Theory · Computer Science 2016-07-07 Pingyi Fan , Yunquan Dong , Jiaxun Lu , Shanyun Liu

Compression of integer sets and sequences has been extensively studied for settings where elements follow a uniform probability distribution. In addition, methods exist that exploit clustering of elements in order to achieve higher…

Information Theory · Computer Science 2014-02-11 N. Jesper Larsson

This paper studies the distribution of a family of rankings, which includes Google's PageRank, on a directed configuration model. In particular, it is shown that the distribution of the rank of a randomly chosen node in the graph converges…

Probability · Mathematics 2014-10-14 Ningyuan Chen , Nelly Litvak , Mariana Olvera-Cravioto

In this paper, we use the framework of mod-$\phi$ convergence to prove precise large or moderate deviations for quite general sequences of real valued random variables $(X_{n})_{n \in \mathbb{N}}$, which can be lattice or non-lattice…

Probability · Mathematics 2017-02-14 Valentin Féray , Pierre-Loïc Méliot , Ashkan Nikeghbali

A two-parameter family of exchangeable partitions with a simple updating rule is introduced. The partition is identified with a randomized version of a standard symmetric Dirichlet species-sampling model with finitely many types. A…

Probability · Mathematics 2010-01-27 Alexander Gnedin

Random forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of…

Quantitative Methods · Quantitative Biology 2007-05-23 Ramon Diaz-Uriarte , Sara Alvarez de Andres

Entropy Estimation is an important problem with many applications in cryptography, statistic,machine learning. Although the estimators optimal with respect to the sample complexity have beenrecently developed, there are still some…

Data Structures and Algorithms · Computer Science 2020-02-24 Maciej Skorski

We introduce a new method to reconstruct the density matrix $\rho$ of a system of $n$-qubits and estimate its rank $d$ from data obtained by quantum state tomography measurements repeated $m$ times. The procedure consists in minimizing the…

Statistics Theory · Mathematics 2015-06-05 Pierre Alquier , Cristina Butucea , Mohamed Hebiri , Katia Meziani , Morimae Tomoyuki

Randomized rankings have been of recent interest to achieve ex-ante fairer exposure and better robustness than deterministic rankings. We propose a set of natural axioms for randomized group-fair rankings and prove that there exists a…

Machine Learning · Computer Science 2023-05-30 Sruthi Gorantla , Amit Deshpande , Anand Louis

We propose a compression-based version of the empirical entropy of a finite string over a finite alphabet. Whereas previously one considers the naked entropy of (possibly higher order) Markov processes, we consider the sum of the…

Information Theory · Computer Science 2011-04-05 Paul M. B. Vitányi

The two-sample problem, which consists in testing whether independent samples on $\mathbb{R}^d$ are drawn from the same (unknown) distribution, finds applications in many areas. Its study in high-dimension is the subject of much attention,…

Statistics Theory · Mathematics 2023-02-09 Stephan Clémençon , Myrto Limnios , Nicolas Vayatis

Real-world complex systems often comprise many distinct types of elements as well as many more types of networked interactions between elements. When the relative abundances of types can be measured well, we often observe heavy-tailed…

Physics and Society · Physics 2025-03-17 P. S. Dodds , J. R. Minot , M. V. Arnold , T. Alshaabi , J. L. Adams , A. J. Reagan , C. M. Danforth

Machine learning practitioners frequently observe tension between predictive accuracy and group fairness constraints -- yet sometimes fairness interventions appear to improve accuracy. We show that both phenomena can be artifacts of…

Machine Learning · Computer Science 2026-02-06 Amir Asiaee , Kaveh Aryan

We consider the problem of finding good low rank approximations of symmetric, positive-definite $A \in \mathbb{R}^{n \times n}$. Chen-Epperly-Tropp-Webber showed, among many other things, that the randomly pivoted partial Cholesky algorithm…

Numerical Analysis · Mathematics 2024-04-18 Stefan Steinerberger

Categorical responses arise naturally within various scientific disciplines. In many circumstances, there is no predetermined order for the response categories, and the response has to be modeled as nominal. In this study, we regard the…

Statistics Theory · Mathematics 2024-03-20 Tianmeng Wang , Jie Yang

Permutation tests are widely used for statistical hypothesis testing when the sampling distribution of the test statistic under the null hypothesis is analytically intractable or unreliable due to finite sample sizes. One critical challenge…

Computation · Statistics 2023-08-29 Yang Shi , Huining Kang , Ji-Hyun Lee , Hui Jiang

Several real-world and abstract structures and systems are characterized by marked hierarchy to the point of being expressed as trees. Because the study of these entities often involves sampling (or discovering) the tree nodes in a specific…

Physics and Society · Physics 2022-04-18 Alexandre Benatti , Luciano da F. Costa

Consider bivariate observations $(X_1,Y_1), \ldots, (X_n,Y_n) \in \mathbb{R}\times \mathbb{R}$ with unknown conditional distributions $Q_x$ of $Y$, given that $X = x$. The goal is to estimate these distributions under the sole assumption…

Statistics Theory · Mathematics 2025-01-31 Alexandre Mösching , Lutz Duembgen

Data analysis and machine learning have become an integrative part of the modern scientific methodology, offering automated procedures for the prediction of a phenomenon based on past observations, unraveling underlying patterns in data and…

Machine Learning · Statistics 2015-06-04 Gilles Louppe

In R\'enyi's representation for exponential order statistics, we replace the iid exponential sequence with any iid sequence, and call the resulting order statistic generalized R\'enyi statistic. We prove that by randomly reordering the…

Statistics Theory · Mathematics 2025-02-24 Péter Kevei , László Viharos