English
Related papers

Related papers: Higher criticism: $p$-values and criticism

200 papers

Public data repositories have enabled researchers to compare results across multiple genomic studies in order to replicate findings. A common approach is to first rank genes according to an hypothesis of interest within each study. Then,…

Applications · Statistics 2012-06-29 Loki Natarajan , Minya Pu , Karen Messer

We consider the approximation of a convolution of possibly different probability measures by (compound) Poisson distributions and also by related signed measures of higher order. We present new total variation bounds having a better…

Probability · Mathematics 2017-03-08 Bero Roos

While P-values are widely abused, they are a useful tool for many purposes; banning them is analogous to banning scalpels because most people do not know how to perform surgery. Many reported P-values are not genuine P-values, for a variety…

Other Statistics · Statistics 2022-04-19 Philip B. Stark

Most NLP datasets are not annotated with protected attributes such as gender, making it difficult to measure classification bias using standard measures of fairness (e.g., equal opportunity). However, manually annotating a large dataset…

Computation and Language · Computer Science 2020-04-28 Kawin Ethayarajh

J. E. Hirsch (2005) introduced the h-index to quantify an individual's scientific research output by the largest number h of a scientist's papers, that received at least h citations. This so-called Hirsch index can be easily modified to…

Physics and Society · Physics 2013-01-31 Michael Schreiber

Online reviews and recommendation systems help users navigate overwhelming choice, but they are vulnerable to self-reinforcing distortions. This paper examines how a single malicious reviewer can exploit popularity-biased rating dynamics…

Social and Information Networks · Computer Science 2026-04-16 Itsuki Fujisaki , Kunhao Yang

Analysis of credibility is a reverse-Bayes technique that has been proposed by Matthews (2001) to overcome some of the shortcomings of significance tests. A significant result is deemed credible if current knowledge about the effect size is…

Methodology · Statistics 2017-12-11 Leonhard Held

We study the properties of several likelihood-based statistics commonly used in testing for the presence of a known signal under a mixture model with known background, but unknown signal fraction. Under the null hypothesis of no signal, all…

Data Analysis, Statistics and Probability · Physics 2018-12-26 Igor Volobouev , A. Alexandre Trindade

Kernel density estimation is a well known method involving a smoothing parameter (the bandwidth) that needs to be tuned by the user. Although this method has been widely used the bandwidth selection remains a challenging issue in terms of…

Statistics Theory · Mathematics 2019-02-05 Suzanne Varet , Claire Lacour , Pascal Massart , Vincent Rivoirard

A fundamental problem in statistics is to compare the outcomes attained by members of subpopulations. This problem arises in the analysis of randomized controlled trials, in the analysis of A/B tests, and in the assessment of fairness and…

Methodology · Statistics 2021-12-02 Mark Tygert

We use the replica method of statistical mechanics to examine a typical performance of correctly reconstructing $N$-dimensional sparse vector $bx=(x_i)$ from its linear transformation $by=bF bx$ of $P$ dimensions on the basis of…

Information Theory · Computer Science 2010-06-03 Yoshiyuki Kabashima , Tadashi Wadayama , Toshiyuki Tanaka

The Hirsch index or h-index is widely used to quantify the impact of an individual's scientific research output, determining the highest number h of a scientist's papers that received at least h citations. Several variants of the index have…

Physics and Society · Physics 2015-05-19 Michael Schreiber

Peer review and citation metrics are two means of gauging the value of scientific research, but the lack of publicly available peer review data makes the comparison of these methods difficult. Mathematics can serve as a useful laboratory…

Digital Libraries · Computer Science 2020-12-17 Lawrence Smolinsky , Daniel S. Sage , Aaron J. Lercher , Aaron Cao

The availability of large datasets requires an improved view on statistical laws in complex systems, such as Zipf's law of word frequencies, the Gutenberg-Richter law of earthquake magnitudes, or scale-free degree distribution in networks.…

Data Analysis, Statistics and Probability · Physics 2019-04-30 Martin Gerlach , Eduardo G. Altmann

The issue of combining individual $p$-values to aggregate multiple small effects is prevalent in many scientific investigations and is a long-standing statistical topic. Many classical methods are designed for combining independent and…

Methodology · Statistics 2021-09-08 Yusi Fang , George C. Tseng , Chung Chang

In many important statistical analyses, the number of covariates $p$ often exceeds the data size $n$, a regime commonly referred to as high-dimensional. While considerable progress has been made in high-dimensional regression under the…

Methodology · Statistics 2026-05-29 Herman Tesso , Georges Nguefack-Tsague

Inference of physical parameters from reference data is a well studied problem with many intricacies (inconsistent sets of data due to experimental systematic errors, approximate physical models...). The complexity is further increased when…

Data Analysis, Statistics and Probability · Physics 2017-09-06 Pascal Pernot , Fabien Cailliez

We consider Gaussian mixture models in high dimensions and concentrate on the twin tasks of detection and feature selection. Under sparsity assumptions on the difference in means, we derive information bounds and establish the performance…

Statistics Theory · Mathematics 2016-10-04 Nicolas Verzelen , Ery Arias-Castro

We present the expected values from p-value hacking as a choice of the minimum p-value among $m$ independents tests, which can be considerably lower than the "true" p-value, even with a single trial, owing to the extreme skewness of the…

Applications · Statistics 2018-01-29 Nassim Nicholas Taleb

Background Most methods of adjusting for multiplicity focus primarily on controlling type I errors and rarely consider type II errors. We propose a new method that considers controlling for false-positive findings while ensuring sufficient…

Applications · Statistics 2025-07-31 Jiale Li , Zimu Wei