English
Related papers

Related papers: When is p-hacking detectable?

200 papers

The Fr\'echet distance is a popular distance measure for curves which naturally lends itself to fundamental computational tasks, such as clustering, nearest-neighbor searching, and spherical range searching in the corresponding metric…

Computational Geometry · Computer Science 2018-08-07 Anne Driemel , Amer Krivošija

Phenomenological (P-type) bifurcations are qualitative changes in stochastic dynamical systems whereby the stationary probability density function (PDF) changes its topology. The current state of the art for detecting these bifurcations…

Algebraic Topology · Mathematics 2024-06-10 Sunia Tanweer , Firas A. Khasawneh

This paper investigates the fundamental limits for detecting a high-dimensional sparse matrix contaminated by white Gaussian noise from both the statistical and computational perspectives. We consider $p\times p$ matrices whose rows and…

Statistics Theory · Mathematics 2018-01-03 T. Tony Cai , Yihong Wu

With the development of information technology and the Internet, recommendation systems have become an important means to solve the problem of information overload. However, recommendation system is greatly fragile as it relies heavily on…

Cryptography and Security · Computer Science 2019-08-21 Wanqiao Yuan , Yingyuan Xiao , Xu Jiao , Wenguang Zheng , Zihao Ling

Designed gene expression micro-array experiments, consisting of several treatment levels with a number of replicates per level, are analyzed by applying simple tests for group differences at the per gene level. The gene level statistics are…

Methodology · Statistics 2017-12-11 Grant Izmirlian

Scatterplots are the most common way for statisticians, scientists, and the public to visually detect relationships between measured variables. At the same time, and despite widely publicized controversy, P-values remain the most commonly…

Applications · Statistics 2014-04-23 Aaron Fisher , G. Brooke Anderson , Roger Peng , Jeff Leek

In contrast to its common definition and calculation, interpretation of p-values diverges among statisticians. Since p-value is the basis of various methodologies, this divergence has led to a variety of test methodologies and evaluations…

Methodology · Statistics 2012-12-27 Tomokazu Konishi

This paper clarifies a fundamental difference between causal inference and traditional statistical inference by formalizing a mathematical distinction between their respective parameters. We connect two major approaches to causal inference,…

Methodology · Statistics 2025-08-29 Muye Liu , Jun Xie

When data are collected subject to a detection limit, observations below the detection limit may be considered censored. In addition, the domain of such observations may be restricted; for example, values may be required to be non-negative.…

Applications · Statistics 2020-06-30 Justin R. Williams , Hyung-Woo Kim , Catherine M. Crespi

Traditional fact checking by expert journalists cannot keep up with the enormous volume of information that is now generated online. Computational fact checking may significantly enhance our ability to evaluate the veracity of dubious…

Computers and Society · Computer Science 2020-07-01 Giovanni Luca Ciampaglia , Prashant Shiralkar , Luis M. Rocha , Johan Bollen , Filippo Menczer , Alessandro Flammini

This paper studies the construction of p-values for nonparametric outlier detection, taking a multiple-testing perspective. The goal is to test whether new independent samples belong to the same distribution as a reference data set or are…

Methodology · Statistics 2024-03-12 Stephen Bates , Emmanuel Candès , Lihua Lei , Yaniv Romano , Matteo Sesia

Good (Frequentist) statistical practice requires that statistical tests be performed in order to determine if the phenomenon being observed could plausibly occur by chance if the null hypothesis is false. Good practice also requires that a…

Computers and Society · Computer Science 2024-08-06 Michael Guerzhoy

This paper develops a framework to study the statistical power of revealed-preference tests. With randomly sampled budgets and mild smoothness of demand, statistical learning implies that any model consistent with the data must approximate…

Theoretical Economics · Economics 2026-02-12 Charles Gauthier , Raghav Malhotra , Agustin Troccoli Moretti

Selective inference is a subfield of statistics that enables valid inference after selection of a data-dependent question. In this paper, we introduce selectively dominant p-values, a class of p-values that allow practitioners to easily…

Methodology · Statistics 2024-11-22 Anav Sood

Increasing accessibility of data to researchers makes it possible to conduct massive amounts of statistical testing. Rather than follow a carefully crafted set of scientific hypotheses with statistical analysis, researchers can now test…

Genomics · Quantitative Biology 2016-09-08 Olga A. Vsevolozhskaya , Chia-Ling Kuo , Gabriel Ruiz , Luda Diatchenko , Dmitri V. Zaykin

Pearson's chi-squared test can detect outliers in the data distribution of a given set of histograms. However, in fields such as demographics (for e.g. birth years), outliers may be more easily found in terms of the histogram smoothness…

Methodology · Statistics 2020-12-22 Nikola Banić , Neven Elezović

The prediction accuracy of machine learning methods is steadily increasing, but the calibration of their uncertainty predictions poses a significant challenge. Numerous works focus on obtaining well-calibrated predictive models, but less is…

Machine Learning · Statistics 2023-12-07 Donghwan Lee , Xinmeng Huang , Hamed Hassani , Edgar Dobriban

Scientific fraud is an increasingly vexing problem. Many current programs for fraud detection focus on image manipulation, while techniques for detection based on anomalous patterns that may be discoverable in the underlying numerical data…

Quantitative Methods · Quantitative Biology 2013-11-22 Joel H Pitt , Helene Z Hill

Statistical inference of directed relations given some unspecified interventions (i.e., the intervention targets are unknown) is challenging. In this article, we test hypothesized directed relations with unspecified interventions. First, we…

Methodology · Statistics 2023-03-02 Chunlin Li , Xiaotong Shen , Wei Pan

The mid-p-value is a proposed improvement on the ordinary p-value for the case where the test statistic is partially or completely discrete. In this case, the ordinary p-value is conservative, meaning that its null distribution is larger…

Statistics Theory · Mathematics 2017-06-02 Patrick Rubin-Delanchy , Nicholas A. Heard , Daniel John Lawson