English
Related papers

Related papers: A Generalized Publication Bias Model

200 papers

Heuristic negative sampling enhances recommendation performance by selecting negative samples of varying hardness levels from predefined candidate pools to guide the model toward learning more accurate decision boundaries. However, our…

Machine Learning · Computer Science 2025-08-12 Chu Zhao , Eneng Yang , Yizhou Dang , Jianzhe Zhao , Guibing Guo , Xingwei Wang

Shrinkage estimation is a fundamental tool of modern statistics, pioneered by Charles Stein upon his discovery of the famous paradox involving the multivariate Gaussian. A large portion of the subsequent literature only considers the…

Statistics Theory · Mathematics 2022-03-30 Max Fathi , Larry Goldstein , Gesine Reinert , Adrien Saumard

Many studies that gather social network data use survey methods that lead to censored, missing or otherwise incomplete information. For example, the popular fixed rank nomination (FRN) scheme, often used in studies of schools and…

Methodology · Statistics 2012-12-27 Peter Hoff , Bailey Fosdick , Alex Volfovsky , Katherine Stovel

The false discovery rate (FDR) and false nondiscovery rate (FNDR) have received considerable attention in the literature on multiple testing. These performance measures are also appropriate for classification, and in this work we develop…

Statistics Theory · Mathematics 2009-01-28 Clayton Scott , Gowtham Bellala , Rebecca Willett

In the setting of entangled single-sample distributions, the goal is to estimate some common parameter shared by a family of $n$ distributions, given one single sample from each distribution. This paper studies mean estimation for entangled…

Machine Learning · Computer Science 2020-07-14 Yingyu Liang , Hui Yuan

In selective classification (SC), a classifier abstains from making predictions that are likely to be wrong to avoid excessive errors. To deploy imperfect classifiers -- either due to intrinsic statistical noise of data or for robustness…

Machine Learning · Computer Science 2024-11-28 Hengyue Liang , Le Peng , Ju Sun

We study the behavior of the posterior distribution in high-dimensional Bayesian Gaussian linear regression models having $p\gg n$, with $p$ the number of predictors and $n$ the sample size. Our focus is on obtaining quantitative finite…

Statistics Theory · Mathematics 2014-01-06 Nate Strawn , Artin Armagan , Rayan Saab , Lawrence Carin , David Dunson

In statistical inference, uncertainty is unknown and all models are wrong. That is to say, a person who makes a statistical model and a prior distribution is simultaneously aware that both are fictional candidates. To study such cases,…

Machine Learning · Computer Science 2023-02-13 Sumio Watanabe

Background: Publication bias is the failure to publish the results of a study based on the direction or strength of the study findings. The existence of publication bias is firmly established in areas like medical research. Recent research…

Software Engineering · Computer Science 2021-06-24 Rolando P. Reyes , Óscar Dieste , Efraín R. Fonseca C. , Natalia Juristo

We consider a multiplicative deconvolution problem, in which the density $f$ or the survival function $S^X$ of a strictly positive random variable $X$ is estimated nonparametrically based on an i.i.d. sample from a noisy observation $Y =…

Statistics Theory · Mathematics 2025-09-30 Sergio Brenner Miguel , Jan Johannes , Maximilian Siebel

In 2015 the Open Science Collaboration (OSC) (Nosek et al 2015) published a highly influential paper which claimed that a large fraction of published results in the psychological sciences were not reproducible. In this article we review…

Applications · Statistics 2026-02-18 Anthony Almudevar , Jacob Almudevar

Publication bias and p-hacking are two well-known phenomena that strongly affect the scientific literature and cause severe problems in meta-analyses. Due to these phenomena, the assumptions of meta-analyses are seriously violated and the…

Methodology · Statistics 2020-02-26 Jonas Moss , Riccardo De Bin

This paper studies the fixed budget formulation of the Ranking and Selection (R&S) problem with independent normal samples, where the goal is to investigate different algorithms' convergence rate in terms of their resulting probability of…

Optimization and Control · Mathematics 2018-11-30 Di Wu , Enlu Zhou

Network meta-analysis (NMA) is a useful tool to compare multiple interventions simultaneously in a single meta-analysis, it can be very helpful for medical decision making when the study aims to find the best therapy among several active…

Methodology · Statistics 2024-02-02 Ao Huang , Yi Zhou , Satoshi Hattori

An established and growing literature on generalized fiducial inference and related fiducial ideas points to the adoption of fiducial inference as a mainstream perspective among modern statisticians. Like Bayesian posteriors, generalized…

Statistics Theory · Mathematics 2026-03-03 J. E. Borgert , Jan Hannig

In recent years the ultrahigh dimensional linear regression problem has attracted enormous attentions from the research community. Under the sparsity assumption most of the published work is devoted to the selection and estimation of the…

Methodology · Statistics 2013-05-01 Randy C. S. Lai , Jan Hannig , Thomas C. M. Lee

Tolerance limits have received considerable attention in the statistical literature, with applications reaching far beyond their initial role in quality control. The well-known formula of Scheff\'e and Tukey (1944) establishes a simple,…

Methodology · Statistics 2026-05-20 James H. McVittie , Martin Lysy , Masoud Asgharian

There is recent interest in estimating the false discovery rate (FDR) with published p-values. However, there is little formal research that addresses the manner and extent to which the presumed selection, or publication, bias model impacts…

Methodology · Statistics 2026-03-03 Tianyu Cao , Sangyoon Yi , Joshua Habiger

We introduce a novel framework for uncertainty quantification in clustering that combines martingale posterior distributions with density-based clustering. Unlike classical model-based approaches, which define clusters at the latent level…

Machine Learning · Statistics 2026-04-20 Nicola Bariletto , Stephen G. Walker

Machine learning systems increasingly face requirements to forget not only individual data points, but entire domains of information, such as toxic language, copyrighted corpora, or demographic biases. This raises a fundamental dilemma of…

Statistics Theory · Mathematics 2026-05-19 Aaradhya Pandey , Sanjeev Kulkarni