English
Related papers

Related papers: A Generalized Publication Bias Model

200 papers

In the double rank analysis of research publications, the local rank position of a country or institution publication is expressed as a function of the world rank position. Excluding some highly or lowly cited publications, the double rank…

Digital Libraries · Computer Science 2018-02-07 Ricardo Brito , Alonso Rodriguez-Navarro

Preferential sampling has attracted considerable attention in geostatistics since the pioneering work of Diggle et al. (2010). A variety of likelihood-based approaches have been developed to correct estimation bias by explicitly modelling…

Methodology · Statistics 2025-11-06 Changqing Lu , Ganggang Xu , Junho Yang , Yongtao Guan

Standard random-effects meta-analysis relies heavily on the assumption that the underlying true effects are normally distributed. In the social sciences, where evidence synthesis increasingly involves large, highly heterogeneous datasets,…

Methodology · Statistics 2026-05-01 Daihe Sui , Elizabeth Tipton

Over the past two decades, shrinkage priors have become increasingly popular, and many proposals can be found in the literature. These priors aim to shrink small effects to zero while maintaining true large effects. Horseshoe-type priors…

Statistics Theory · Mathematics 2025-01-14 Maria De Iorio , Andreas Heinecke , Beatrice Franzolini , Rafael Cabral

Traditional methods for linear regression generally assume that the underlying error distribution, equivalently the distribution of the responses, is normal. Yet, sometimes real life response data may exhibit a skewed pattern, and assuming…

Methodology · Statistics 2025-01-07 Amarnath Nandy , Ayanendranath Basu , Abhik Ghosh

Despite the strong predictive performance achieved by machine learning models across many application domains, assessing their trustworthiness through reliable estimates of predictive confidence remains a critical challenge. This issue…

Machine Learning · Computer Science 2026-03-25 Abolfazl Mohammadi-Seif , Carlos Soares , Rita P. Ribeiro , Ricardo Baeza-Yates

Standard Bayesian analyses can be difficult to perform when the full likelihood, and consequently the full posterior distribution, is too complex and difficult to specify or if robustness with respect to data or to model misspecifications…

Methodology · Statistics 2019-01-08 Federica Giummolè , Valentina Mameli , Erlis Ruli , Laura Ventura

This paper extends the theory of false discovery rates (FDR) pioneered by Benjamini and Hochberg [J. Roy. Statist. Soc. Ser. B 57 (1995) 289-300]. We develop a framework in which the False Discovery Proportion (FDP)--the number of false…

Statistics Theory · Mathematics 2007-06-13 Christopher Genovese , Larry Wasserman

When P indistinguishable balls are randomly distributed among L distinguishable boxes, and considering the dense system in which P much greater than L, our natural intuition tells us that the box with the average number of balls has the…

Physics and Society · Physics 2016-06-14 Oded Kafri

Since the National Academy of Sciences released their report outlining paths for improving reliability, standards, and policies in the forensic sciences NAS (2009), there has been heightened interest in evaluating and improving the…

In an earlier work we had considered a Gaussian ensemble of random matrices in the presence of a given external matrix source. The measure is no longer unitary invariant and the usual techniques based on orthogonal polynomials, or on the…

Statistical Mechanics · Physics 2009-10-31 E. Brezin , S. Hikami

Generalized likelihood ratio statistics have been proposed in Fan, Zhang and Zhang [Ann. Statist. 29 (2001) 153-193] as a generally applicable method for testing nonparametric hypotheses about nonparametric functions. The likelihood ratio…

Statistics Theory · Mathematics 2007-06-13 Jianqing Fan , Jian Zhang

Large-scale datasets are increasingly being used to inform decision making. While this effort aims to ground policy in real-world evidence, challenges have arisen as selection bias and other forms of distribution shifts often plague…

Methodology · Statistics 2023-11-07 Santiago Cortes-Gomez , Mateo Dulce , Carlos Patino , Bryan Wilder

The influential claim that most published results are false raised concerns about the trustworthiness and integrity of science. Since then, there have been numerous attempts to examine the rate of false-positive results that have failed to…

Applications · Statistics 2023-09-19 Ulrich Schimmack , František Bartoš

We find the value of constants related to constraints in characterization of some known statistical distributions and then we proceed to use the idea behind maximum entropy principle to derive generalized version of this distributions using…

Statistical Mechanics · Physics 2007-05-23 Oscar Sotolongo-Costa , Alejandro Gonzalez Gonzalez , Francois Brouers

A learned generative model often produces biased statistics relative to the underlying data distribution. A standard technique to correct this bias is importance sampling, where samples from the model are weighted by the likelihood ratio…

Machine Learning · Statistics 2019-11-05 Aditya Grover , Jiaming Song , Alekh Agarwal , Kenneth Tran , Ashish Kapoor , Eric Horvitz , Stefano Ermon

These lecture notes consist of three chapters. In the first chapter we present oracle inequalities for the prediction error of the Lasso and square-root Lasso and briefly describe the scaled Lasso. In the second chapter we establish…

Statistics Theory · Mathematics 2014-10-01 Sara van de Geer

Most NLP datasets are not annotated with protected attributes such as gender, making it difficult to measure classification bias using standard measures of fairness (e.g., equal opportunity). However, manually annotating a large dataset…

Computation and Language · Computer Science 2020-04-28 Kawin Ethayarajh

Machine learning systems increasingly face requirements to remove entire domains of information--such as toxic language or biases--rather than individual user data. This task presents a dilemma: full removal of the unwanted domain data is…

Machine Learning · Computer Science 2026-01-15 Youssef Allouah , Rachid Guerraoui , Sanmi Koyejo

Statistical dependence between hypotheses poses a significant challenge to the stability of large scale multiple hypotheses testing. Ignoring it often results in an unacceptably large spread in the false positive proportion even though the…

Methodology · Statistics 2018-10-15 Sairam Rayaprolu , Zhiyi Chi