English
Related papers

Related papers: A Generalized Publication Bias Model

200 papers

We investigate theoretical guarantees for the false-negative rate (FNR) -- the fraction of true causal edges whose orientation is not recovered, under single-variable random interventions and an $\epsilon$-interventional faithfulness…

Machine Learning · Computer Science 2025-11-05 Mathieu Chevalley , Arash Mehrjou , Patrick Schwab

In research policy, effective measures that lead to improvements in the generation of knowledge must be based on reliable methods of research assessment, but for many countries and institutions this is not the case. Publication and citation…

Digital Libraries · Computer Science 2018-07-20 Alonso Rodriguez-Navarro , Ricardo Brito

In machine learning, it is commonly assumed that training and test data share the same population distribution. However, this assumption is often violated in practice because the sample selection bias may induce the distribution shift from…

Machine Learning · Computer Science 2020-06-09 Kun Kuang , Hengtao Zhang , Fei Wu , Yueting Zhuang , Aijun Zhang

The Open Science Collaboration recently reported that 36% of published findings from psychological studies were reproducible by independent researchers. We can use this information together with Bayes theorem to estimate the statistical…

Physics and Society · Physics 2016-09-13 Michael Ingre

At the present time reliably established that probability density functions of gene expression of microarray experiments possess a number of universal properties. First of all these distributions have power asymptotic and secondly the shape…

Statistics Theory · Mathematics 2015-06-08 Viacheslav Saenko , Yurij Saenko

Publication bias occurs when the publication of research results depends not only on the quality of the research but also on its nature and direction. The consequence is that published studies may not be truly representative of all valid…

Methodology · Statistics 2020-02-13 Chuan Hong , Jing Zhang , Yang Li , Elena Elia , Richard Riley , Yong Chen

Failure of machine learning models to generalize to new data is a core problem limiting the reliability of AI systems, partly due to the lack of simple and robust methods for comparing new data to the original training dataset. We propose a…

Machine Learning · Computer Science 2025-02-26 W. Max Schreyer , Christopher Anderson , Reid F. Thompson

Recommender systems leverage extensive user interaction data to model preferences; however, directly modeling these data may introduce biases that disproportionately favor popular items. In this paper, we demonstrate that popularity bias…

Information Retrieval · Computer Science 2025-04-21 Jiahao Liu , Dongsheng Li , Hansu Gu , Peng Zhang , Tun Lu , Li Shang , Ning Gu

This monograph presents a unified treatment of single- and multi-user problems in Shannon's information theory where we depart from the requirement that the error probability decays asymptotically in the blocklength. Instead, the error…

Information Theory · Computer Science 2015-04-13 Vincent Y. F. Tan

Garcia-Donato et al. (2025) present a methodology for handling missing data in a model selection problem using an objective Bayesian approach. The current comment discusses an alternative, existing objective Bayesian method for this…

Methodology · Statistics 2025-12-25 Joris Mulder

Progressive censoring scheme has received considerable attention in recent years. In this paper we introduce a new type-II progressive censoring scheme for two samples. It is observed that the proposed censoring scheme is analytically more…

Methodology · Statistics 2016-09-20 Shuvashree Mondal , Debasis Kundu

Residuals in normal regression are used to assess a model's goodness-of-fit (GOF) and discover directions for improving the model. However, there is a lack of residuals with a characterized reference distribution for censored regression. In…

Methodology · Statistics 2022-01-04 Longhai Li , Tingxuan Wu , Cindy Feng

Explicit finite-sample statistical guarantees on model performance are an important ingredient in responsible machine learning. Previous work has focused mainly on bounding either the expected loss of a predictor or the probability that an…

Machine Learning · Computer Science 2024-03-07 Zhun Deng , Thomas P. Zollo , Jake C. Snell , Toniann Pitassi , Richard Zemel

Within the machine learning community, the widely-used uniform convergence framework has been used to answer the question of how complex, over-parameterized models can generalize well to new data. This approach bounds the test error of the…

Machine Learning · Statistics 2021-03-05 Ryan Theisen , Jason M. Klusowski , Michael W. Mahoney

In meta-analyses, publication bias is a well-known, important and challenging issue because the validity of the results from a meta-analysis is threatened if the sample of studies retrieved for review is biased. One popular method to deal…

Methodology · Statistics 2020-07-03 Rui Duan , Jin Piao , Arielle Marks-Anglin , Jiayi Tong , Lifeng Lin , Haitao Chu , Jing Ning , Yong Chen

A leading explanation for widespread replication failures is publication bias. I show in a simple model of selective publication that, contrary to common perceptions, the replication rate is unaffected by the suppression of insignificant…

General Economics · Economics 2022-07-04 Patrick Vu

The progressive censoring scheme has received considerable amount of attention in the last fifteen years. During the last few years joint progressive censoring scheme has gained some popularity. Recently, the authors Mondal and Kundu ("A…

Applications · Statistics 2018-01-03 Shuvashree Mondal , Debasis Kundu

In the Gaussian linear regression model (with unknown mean and variance), we show that the standard confidence set for one or two regression coefficients is admissible in the sense of Joshi (1969). This solves a long-standing open problem…

Statistics Theory · Mathematics 2018-09-25 Hannes Leeb , Paul Kabaila

Building on the recent development of the model-free generalized fiducial (MFGF) paradigm (Williams, 2023) for predictive inference with finite-sample frequentist validity guarantees, in this paper, we develop an MFGF-based approach to…

Statistics Theory · Mathematics 2024-05-20 Jonathan P Williams , Yang Liu

Federated Semi-Supervised Learning (FSSL) leverages both labeled and unlabeled data on clients to collaboratively train a model.In FSSL, the heterogeneous data can introduce prediction bias into the model, causing the model's prediction to…

Machine Learning · Computer Science 2024-05-31 Guogang Zhu , Xuefeng Liu , Xinghao Wu , Shaojie Tang , Chao Tang , Jianwei Niu , Hao Su