English
Related papers

Related papers: A Generalized Publication Bias Model

200 papers

One of the classic ways to measure the success of a scientific facility is the publication return, which is defined as the number of refereed papers produced per unit of allocated resources (for example, telescope time or proposals). The…

Instrumentation and Methods for Astrophysics · Physics 2018-03-07 F. Patat , H. M. J. Boffin , D. Bordelon , U. Grothkopf , S. Meakins , S. Mieske , M. Rejkuba

Probability density functions (PDF) of statistical distributions of cluster sizes N, where N is the number of particles in the cluster, often seem to have less freedom than expected from considering the number of degrees of freedom at the…

Data Analysis, Statistics and Probability · Physics 2011-03-08 Sascha Vongehr , Shaochun Tang , Xiangkang Meng

Motivated by parametric models for which the likelihood is analytically unavailable, numerically unstable, or prohibitively expensive to compute or optimize, we develop a prior- and likelihood-free framework for fully probabilistic…

Methodology · Statistics 2026-03-17 Leonardo Cella , Emily C. Hector

There has been considerable interest in modelling the spread of information on X (formerly Twitter) using machine learning models. Here, we consider the problem of predicting the reposting of new information, i.e., when a user propagates…

Social and Information Networks · Computer Science 2026-01-19 Ziming Xu , Shi Zhou , Vasileios Lampos , Ingemar J. Cox

Bayesian Neural Networks (BNN) have emerged as a crucial approach for interpreting ML predictions. By sampling from the posterior distribution, data scientists may estimate the uncertainty of an inference. Unfortunately many inference…

Machine Learning · Computer Science 2023-11-23 Thomas D. Ahle , Sahar Karimi , Peter Tak Peter Tang

This paper presents a theoretical analysis of sample selection bias correction. The sample bias correction technique commonly used in machine learning consists of reweighting the cost of an error on each training point of a biased sample to…

Machine Learning · Computer Science 2008-12-18 Corinna Cortes , Mehryar Mohri , Michael Riley , Afshin Rostamizadeh

In 1957, Lindley published "A statistical paradox" in Biometrika, revealing a fundamental conflict between frequentist and Bayesian inference as sample size approaches infinity. We present a new paradox of a different kind: a conflict…

Methodology · Statistics 2025-12-01 Miodrag M. Lovric

We propose a novel method for selective classification (SC), a problem which allows a classifier to abstain from predicting some instances, thus trading off accuracy against coverage (the fraction of instances predicted). In contrast to…

Machine Learning · Computer Science 2021-10-26 Aditya Gangrade , Anil Kag , Venkatesh Saligrama

We study the distributions of the random Dirichlet series with parameters $(s, \beta)$ defined by $$ S=\sum_{n=1}^{\infty}\frac{I_n}{n^s}, $$ where $(I_n)$ is a sequence of independent Bernoulli random variables, $I_n$ taking value $1$ with…

Probability · Mathematics 2016-01-01 Ron Peled , Yuval Peres , Jim Pitman , Ryokichi Tanaka

Feedforward neural networks (FNNs) can be viewed as non-linear regression models, where covariates enter the model through a combination of weighted summations and non-linear functions. Although these models have some similarities to the…

Methodology · Statistics 2024-05-02 Andrew McInerney , Kevin Burke

Understanding the uncertainty of a neural network's (NN) predictions is essential for many purposes. The Bayesian framework provides a principled approach to this, however applying it to NNs is challenging due to large numbers of parameters…

Machine Learning · Statistics 2020-02-27 Tim Pearce , Felix Leibfried , Alexandra Brintrup , Mohamed Zaki , Andy Neely

The linear exponential distribution is a generalization of the exponential and Rayleigh distributions. This distribution is one of the best models to fit data with increasing failure rate (IFR). But it does not provide a reasonable fit for…

Statistics Theory · Mathematics 2021-07-21 M. Arshad , M. Khetan , V. Kumar , A. K. Pathak

Prior-data fitted networks (PFNs) have emerged as promising foundation models for prediction from tabular datasets, achieving state-of-the-art performance on small to moderate data sizes without tuning. While PFNs are motivated by Bayesian…

Methodology · Statistics 2026-05-11 Thomas Nagler , David Rügamer

Classification is a fundamental task in supervised learning, while achieving valid misclassification rate control remains challenging due to possibly the limited predictive capability of the classifiers or the intrinsic complexity of the…

Methodology · Statistics 2025-09-16 Yinrui Sun , Yin Xia

This paper discusses and analyzes a class of likelihood models which are based on two distributional innovations in financial models for stock returns. That is, the notion that the marginal distribution of aggregate returns of log-stock…

Statistics Theory · Mathematics 2007-06-13 Lancelot F. James , John W. Lau

A random balanced sample (RBS) is a multivariate distribution with n components X_1,...,X_n, each uniformly distributed on [-1, 1], such that the sum of these components is precisely 0. The corresponding vectors X lie in an…

Statistics Theory · Mathematics 2007-06-18 Peter Bubenik , John Holbrook

This paper develops a new framework for indirect statistical inference with guaranteed necessity and sufficiency, applicable to continuous random variables. We prove that when comparing exponentially transformed order statistics from an…

Statistics Theory · Mathematics 2025-09-25 Z Zhang , X Hu , C Lu , T Liu

The distribution of scientific citations for publications selected with different rules (author, topic, institution, country, journal, etc.) collapse on a single curve if one plots the citations relative to their mean value. We find that…

Physics and Society · Physics 2017-11-01 Zoltán Néda , Levente Varga , Tamás S. Biró

A major challenge in Semi-Supervised Learning (SSL) is the limited information available about the class distribution in the unlabeled data. In many real-world applications this arises from the prevalence of long-tailed distributions, where…

Machine Learning · Computer Science 2025-02-04 Khiem Pham , Charles Herrmann , Ramin Zabih

Consider a researcher estimating the parameters of a regression function based on data for all 50 states in the United States or on data for all visits to a website. What is the interpretation of the estimated parameters and the standard…

Statistics Theory · Mathematics 2019-06-25 Alberto Abadie , Susan Athey , Guido W. Imbens , Jeffrey M. Wooldridge