Related papers: A Generalized Publication Bias Model
One of the classic ways to measure the success of a scientific facility is the publication return, which is defined as the number of refereed papers produced per unit of allocated resources (for example, telescope time or proposals). The…
Probability density functions (PDF) of statistical distributions of cluster sizes N, where N is the number of particles in the cluster, often seem to have less freedom than expected from considering the number of degrees of freedom at the…
Motivated by parametric models for which the likelihood is analytically unavailable, numerically unstable, or prohibitively expensive to compute or optimize, we develop a prior- and likelihood-free framework for fully probabilistic…
There has been considerable interest in modelling the spread of information on X (formerly Twitter) using machine learning models. Here, we consider the problem of predicting the reposting of new information, i.e., when a user propagates…
Bayesian Neural Networks (BNN) have emerged as a crucial approach for interpreting ML predictions. By sampling from the posterior distribution, data scientists may estimate the uncertainty of an inference. Unfortunately many inference…
This paper presents a theoretical analysis of sample selection bias correction. The sample bias correction technique commonly used in machine learning consists of reweighting the cost of an error on each training point of a biased sample to…
In 1957, Lindley published "A statistical paradox" in Biometrika, revealing a fundamental conflict between frequentist and Bayesian inference as sample size approaches infinity. We present a new paradox of a different kind: a conflict…
We propose a novel method for selective classification (SC), a problem which allows a classifier to abstain from predicting some instances, thus trading off accuracy against coverage (the fraction of instances predicted). In contrast to…
We study the distributions of the random Dirichlet series with parameters $(s, \beta)$ defined by $$ S=\sum_{n=1}^{\infty}\frac{I_n}{n^s}, $$ where $(I_n)$ is a sequence of independent Bernoulli random variables, $I_n$ taking value $1$ with…
Feedforward neural networks (FNNs) can be viewed as non-linear regression models, where covariates enter the model through a combination of weighted summations and non-linear functions. Although these models have some similarities to the…
Understanding the uncertainty of a neural network's (NN) predictions is essential for many purposes. The Bayesian framework provides a principled approach to this, however applying it to NNs is challenging due to large numbers of parameters…
The linear exponential distribution is a generalization of the exponential and Rayleigh distributions. This distribution is one of the best models to fit data with increasing failure rate (IFR). But it does not provide a reasonable fit for…
Prior-data fitted networks (PFNs) have emerged as promising foundation models for prediction from tabular datasets, achieving state-of-the-art performance on small to moderate data sizes without tuning. While PFNs are motivated by Bayesian…
Classification is a fundamental task in supervised learning, while achieving valid misclassification rate control remains challenging due to possibly the limited predictive capability of the classifiers or the intrinsic complexity of the…
This paper discusses and analyzes a class of likelihood models which are based on two distributional innovations in financial models for stock returns. That is, the notion that the marginal distribution of aggregate returns of log-stock…
A random balanced sample (RBS) is a multivariate distribution with n components X_1,...,X_n, each uniformly distributed on [-1, 1], such that the sum of these components is precisely 0. The corresponding vectors X lie in an…
This paper develops a new framework for indirect statistical inference with guaranteed necessity and sufficiency, applicable to continuous random variables. We prove that when comparing exponentially transformed order statistics from an…
The distribution of scientific citations for publications selected with different rules (author, topic, institution, country, journal, etc.) collapse on a single curve if one plots the citations relative to their mean value. We find that…
A major challenge in Semi-Supervised Learning (SSL) is the limited information available about the class distribution in the unlabeled data. In many real-world applications this arises from the prevalence of long-tailed distributions, where…
Consider a researcher estimating the parameters of a regression function based on data for all 50 states in the United States or on data for all visits to a website. What is the interpretation of the estimated parameters and the standard…