English
Related papers

Related papers: Automatic Biases Correction

200 papers

The most fundamental problem in statistics is the inference of an unknown probability distribution from a finite number of samples. For a specific observed data set, answers to the following questions would be desirable: (1) Estimation:…

Statistics Theory · Mathematics 2013-01-23 Ali Kinkhabwala

This paper presents some results on the maximum likelihood (ML) estimation from incomplete data. Finite sample properties of conditional observed information matrices are established. They possess positive definiteness and the same Loewner…

Methodology · Statistics 2022-07-26 Budhi Arta Surya

Optimum designs for parameter estimation in generalized regression models are standardly based on the Fisher information matrix (cf. Atkinson et al (2014) for a recent exposition). The corresponding optimality criteria are related to the…

Statistics Theory · Mathematics 2015-07-28 Katarína Burclová , Andrej Pázman

This article introduces a framework for evaluating statistical decisions under both prior ambiguity and likelihood misspecification. We begin with an ambiguity set - a frequentist model that pairs a possibly misspecified likelihood with…

Econometrics · Economics 2026-05-14 Karun Adusumilli

Diffusion Models (DMs) iteratively denoise random samples to produce high-quality data. The iterative sampling process is derived from Stochastic Differential Equations (SDEs), allowing a speed-quality trade-off chosen at inference. Another…

Machine Learning · Computer Science 2024-09-27 Mattias Cross , Anton Ragni

Preferential attachment is an appealing mechanism for modeling power-law behavior of the degree distributions in directed social networks. In this paper, we consider methods for fitting a 5-parameter linear preferential model to network…

Methodology · Statistics 2017-08-29 Phyllis Wan , Tiandong Wang , Richard A. Davis , Sidney I. Resnick

Bayesian models quantify uncertainty and facilitate optimal decision-making in downstream applications. For most models, however, practitioners are forced to use approximate inference techniques that lead to sub-optimal decisions due to…

Machine Learning · Statistics 2019-09-12 Tomasz Kuśmierczyk , Joseph Sakaya , Arto Klami

The stationary distribution of allele frequencies under a variety of Wright--Fisher $k$-allele models with selection and parent independent mutation is well studied. However, the statistical properties of maximum likelihood estimates of…

Applications · Statistics 2009-10-12 Erkan Ozge Buzbas , Paul Joyce

The main features of the statistical approach to inverse problems are described on the example of a linear model with additive noise. The approach does not use any Bayesian hypothesis regarding an unknown object; instead, the standard…

Methodology · Statistics 2017-05-05 V. Yu. Terebizh

Binary classifiers trained on a certain proportion of positive items introduce a bias when applied to data sets with different proportions of positive items. Most solutions for dealing with this issue assume that some information on the…

Machine Learning · Statistics 2021-02-18 Marco J. H. Puts , Piet J. H. Daas

In many cases, the values of some model parameters are determined by maximising the likelihood of a set of data points given the parameter values. The presence of outliers in the data and correlations between data points complicate this…

Numerical Analysis · Computer Science 2017-08-28 M. de Jong

Maximum likelihood estimation (MLE) is a fundamental problem in statistics. Characteristics of the MLE problem for discrete algebraic statistical models are reflected in the geometry of the $\textit{likelihood correspondence}$, a variety…

Statistics Theory · Mathematics 2024-11-19 David Barnhill , John Cobb , Matthew Faust

The Tully-Fisher (TF) and Fundamental Plane (FP) relations are widely used to infer extragalactic distances and peculiar velocities, enabling measurements of large-scale velocity statistics and cosmological parameters. Using the…

Cosmology and Nongalactic Astrophysics · Physics 2026-02-27 Tyann Dumerchat , Raul E. Angulo , Julian Bautista , Cesar Aguayo , Sownak Bose , Lars Hernquist

The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limited by a foundational issue: they are trained on discrete…

Infrared I band photometry and velocity widths for galaxies in 24 clusters, with radial velocities between 1,000 and 10,000 \kms, are used to construct a template Tully--Fisher (TF) relation. The sources of scatter in the TF diagram are…

Astrophysics · Physics 2009-10-28 R. Giovanelli , M. Haynes , T. Herter , N. Vogt , L. da Costa , W. Freudling , J. Salzer , G. Wegner

Threshold selection is a fundamental problem in any threshold-based extreme value analysis. While models are asymptotically motivated, selecting an appropriate threshold for finite samples is difficult and highly subjective through standard…

Methodology · Statistics 2024-10-30 Conor Murphy , Jonathan A. Tawn , Zak Varty

With the growing adoption of machine learning (ML) systems in areas like law enforcement, criminal justice, finance, hiring, and admissions, it is increasingly critical to guarantee the fairness of decisions assisted by ML. In this paper,…

Machine Learning · Computer Science 2024-05-17 Meiyu Zhong , Ravi Tandon

Every student in statistics or data science learns early on that when the sample size largely exceeds the number of variables, fitting a logistic model produces estimates that are approximately unbiased. Every student also learns that there…

Statistics Theory · Mathematics 2022-06-08 Pragya Sur , Emmanuel J. Candes

Machine learning (ML) solutions are prevalent in many applications. However, many challenges exist in making these solutions business-grade. For instance, maintaining the error rate of the underlying ML models at an acceptably low level.…

Machine Learning · Computer Science 2023-05-16 Samuel Ackerman , Axel Bendavid , Eitan Farchi , Orna Raz

This paper considers the maximum likelihood estimation of factor models of high dimension, where the number of variables (N) is comparable with or even greater than the number of observations (T). An inferential theory is developed. We…

Statistics Theory · Mathematics 2012-05-31 Jushan Bai , Kunpeng Li