English
Related papers

Related papers: Counting on count regression: overlooked aspects o…

200 papers

The computation and inversion of the binomial and negative binomial cumulative distribution functions play a key role in many applications. In this paper, we explain how methods used for the central beta distribution function (described in…

Classical Analysis and ODEs · Mathematics 2020-01-14 A. Gil , J. Segura , N. M. Temme

Binary classification involves predicting the label of an instance based on whether the model score for the positive class exceeds a threshold chosen based on the application requirements (e.g., maximizing recall for a precision bound).…

Machine Learning · Computer Science 2023-11-21 Gundeep Arora , Srujana Merugu , Anoop Saladi , Rajeev Rastogi

In practice functional data are sampled on a discrete set of observation points and often susceptible to noise. We consider in this paper the setting where such data are used as explanatory variables in a regression problem. If the primary…

Methodology · Statistics 2021-12-14 Siegfried Hörmann , Fatima Jammoul

This study introduces an approach to estimate the uncertainty in bibliometric indicator values that is caused by data errors. This approach utilizes Bayesian regression models, estimated from empirical data samples, which are used to…

Digital Libraries · Computer Science 2024-12-11 Paul Donner

Omitted variable bias occurs when a statistical model leaves out variables that are relevant determinants of the effects under study. This results in the model attributing the missing variables' effect to some of the included variables --…

Software Engineering · Computer Science 2026-04-02 Carlo A. Furia , Richard Torkar

The Fisher information matrix is a quantity of fundamental importance for information geometry and asymptotic statistics. In practice, it is widely used to quickly estimate the expected information available in a data set and guide…

Methodology · Statistics 2023-06-06 William R. Coulton , Benjamin D. Wandelt

We analyze different types of simulations that applied researchers can use to assess whether their inference methods reliably control false-positive rates. We show that different assessments involve trade-offs, varying in the types of…

Econometrics · Economics 2025-10-03 Bruno Ferman

Paraphrase Identification is a fundamental task in Natural Language Processing. While much progress has been made in the field, the performance of many state-of-the-art models often suffer from distribution shift during inference time. We…

Computation and Language · Computer Science 2022-10-06 Yifei Zhou , Renyu Li , Hayden Housen , Ser-Nam Lim

Pragmatic clinical trials evaluate the effectiveness of health interventions in real-world settings. Negative spillover can arise in a pragmatic trial if the study intervention affects how scarce resources are allocated between patients in…

Methodology · Statistics 2024-12-20 Sean Mann

In this article, we propose a new three parameter distribution by compounding negative binomial with reciprocal inverse Gaussian model called negative binomial-reciprocal inverse Gaussian distribution. This model is tractable with some…

Methodology · Statistics 2019-06-10 Ishfaq Shah Ahmad , Anwar Hassan , Peer Bilal Ahmad

Quantifying the influence of infinitesimal changes in training data on model performance is crucial for understanding and improving machine learning models. In this work, we reformulate this problem as a weighted empirical risk minimization…

Machine Learning · Computer Science 2025-04-11 Omri Lev , Ashia C. Wilson

The fundamental problem of causal inference -- that we never observe counterfactuals -- prevents us from identifying how many might be negatively affected by a proposed intervention. If, in an A/B test, half of users click (or buy, or…

Methodology · Statistics 2022-11-22 Nathan Kallus

I present a critique of the methods used in a typical paper. This leads to three broad conclusions about the conventional use of statistical methods. First, results are often reported in an unnecessarily obscure manner. Second, the null…

Applications · Statistics 2013-03-05 Michael Wood

Count data take on non-negative integer values and are challenging to properly analyze using standard linear-Gaussian methods such as linear regression and principal components analysis. Generalized linear models enable direct modeling of…

Methodology · Statistics 2020-01-14 F. William Townes

Negative binomial related distributions have been widely used in practice. The calculation of the corresponding Fisher information matrices involves the expectation of trigamma function values which can only be calculated numerically and…

Computation · Statistics 2024-01-22 Zhou Yu , Niloufar Dousti Mousavi , Jie Yang

When estimating a regression model, we might have data where some labels are missing, or our data might be biased by a selection mechanism. When the response or selection mechanism is ignorable (i.e., independent of the response variable…

Statistics Theory · Mathematics 2023-08-22 Philip Boeken , Noud de Kroon , Mathijs de Jong , Joris M. Mooij , Onno Zoeter

We study the estimation of causal parameters when not all confounders are observed and instead negative controls are available. Recent work has shown how these can enable identification and efficient estimation via two so-called bridge…

Machine Learning · Statistics 2022-10-11 Nathan Kallus , Xiaojie Mao , Masatoshi Uehara

Anomaly detection presents a unique challenge in machine learning, due to the scarcity of labeled anomaly data. Recent work attempts to mitigate such problems by augmenting training of deep anomaly detection models with additional labeled…

Machine Learning · Computer Science 2021-05-18 Ziyu Ye , Yuxin Chen , Haitao Zheng

Matrix completion focuses on recovering missing or incomplete information in matrices. This problem arises in various applications, including image processing and network analysis. Previous research proposed Poisson matrix completion for…

Machine Learning · Computer Science 2024-08-30 Yu Lu , Kevin Bui , Roummel F. Marcia

Imputation of missing values is a strategy for handling non-responses in surveys or data loss in measurement processes, which may be more effective than ignoring them. When the variable represents a count, the literature dealing with this…

Applications · Statistics 2020-07-31 Gilma Hernández-Herrera , Albert Navarro , David Moriña