English
Related papers

Related papers: PPI is the Difference Estimator: Recognizing the S…

200 papers

To estimate accurately the parameters of a regression model, the sample size must be large enough relative to the number of possible predictors for the model. In practice, sufficient data is often lacking, which can lead to overfitting of…

Applications · Statistics 2024-09-25 Marianne A Jonker , Hassan Pazira , Anthony CC Coolen

Recidivism prediction instruments (RPI's) provide decision makers with an assessment of the likelihood that a criminal defendant will reoffend at a future point in time. While such instruments are gaining increasing popularity across the…

Applications · Statistics 2017-03-02 Alexandra Chouldechova

The spatial scan statistic is widely used to detect disease clusters in epidemiological surveillance. Since the seminal work by~\cite{kulldorff1997}, numerous extensions have emerged, including methods for defining scan regions, detecting…

Methodology · Statistics 2025-02-11 Takayuki Kawashima , Daisuke Yoneoka , Yuta Tanoue , Akifumi Eguchi , Shuhei Nomura

Consider a researcher estimating the parameters of a regression function based on data for all 50 states in the United States or on data for all visits to a website. What is the interpretation of the estimated parameters and the standard…

Statistics Theory · Mathematics 2019-06-25 Alberto Abadie , Susan Athey , Guido W. Imbens , Jeffrey M. Wooldridge

Several well known estimators of finite population mean and its functions are investigated under some standard sampling designs. Such functions of mean include the variance, the correlation coefficient and the regression coefficient in the…

Statistics Theory · Mathematics 2023-05-25 Anurag Dey , Probal Chaudhuri

Panels with large time $(T)$ and cross-sectional $(N)$ dimensions are a key data structure in social sciences and other fields. A central question in panel data analysis is whether to pool data across individuals or to estimate separate…

Methodology · Statistics 2025-12-18 Tim Kutta , Martin Schumann , Holger Dette

Statistical models are central to machine learning with broad applicability across a range of downstream tasks. The models are controlled by free parameters that are typically estimated from data by maximum-likelihood estimation or…

Machine Learning · Computer Science 2023-08-16 Vaidotas Simkus , Benjamin Rhodes , Michael U. Gutmann

In order to estimate the population mean in the presence of both non-response and measurement errors that are uncorrelated, the paper presents some novel estimators employing ranked set sampling by utilizing auxiliary information.Up to the…

Methodology · Statistics 2023-11-06 Rajesh Singh , Anamika Kumari

A good prediction is very important for scientific, economic, and administrative purposes. It is therefore necessary to know whether a predictor is skillful enough to predict the future. Given the increased reliance on predictions in…

General Economics · Economics 2022-09-13 Thitithep Sitthiyot , Kanyarat Holasut

We establish a general framework for statistical inferences with non-probability survey samples when relevant auxiliary information is available from a probability survey sample. We develop a rigorous procedure for estimating the propensity…

Methodology · Statistics 2018-05-17 Yilin Chen , Pengfei Li , Changbao Wu

While Gaussian processes (GPs) are the method of choice for regression tasks, they also come with practical difficulties, as inference cost scales cubic in time and quadratic in memory. In this paper, we introduce a natural and expressive…

Machine Learning · Computer Science 2018-09-13 Martin Trapp , Robert Peharz , Carl E. Rasmussen , Franz Pernkopf

A traditional approach to assessing emerging intelligence in the theory of intelligent systems is based on the similarity, "imitation" of human-like actions and behaviors, benchmarking the performance of intelligent systems on the scale of…

Neural and Evolutionary Computing · Computer Science 2025-05-28 Serge Dolgikh

Negative controls are increasingly used to evaluate the presence of potential unmeasured confounding in observational studies. Beyond the use of negative controls to detect the presence of residual confounding, proximal causal inference…

Methodology · Statistics 2024-06-06 Jiewen Liu , Chan Park , Kendrick Li , Eric J. Tchetgen Tchetgen

Process performance indicators (PPIs) are metrics to quantify the degree with which organizational goals defined based on business processes are fulfilled. They exploit the event logs recorded by information systems during the execution of…

Cryptography and Security · Computer Science 2021-03-23 Martin Kabierski , Stephan Fahrenkrog-Petersen , Matthias Weidlich

Scientists and practitioners increasingly rely on machine learning to model data and draw conclusions. Compared to statistical modeling approaches, machine learning makes fewer explicit assumptions about data structures, such as linearity.…

Artificial intelligence (AI) and machine learning (ML) are increasingly used to generate data for downstream analyses, yet naively treating these predictions as true observations can lead to biased results and incorrect inference. Wang et…

We introduce a novel framework for human-AI collaboration in prediction and decision tasks. Our approach leverages human judgment to distinguish inputs which are algorithmically indistinguishable, or "look the same" to any feasible…

Machine Learning · Computer Science 2024-10-21 Rohan Alur , Loren Laine , Darrick K. Li , Dennis Shung , Manish Raghavan , Devavrat Shah

The authors derive likelihood-based exact inference methods for the multivariate regression model, for singly imputed synthetic data generated via Posterior Predictive Sampling (PPS) and for multiply imputed synthetic data generated via a…

Statistics Theory · Mathematics 2017-07-26 Ricardo Moura , Martin Klein , Carlos A. Coelho , Bimal Sinha

Modern multi-modal and multi-site data frequently suffer from blockwise missingness, where subsets of features are missing for groups of individuals, creating complex patterns that challenge standard inference methods. Existing approaches…

Methodology · Statistics 2025-09-18 Sarah Zhao , Emmanuel Candès

Gradient boosting of regression trees is a competitive procedure for learning predictive models of continuous data that fits the data with an additive non-parametric model. The classic version of gradient boosting assumes that the data is…

Machine Learning · Computer Science 2016-07-04 Iman Alodah , Jennifer Neville