中文
相关论文

相关论文: PPI is the Difference Estimator: Recognizing the S…

200 篇论文

To estimate accurately the parameters of a regression model, the sample size must be large enough relative to the number of possible predictors for the model. In practice, sufficient data is often lacking, which can lead to overfitting of…

应用统计 · 统计学 2024-09-25 Marianne A Jonker , Hassan Pazira , Anthony CC Coolen

Recidivism prediction instruments (RPI's) provide decision makers with an assessment of the likelihood that a criminal defendant will reoffend at a future point in time. While such instruments are gaining increasing popularity across the…

应用统计 · 统计学 2017-03-02 Alexandra Chouldechova

The spatial scan statistic is widely used to detect disease clusters in epidemiological surveillance. Since the seminal work by~\cite{kulldorff1997}, numerous extensions have emerged, including methods for defining scan regions, detecting…

统计方法学 · 统计学 2025-02-11 Takayuki Kawashima , Daisuke Yoneoka , Yuta Tanoue , Akifumi Eguchi , Shuhei Nomura

Consider a researcher estimating the parameters of a regression function based on data for all 50 states in the United States or on data for all visits to a website. What is the interpretation of the estimated parameters and the standard…

统计理论 · 数学 2019-06-25 Alberto Abadie , Susan Athey , Guido W. Imbens , Jeffrey M. Wooldridge

Several well known estimators of finite population mean and its functions are investigated under some standard sampling designs. Such functions of mean include the variance, the correlation coefficient and the regression coefficient in the…

统计理论 · 数学 2023-05-25 Anurag Dey , Probal Chaudhuri

Panels with large time $(T)$ and cross-sectional $(N)$ dimensions are a key data structure in social sciences and other fields. A central question in panel data analysis is whether to pool data across individuals or to estimate separate…

统计方法学 · 统计学 2025-12-18 Tim Kutta , Martin Schumann , Holger Dette

Statistical models are central to machine learning with broad applicability across a range of downstream tasks. The models are controlled by free parameters that are typically estimated from data by maximum-likelihood estimation or…

机器学习 · 计算机科学 2023-08-16 Vaidotas Simkus , Benjamin Rhodes , Michael U. Gutmann

In order to estimate the population mean in the presence of both non-response and measurement errors that are uncorrelated, the paper presents some novel estimators employing ranked set sampling by utilizing auxiliary information.Up to the…

统计方法学 · 统计学 2023-11-06 Rajesh Singh , Anamika Kumari

A good prediction is very important for scientific, economic, and administrative purposes. It is therefore necessary to know whether a predictor is skillful enough to predict the future. Given the increased reliance on predictions in…

综合经济学 · 经济学 2022-09-13 Thitithep Sitthiyot , Kanyarat Holasut

We establish a general framework for statistical inferences with non-probability survey samples when relevant auxiliary information is available from a probability survey sample. We develop a rigorous procedure for estimating the propensity…

统计方法学 · 统计学 2018-05-17 Yilin Chen , Pengfei Li , Changbao Wu

While Gaussian processes (GPs) are the method of choice for regression tasks, they also come with practical difficulties, as inference cost scales cubic in time and quadratic in memory. In this paper, we introduce a natural and expressive…

机器学习 · 计算机科学 2018-09-13 Martin Trapp , Robert Peharz , Carl E. Rasmussen , Franz Pernkopf

A traditional approach to assessing emerging intelligence in the theory of intelligent systems is based on the similarity, "imitation" of human-like actions and behaviors, benchmarking the performance of intelligent systems on the scale of…

神经与进化计算 · 计算机科学 2025-05-28 Serge Dolgikh

Negative controls are increasingly used to evaluate the presence of potential unmeasured confounding in observational studies. Beyond the use of negative controls to detect the presence of residual confounding, proximal causal inference…

统计方法学 · 统计学 2024-06-06 Jiewen Liu , Chan Park , Kendrick Li , Eric J. Tchetgen Tchetgen

Process performance indicators (PPIs) are metrics to quantify the degree with which organizational goals defined based on business processes are fulfilled. They exploit the event logs recorded by information systems during the execution of…

密码学与安全 · 计算机科学 2021-03-23 Martin Kabierski , Stephan Fahrenkrog-Petersen , Matthias Weidlich

Scientists and practitioners increasingly rely on machine learning to model data and draw conclusions. Compared to statistical modeling approaches, machine learning makes fewer explicit assumptions about data structures, such as linearity.…

Artificial intelligence (AI) and machine learning (ML) are increasingly used to generate data for downstream analyses, yet naively treating these predictions as true observations can lead to biased results and incorrect inference. Wang et…

统计方法学 · 统计学 2025-07-15 Stephen Salerno , Kentaro Hoffman , Awan Afiaz , Anna Neufeld , Tyler H. McCormick , Jeffrey T. Leek

We introduce a novel framework for human-AI collaboration in prediction and decision tasks. Our approach leverages human judgment to distinguish inputs which are algorithmically indistinguishable, or "look the same" to any feasible…

机器学习 · 计算机科学 2024-10-21 Rohan Alur , Loren Laine , Darrick K. Li , Dennis Shung , Manish Raghavan , Devavrat Shah

The authors derive likelihood-based exact inference methods for the multivariate regression model, for singly imputed synthetic data generated via Posterior Predictive Sampling (PPS) and for multiply imputed synthetic data generated via a…

统计理论 · 数学 2017-07-26 Ricardo Moura , Martin Klein , Carlos A. Coelho , Bimal Sinha

Modern multi-modal and multi-site data frequently suffer from blockwise missingness, where subsets of features are missing for groups of individuals, creating complex patterns that challenge standard inference methods. Existing approaches…

统计方法学 · 统计学 2025-09-18 Sarah Zhao , Emmanuel Candès

Gradient boosting of regression trees is a competitive procedure for learning predictive models of continuous data that fits the data with an additive non-parametric model. The classic version of gradient boosting assumes that the data is…

机器学习 · 计算机科学 2016-07-04 Iman Alodah , Jennifer Neville