中文
相关论文

相关论文: How many imputations do you need? A two-stage calc…

200 篇论文

The increased availability of massive data sets provides a unique opportunity to discover subtle patterns in their distributions, but also imposes overwhelming computational challenges. To fully utilize the information contained in big…

统计理论 · 数学 2018-04-12 Stanislav Volgushev , Shih-Kang Chao , Guang Cheng

The multivariate probit is popular for modeling correlated binary data, with an attractive balance of flexibility and simplicity. However, considerable challenges remain in computation and in devising a clear statistical framework. Interest…

统计方法学 · 统计学 2020-04-22 Bryan W. Ting , Fred A. Wright , Yi-Hui Zhou

Annotating data via crowdsourcing is time-consuming and expensive. Due to these costs, dataset creators often have each annotator label only a small subset of the data. This leads to sparse datasets with examples that are marked by few…

计算与语言 · 计算机科学 2023-10-06 London Lowmanstone , Ruyuan Wan , Risako Owan , Jaehyung Kim , Dongyeop Kang

Given a small training data set and a learning algorithm, how much more data is necessary to reach a target validation or test performance? This question is of critical importance in applications such as autonomous driving or medical…

计算机视觉与模式识别 · 计算机科学 2022-07-14 Rafid Mahmood , James Lucas , David Acuna , Daiqing Li , Jonah Philion , Jose M. Alvarez , Zhiding Yu , Sanja Fidler , Marc T. Law

Counterfactual explanations constitute among the most popular methods for analyzing black-box systems since they can recommend cost-efficient and actionable changes to the input of a system to obtain the desired system output. While most of…

机器学习 · 计算机科学 2024-05-22 André Artelt , Andreas Gregoriades

A common approach to detect multiple changepoints is to minimise a measure of data fit plus a penalty that is linear in the number of changepoints. This paper shows that the general finite sample behaviour of such a method can be related to…

统计理论 · 数学 2022-08-15 Chao Zheng , Idris A. Eckley , Paul Fearnhead

In this paper, we focus on regression estimation in both the inductive and the transductive case. We assume that we are given a set of features (which can be a base of functions, but not necessarily). We begin by giving a deviation…

统计理论 · 数学 2015-06-26 Pierre Alquier

Objective: Researchers often use model-based multiple imputation to handle missing at random data to minimize bias while making the best use of all available data. However, there are sometimes constraints within the data that make…

统计方法学 · 统计学 2020-11-03 Chinchin Wang , Tyrel Stokes , Russell Steele , Niels Wedderkopp , Ian Shrier

Interpretability and efficiency are two important considerations for the adoption of neural automatic metrics. In this work, we develop strong-performing automatic metrics for reference-based summarization evaluation, based on a two-stage…

计算与语言 · 计算机科学 2023-11-17 Yixin Liu , Alexander R. Fabbri , Yilun Zhao , Pengfei Liu , Shafiq Joty , Chien-Sheng Wu , Caiming Xiong , Dragomir Radev

Generalized linear mixed models are useful in studying hierarchical data with possibly non-Gaussian responses. However, the intractability of likelihood functions poses challenges for estimation. We develop a new method suitable for this…

统计方法学 · 统计学 2022-01-26 Zexi Song , Zhiqiang Tan

Missing data imputation can help improve the performance of prediction models in situations where missing data hide useful information. This paper compares methods for imputing missing categorical data for supervised classification tasks.…

机器学习 · 统计学 2020-08-11 Jason Poulos , Rafael Valle

Iterative phase estimation has long been used in quantum computing to estimate Hamiltonian eigenvalues. This is done by applying many repetitions of the same fundamental simulation circuit to an initial state, and using statistical…

量子物理 · 物理学 2019-07-25 Ian D. Kivlichan , Christopher E. Granade , Nathan Wiebe

Resampling techniques have become increasingly popular for estimation of uncertainty in data collected via surveys. Survey data are also frequently subject to missing data which are often imputed. This note addresses the issue of using…

统计方法学 · 统计学 2023-11-27 Michael W. Robbins , Lane Burgette , Sebastian Bauhoff

When training predictive models on data with missing entries, the most widely used and versatile approach is a pipeline technique where we first impute missing entries and then compute predictions. In this paper, we view prediction with…

机器学习 · 计算机科学 2025-02-25 Dimitris Bertsimas , Arthur Delarue , Jean Pauphilet

Inference for high-dimensional logistic regression models using penalized methods has been a challenging research problem. As an illustration, a major difficulty is the significant bias of the Lasso estimator, which limits its direct…

统计方法学 · 统计学 2024-10-29 Yuming Zhang , Stéphane Guerrier , Runze Li

Machine learning techniques have been developed to learn from complete data. When missing values exist in a dataset, the incomplete data should be preprocessed separately by removing data points with missing values or imputation. In this…

机器学习 · 计算机科学 2020-12-25 Hadi A. Khorshidi , Michael Kirley , Uwe Aickelin

Inference problems with incomplete observations often aim at estimating population properties of unobserved quantities. One simple way to accomplish this estimation is to impute the unobserved quantities of interest at the individual level…

统计方法学 · 统计学 2012-10-16 Vladimir N. Minin , John D. O'Brien , Arseni Seregin

Ratings are frequently used to evaluate and compare subjects in various applications, from education to healthcare, because ratings provide succinct yet credible measures for comparing subjects. However, when multiple rating lists are…

机器学习 · 统计学 2023-12-05 Young Woong Park , Jinhak Kim , Dan Zhu

We investigate the problem of calibration and assessment of predictive rules in prognostic designs when missing values are present in the predictors. Our paper has two key objectives which are entwined. The first is to investigate how the…

统计方法学 · 统计学 2018-10-12 B. J. A. Mertens , E. Banzato , L. C. de Wreede

Multiple imputation is a well-established general technique for analyzing data with missing values. A convenient way to implement multiple imputation is sequential regression multiple imputation (SRMI), also called chained equations…