中文
相关论文

相关论文: An Empirical Comparison of Multiple Imputation Met…

200 篇论文

This work proposes a non-iterative strategy for missing value imputations which is guided by similarity between observations, but instead of explicitly determining distances or nearest neighbors, it assigns observations to overlapping…

机器学习 · 统计学 2019-11-25 David Cortes

Multivariate time-series data are used in many classification and regression predictive tasks, and recurrent models have been widely used for such tasks. Most common recurrent models assume that time-series data elements are of equal length…

机器学习 · 计算机科学 2020-09-21 Mehak Gupta , Rahmatollah Beheshti

Missing values pose a persistent challenge in modern data science. Consequently, there is an ever-growing number of publications introducing new imputation methods in various fields. While many studies compare imputation approaches, they…

统计计算 · 统计学 2025-11-10 Krystyna Grzesiak , Christophe Muller , Julie Josse , Jeffrey Näf

In many applications, researchers seek to identify overlapping entities across multiple data files. Record linkage algorithms facilitate this task, in the absence of unique identifiers. As these algorithms rely on semi-identifying…

统计方法学 · 统计学 2026-04-24 Gauri Kamat , Roee Gutman

Missing values in tabular data restrict the use and performance of machine learning, requiring the imputation of missing values. The most popular imputation algorithm is arguably multiple imputations using chains of equations (MICE), which…

机器学习 · 计算机科学 2022-03-01 Manar D Samad , Sakib Abrar , Norou Diawara

This paper presents algorithm for missing values imputation in categorical data. The algorithm is based on using association rules and is presented in three variants. Experimental shows better accuracy of missing values imputation using the…

机器学习 · 计算机科学 2012-11-09 Jiří Kaiser

Iterative imputation, in which variables are imputed one at a time each given a model predicting from all the others, is a popular technique that can be convenient and flexible, as it replaces a potentially difficult multivariate modeling…

统计理论 · 数学 2012-04-04 Jingchen Liu , Andrew Gelman , Jennifer Hill , Yu-Sung Su

Statistical matching is a technique for integrating two or more data sets when information available for matching records for individual participants across data sets is incomplete. Statistical matching can be viewed as a missing data…

统计方法学 · 统计学 2015-10-14 Jae-kwang Kim , Emily Berg , Taesung Park

By filling in missing values in datasets, imputation allows these datasets to be used with algorithms that cannot handle missing values by themselves. However, missing values may in principle contribute useful information that is lost…

机器学习 · 计算机科学 2024-10-31 Oliver Urs Lenz , Daniel Peralta , Chris Cornelis

Ecological Momentary Assessments (EMA) capture real-time thoughts and behaviors in natural settings, producing rich longitudinal data for statistical and physiological analyses. However, the robustness of these analyses can be compromised…

统计方法学 · 统计学 2023-11-21 Yiheng Wei , Donald Hedeker

Causal discovery algorithms estimate causal graphs from observational data. This can provide a valuable complement to analyses focussing on the causal relation between individual treatment-outcome pairs. Constraint-based causal discovery…

统计方法学 · 统计学 2021-08-31 Janine Witte , Ronja Foraita , Vanessa Didelez

We present and compare multiple imputation methods for multilevel continuous and binary data where variables are systematically and sporadically missing. The methods are compared from a theoretical point of view and through an extensive…

In binary-transaction data-mining, traditional frequent itemset mining often produces results which are not straightforward to interpret. To overcome this problem, probability models are often used to produce more compact and conclusive…

机器学习 · 计算机科学 2012-09-27 Ruefei He , Jonathan Shapiro

Multiple imputation (MI) is a technique especially designed for handling missing data in public-use datasets. It allows analysts to perform incomplete-data inference straightforwardly by using several already imputed datasets released by…

统计方法学 · 统计学 2022-01-03 Kin Wai Chan

We develop a Bayesian framework for tackling the supervised clustering problem, the generic problem encountered in tasks such as reference matching, coreference resolution, identity uncertainty and record linkage. Our clustering model is…

机器学习 · 计算机科学 2009-07-07 Hal Daumé , Daniel Marcu

In many machine learning applications, we are faced with incomplete datasets. In the literature, missing data imputation techniques have been mostly concerned with filling missing values. However, the existence of missing values is…

机器学习 · 计算机科学 2020-09-07 Mohammad Kachuee , Kimmo Karkkainen , Orpaz Goldstein , Sajad Darabi , Majid Sarrafzadeh

Missing data is a common problem faced with real-world datasets. Imputation is a widely used technique to estimate the missing data. State-of-the-art imputation approaches, such as Generative Adversarial Imputation Nets (GAIN), model the…

机器学习 · 计算机科学 2020-12-02 Saqib Ejaz Awan , Mohammed Bennamoun , Ferdous Sohel , Frank M Sanfilippo , Girish Dwivedi

Most datasets suffer from partial or complete missing values, which has downstream limitations on the available models on which to test the data and on any statistical inferences that can be made from the data. Several imputation techniques…

机器学习 · 统计学 2023-02-09 Adrienne Kline , Yuan Luo

Multiple imputation (MI) is a method for repairing and analyzing data with missing values. MI replaces missing values with a sample of random values drawn from an imputation model. The most popular form of MI, which we call posterior draw…

统计方法学 · 统计学 2019-11-18 Paul T. von Hippel , Jonathan Bartlett

The ICH E9(R1) Addendum (International Council for Harmonization 2019) suggests treatment-policy as one of several strategies for addressing intercurrent events such as treatment withdrawal when defining an estimand. This strategy requires…

统计方法学 · 统计学 2023-08-28 Suzie Cro , James H Roger , James R Carpenter