中文
相关论文

相关论文: Are deep learning models superior for missing data…

200 篇论文

Multiple imputation is widely used for handling missing data in real-world applications. For variable selection on multiply-imputed datasets, however, if selection is performed on each imputed dataset separately, it can result in different…

统计方法学 · 统计学 2025-08-07 Jungang Zou , Sijian Wang , Qixuan Chen

Decision trees are a popular family of models due to their attractive properties such as interpretability and ability to handle heterogeneous data. Concurrently, missing data is a prevalent occurrence that hinders performance of machine…

机器学习 · 计算机科学 2020-07-01 Pasha Khosravi , Antonio Vergari , YooJung Choi , Yitao Liang , Guy Van den Broeck

Presence of missing values in a dataset can adversely affect the performance of a classifier. Single and Multiple Imputation are normally performed to fill in the missing values. In this paper, we present several variants of combining…

机器学习 · 计算机科学 2019-10-16 Shehroz S. Khan , Amir Ahmad , Alex Mihailidis

Data imputation, the process of filling in missing feature elements for incomplete data sets, plays a crucial role in data-driven learning. A fundamental belief is that data imputation is helpful for learning performance, and it follows…

机器学习 · 计算机科学 2025-09-30 Ruikai Yang , Fan He , Mingzhen He , Kaijie Wang , Xiaolin Huang

Deep Learning (DL) methods have dramatically increased in popularity in recent years. While its initial success was demonstrated in the classification and manipulation of image data, there has been significant growth in the application of…

机器学习 · 计算机科学 2022-06-22 David K. Lim , Naim U. Rashid , Junier B. Oliva , Joseph G. Ibrahim

Datasets with missing values are very common on industry applications, and they can have a negative impact on machine learning models. Recent studies introduced solutions to the problem of imputing missing values based on deep generative…

机器学习 · 计算机科学 2019-02-28 Ramiro D. Camino , Christian A. Hammerschmidt , Radu State

Understanding the structural mechanisms of multi-layer networks is essential for analyzing complex systems characterized by multiple interacting layers. This work studies the problem of estimating connection probabilities in multi-layer…

统计方法学 · 统计学 2026-02-26 Dingzi Guo , Diqing Li , Jingyi Wang , Wen-Xin Zhou

Multiple imputation has become one of the standard methods in drawing inferences in many incomplete data applications. Applications of multiple imputation in relatively more complex settings, such as high-dimensional clustered data, require…

统计方法学 · 统计学 2025-04-08 Qiushuang Li , Recai Yucel

Many recent methods for unsupervised or self-supervised representation learning train feature extractors by maximizing an estimate of the mutual information (MI) between different views of the data. This comes with several immediate…

机器学习 · 计算机科学 2020-01-24 Michael Tschannen , Josip Djolonga , Paul K. Rubenstein , Sylvain Gelly , Mario Lucic

Missing covariates in regression or classification problems can prohibit the direct use of advanced tools for further analysis. Recent research has realized an increasing trend towards the usage of modern Machine Learning algorithms for…

机器学习 · 统计学 2022-03-23 Burim Ramosaj , Justus Tulowietzki , Markus Pauly

The rapid adoption of complex Artificial Intelligence (AI) and Machine Learning (ML) models has led to their characterization as black boxes due to the difficulty of explaining their internal decision-making processes. This lack of…

Background: Multiple imputation is often used to reduce bias and gain efficiency when there is missing data. The most appropriate imputation method depends on the model the analyst is interested in fitting. Several imputation approaches…

统计方法学 · 统计学 2022-11-29 Matthew J. Smith , Matteo Quartagno , Edmund Njeru Njagi

The challenge of missing data remains a significant obstacle across various scientific domains, necessitating the development of advanced imputation techniques that can effectively address complex missingness patterns. This study introduces…

机器学习 · 计算机科学 2025-01-22 Harsh Joshi , Rajeshwari Mistri , Manasi Mali , Nachiket Kapure , Parul Kumari

The National Health and Nutrition Examination Survey (NHANES) studies the nutritional and health status over the whole U.S. population with comprehensive physical examinations and questionnaires. However, survey data analyses become…

统计方法学 · 统计学 2019-08-06 Xiaojun Mao , Zhonglei Wang , Shu Yang

Identifying leading measurement units from a large collection is a common inference task in various domains of large-scale inference. Testing approaches, which measure evidence against a null hypothesis rather than effect magnitude, tend to…

统计方法学 · 统计学 2020-11-17 Nicholas C. Henderson , Michael A. Newton

When treatment policy estimands are of interest, clinical trials often attempt to collect patient data after intercurrent events (ICEs), although such data are often limited. Retrieved dropout imputation methods, which use pre-ICE and…

统计方法学 · 统计学 2026-03-31 Brendah Nansereko , Marcel Wolbers , James R. Carpenter , Jonathan W. Bartlett

Pre-trained machine learning (ML) predictions have been increasingly used to complement incomplete data to enable downstream scientific inquiries, but their naive integration risks biased inferences. Recently, multiple methods have been…

统计方法学 · 统计学 2025-11-12 Xingran Chen , Tyler McCormick , Bhramar Mukherjee , Zhenke Wu

Missing data is a common concern in health datasets, and its impact on good decision-making processes is well documented. Our study's contribution is a methodology for tackling missing data problems using a combination of synthetic dataset…

机器学习 · 计算机科学 2022-11-08 Gift Khangamwa , Terence L. van Zyl , Clint J. van Alten

Mutual Information (MI) plays an important role in representation learning. However, MI is unfortunately intractable in continuous and high-dimensional settings. Recent advances establish tractable and scalable MI estimators to discover…

机器学习 · 统计学 2020-05-05 Liangjian Wen , Yiji Zhou , Lirong He , Mingyuan Zhou , Zenglin Xu

Often in real-world datasets, especially in high dimensional data, some feature values are missing. Since most data analysis and statistical methods do not handle gracefully missing values, the first step in the analysis requires the…

机器学习 · 统计学 2016-12-08 Yehezkel S. Resheff , Daphna Weinshall