中文
相关论文

相关论文: Imputation of missing values in multi-view data

200 篇论文

Time series data are ubiquitous in real-world applications. However, one of the most common problems is that the time series data could have missing values by the inherent nature of the data collection process. So imputing missing values…

机器学习 · 计算机科学 2022-09-23 Eunkyu Oh , Taehun Kim , Yunhu Ji , Sushil Khyalia

Imputing missing values is an important preprocessing step in data analysis, but the literature offers little guidance on how to choose between different imputation models. This letter suggests adopting the imputation model that generates a…

统计方法学 · 统计学 2021-07-13 Moritz Marbach

Tree-based learning methods such as Random Forest and XGBoost are still the gold-standard prediction methods for tabular data. Feature importance measures are usually considered for feature selection as well as to assess the effect of…

应用统计 · 统计学 2024-12-19 Jakob Schwerter , Andrés Romero , Florian Dumpert , Markus Pauly

Missing values in tabular data restrict the use and performance of machine learning, requiring the imputation of missing values. The most popular imputation algorithm is arguably multiple imputations using chains of equations (MICE), which…

机器学习 · 计算机科学 2022-03-01 Manar D Samad , Sakib Abrar , Norou Diawara

Multiple imputation by chained equations (MICE) has emerged as a popular approach for handling missing data. A central challenge for applying MICE is determining how to incorporate outcome information into covariate imputation models,…

统计方法学 · 统计学 2019-10-11 Lauren Beesley , Jeremy M G Taylor

This work considers the problem of fitting functional models with sparsely and irregularly sampled functional data. It overcomes the limitations of the state-of-the-art methods, which face major challenges in the fitting of more complex…

统计方法学 · 统计学 2023-05-02 Aniruddha Rajendra Rao , Matthew Reimherr

Supervised learning with missing data aims at building the best prediction of a target output based on partially-observed inputs. Major approaches to address this problem can be decomposed into $(i)$ impute-then-predict strategies, which…

统计理论 · 数学 2024-10-14 Angel D Reyero Lobo , Alexis Ayme , Claire Boyer , Erwan Scornet

Imputation of missing data in large regions of satellite imagery is necessary when the acquired image has been damaged by shadows due to clouds, or information gaps produced by sensor failure. The general approach for imputation of missing…

应用统计 · 统计学 2010-06-23 Valeria Rulloni , Oscar Bustos , Ana Georgina Flesia

Missing data is a crucial issue when applying machine learning algorithms to real-world datasets. Starting from the simple assumption that two batches extracted randomly from the same dataset should share the same distribution, we leverage…

机器学习 · 统计学 2020-07-02 Boris Muzellec , Julie Josse , Claire Boyer , Marco Cuturi

Incomplete multi-view clustering is a challenging and non-trivial task to provide effective data analysis for large amounts of unlabeled data in the real world. All incomplete multi-view clustering methods need to address the problem of how…

机器学习 · 计算机科学 2023-05-22 Sifan Fang

Missing data is a widespread problem in many domains, creating challenges in data analysis and decision making. Traditional techniques for dealing with missing data, such as excluding incomplete records or imputing simple estimates (e.g.,…

数据库 · 计算机科学 2024-01-09 Massimo Perini , Milos Nikolic

In the era of big data, it is common to have data with multiple modalities or coming from multiple sources, known as "multi-view data". Multi-view clustering provides a natural way to generate clusters from such data. Since different views…

机器学习 · 计算机科学 2016-11-08 Weixiang Shao , Lifang He , Chun-Ta Lu , Philip S. Yu

We present a nonparametric Bayesian joint model for multivariate continuous and categorical variables, with the intention of developing a flexible engine for multiple imputation of missing values. The model fuses Dirichlet process mixtures…

应用统计 · 统计学 2015-10-14 Jared S. Murray , Jerome P. Reiter

Multi-view unsupervised feature selection has been proven to be efficient in reducing the dimensionality of multi-view unlabeled data with high dimensions. The previous methods assume all of the views are complete. However, in real…

机器学习 · 计算机科学 2023-01-02 Yanyong Huang , Kejun Guo , Xiuwen Yi , Zhong Li , Tianrui Li

Missing data are a concern in many real world data sets and imputation methods are often needed to estimate the values of missing data, but data sets with excessive missingness and high dimensionality challenge most approaches to…

机器学习 · 统计学 2021-04-22 Andrew J. Becker , James P. Bagrow

A wide range of systems exhibit high dimensional incomplete data. Accurate estimation of the missing data is often desired, and is crucial for many downstream analyses. Many state-of-the-art recovery methods involve supervised learning…

计算机视觉与模式识别 · 计算机科学 2019-03-15 Adrian V. Dalca , John Guttag , Mert R. Sabuncu

Feature and instance co-selection, which aims to reduce both feature dimensionality and sample size by identifying the most informative features and instances, has attracted considerable attention in recent years. However, when dealing with…

机器学习 · 计算机科学 2025-12-18 Yuxin Cai , Yanyong Huang , Jinyuan Chang , Dongjie Wang , Tianrui Li , Xiaoyi Jiang

As systems are getting more autonomous with the development of artificial intelligence, it is important to discover the causal knowledge from observational sensory inputs. By encoding a series of cause-effect relations between events,…

机器学习 · 计算机科学 2020-01-16 Yuhao Wang , Vlado Menkovski , Hao Wang , Xin Du , Mykola Pechenizkiy

This work focuses on designing a pipeline for the prediction of bankruptcy. The presence of missing values, high dimensional data, and highly class-imbalance databases are the major challenges in the said task. A new method for missing data…

机器学习 · 计算机科学 2024-04-02 Debarati Chakraborty , Ravi Ranjan

Data collected in clinical trials are often composed of multiple types of variables. For example, laboratory measurements and vital signs are longitudinal data of continuous or categorical variables, adverse events may be recurrent events,…

统计方法学 · 统计学 2023-01-12 Tuo Wang , Rachel Zilinskas , Ying Li , Yongming Qu