中文
相关论文

相关论文: Imputation techniques on missing values in breast …

200 篇论文

Longitudinal passive sensing studies for health and behavior outcomes often have missing and incomplete data. Handling missing data effectively is thus a critical data processing and modeling step. Our formative interviews with researchers…

统计方法学 · 统计学 2024-12-10 Akshat Choube , Rahul Majethia , Sohini Bhattacharya , Vedant Das Swain , Jiachen Li , Varun Mishra

The objectives of this paper are to explore ways to analyze breast cancer dataset in the context of unsupervised learning without prior training model. The paper investigates different ways of clustering techniques as well as preprocessing.…

机器学习 · 计算机科学 2021-09-06 Somenath Chakraborty , Beddhu Murali

In clinical settings, we often face the challenge of building prediction models based on small observational data sets. For example, such a data set might be from a medical center in a multi-center study. Differences between centers might…

Multiple imputation is a straightforward method for handling missing data in a principled fashion. This paper presents an overview of multiple imputation, including important theoretical results and their practical implications for…

统计方法学 · 统计学 2018-01-15 Jared S. Murray

This work proposes a non-iterative strategy for missing value imputations which is guided by similarity between observations, but instead of explicitly determining distances or nearest neighbors, it assigns observations to overlapping…

机器学习 · 统计学 2019-11-25 David Cortes

Data corruption, including missing and noisy data, poses significant challenges in real-world machine learning. This study investigates the effects of data corruption on model performance and explores strategies to mitigate these effects…

机器学习 · 计算机科学 2025-05-22 Qi Liu , Wanjing Ma

Missing data frequently occurs in datasets across various domains, such as medicine, sports, and finance. In many cases, to enable proper and reliable analyses of such data, the missing values are often imputed, and it is necessary that the…

We aim to incorporate variable selection routines into variable-by-variable (or sequential) imputation in clustered data to achieve computational improvement in applications with large-scale health data. Specifically, we utilize variable…

统计方法学 · 统计学 2025-04-08 Qiushuang Li , Recai Yucel

An approach to amputation, the process of introducing missing values to a complete dataset, is presented. It allows to construct missingness indicators in a flexible and principled way via copulas and Bernoulli margins and to incorporate…

应用统计 · 统计学 2025-07-28 Marius Hofert , James Jackson , Niels Hagenbuch

Missing values challenge data analysis because many supervised and unsupervised learning methods cannot be applied directly to incomplete data. Matrix completion based on low-rank assumptions are very powerful solution for dealing with…

机器学习 · 统计学 2020-01-30 Aude Sportisse , Claire Boyer , Julie Josse

Emerging evidence indicates that human cancers are intricately linked to human microbiomes, forming an inseparable connection. However, due to limited sample sizes and significant data loss during collection for various reasons, some…

基因组学 · 定量生物学 2024-08-16 Xinyuan Shi , Fangfang Zhu , Wenwen Min

Most practical data science problems encounter missing data. A wide variety of solutions exist, each with strengths and weaknesses that depend upon the missingness-generating process. Here we develop a theoretical framework for training and…

机器学习 · 计算机科学 2022-11-15 Jahan C. Penny-Dimri , Christoph Bergmeir , Julian Smith

Blood lactate concentration is a strong indicator of mortality risk in critically ill patients. While frequent lactate measurements are necessary to assess patient's health state, the measurement is an invasive procedure that can increase…

机器学习 · 计算机科学 2019-10-04 Behrooz Mamandipoor , Mahshid Majd , Monica Moz , Venet Osmani

Missing data is a common challenge when analyzing epidemiological data, and imputation is often used to address this issue. Here, we investigate the scenario where a covariate used in an analysis has missingness and will be imputed. There…

统计方法学 · 统计学 2024-03-04 Lucy D'Agostino McGowan , Sarah C. Lotspeich , Staci A. Hepler

Missing data is a major challenge in clinical research. In electronic medical records, often a large fraction of the values in laboratory tests and vital signs are missing. The missingness can lead to biased estimates and limit our ability…

机器学习 · 计算机科学 2023-04-18 Omer Noy , Ron Shamir

Missing values in multivariate time series data can harm machine learning performance and introduce bias. These gaps arise from sensor malfunctions, blackouts, and human error and are typically addressed by data imputation. Previous work…

机器学习 · 计算机科学 2025-03-04 Mohammad Rafid Ul Islam , Prasad Tadepalli , Alan Fern

Not-at-random missingness presents a challenge in addressing missing data in many health research applications. In this paper, we propose a new approach to account for not-at-random missingness after multiple imputation through weighted…

统计方法学 · 统计学 2021-01-21 Lauren J Beesley , Jeremy M G Taylor

Missing time-series data is a prevalent practical problem. Imputation methods in time-series data often are applied to the full panel data with the purpose of training a model for a downstream out-of-sample task. For example, in finance,…

机器学习 · 统计学 2023-04-13 Jose Blanchet , Fernando Hernandez , Viet Anh Nguyen , Markus Pelger , Xuhui Zhang

Data collected in clinical trials are often composed of multiple types of variables. For example, laboratory measurements and vital signs are longitudinal data of continuous or categorical variables, adverse events may be recurrent events,…

统计方法学 · 统计学 2023-01-12 Tuo Wang , Rachel Zilinskas , Ying Li , Yongming Qu

Missing value imputation in machine learning is the task of estimating the missing values in the dataset accurately using available information. In this task, several deep generative modeling methods have been proposed and demonstrated…

机器学习 · 计算机科学 2023-03-14 Shuhan Zheng , Nontawat Charoenphakdee