中文
相关论文

相关论文: Missing Data Infill with Automunge

200 篇论文

The use of flexible machine-learning (ML) models to generate imputations of missing data within the framework of Multiple Imputation (MI) has recently gained traction, particularly in observational settings. For randomised controlled trials…

统计方法学 · 统计学 2025-10-07 Mia S. Tackney , Jonathan W. Bartlett , Elizabeth Williamson , Kim May Lee

Traffic data imputation is a critical preprocessing step in intelligent transportation systems, underpinning the reliability of downstream transportation services. Despite substantial progress in imputation models, model selection and…

机器学习 · 计算机科学 2025-10-21 Shengnan Guo , Tonglong Wei , Yiheng Huang , Yan Lin , Zekai Shen , Yujuan Dong , Junliang Lin , Youfang Lin , Huaiyu Wan

We consider the topic of data imputation, a foundational task in machine learning that addresses issues with missing data. To that end, we propose MCFlow, a deep framework for imputation that leverages normalizing flow generative models and…

机器学习 · 计算机科学 2020-03-31 Trevor W. Richardson , Wencheng Wu , Lei Lin , Beilei Xu , Edgar A. Bernal

The quality of underlying training data is very crucial for building performant machine learning models with wider generalizabilty. However, current machine learning (ML) tools lack streamlined processes for improving the data quality. So,…

机器学习 · 计算机科学 2021-12-16 Atindriyo Sanyal , Vikram Chatterji , Nidhi Vyas , Ben Epstein , Nikita Demir , Anthony Corletti

In clinical trials, mixed effects models for repeated measures (MMRM) and pattern mixture models (PMM) are often used to analyze longitudinal continuous outcomes. We describe a simple missing data imputation algorithm for the MMRM that can…

统计方法学 · 统计学 2016-10-13 Yongqiang Tang

Although data may be abundant, complete data is less so, due to missing columns or rows. This missingness undermines the performance of downstream data products that either omit incomplete cases or create derived completed data for…

机器学习 · 计算机科学 2020-06-26 Haw-minn Lu , Giancarlo Perrone , José Unpingco

We study the problem of using low computational cost to automate the choices of learners and hyperparameters for an ad-hoc training dataset and error metric, by conducting trials of different configurations on the given training data. We…

机器学习 · 计算机科学 2021-05-20 Chi Wang , Qingyun Wu , Markus Weimer , Erkang Zhu

Missing attribute values are quite common in the datasets available in the literature. Missing values are also possible because all attributes values may not be recorded and hence unavailable due to several practical reasons. For all these…

信息检索 · 计算机科学 2016-05-04 Yelipe UshaRani , P. Sammulal

Advancements in data collection techniques and the heterogeneity of data resources can yield high percentages of missing observations on variables, such as block-wise missing data. Under missing-data scenarios, traditional methods such as…

统计方法学 · 统计学 2022-05-17 Wei Lan , Xuerong Chen , Tao Zou , Chih-Ling Tsai

Healthcare data frequently contain a substantial proportion of missing values, necessitating effective time series imputation to support downstream disease diagnosis tasks. However, existing imputation methods focus on discrete data points…

机器学习 · 计算机科学 2025-05-19 Mengxuan Li , Ke Liu , Jialong Guo , Jiajun Bu , Hongwei Wang , Haishuai Wang

In practice, we are often faced with small-sized tabular data. However, current tabular benchmarks are not geared towards data-scarce applications, making it very difficult to derive meaningful conclusions from empirical comparisons. We…

机器学习 · 计算机科学 2024-09-04 Ricardo Knauer , Marvin Grimm , Erik Rodner

Many datasets suffer from missing values due to various reasons,which not only increases the processing difficulty of related tasks but also reduces the accuracy of classification. To address this problem, the mainstream approach is to use…

机器学习 · 计算机科学 2024-08-14 Cong Guo , Chun Liu , Wei Yang

Missing value imputation is a crucial preprocessing step for many machine learning problems. However, it is often considered as a separate subtask from downstream applications such as classification, regression, or clustering, and thus is…

机器学习 · 计算机科学 2024-05-02 Adam Catto , Nan Jia , Ansaf Salleb-Aouissi , Anita Raja

Time series classification with missing data is a prevalent issue in time series analysis, as temporal data often contain missing values in practical applications. The traditional two-stage approach, which handles imputation and…

机器学习 · 计算机科学 2024-08-13 Pengshuai Yao , Mengna Liu , Xu Cheng , Fan Shi , Huan Li , Xiufeng Liu , Shengyong Chen

Tabular Foundation Models (TFMs) achieve state-of-the-art zero-shot accuracy on small tabular datasets by meta-learning over synthetic data-generating processes -- making them highly attractive for practitioners who cannot afford large…

机器学习 · 计算机科学 2026-04-29 Laure Berti-Equille

As more Intensive Care Unit (ICU) data becomes available, the interest in developing clinical prediction models to improve healthcare protocols increases. However, the lack of data quality still hinders clinical prediction using Machine…

High-quality, error-free datasets are a key ingredient in building reliable, accurate, and unbiased machine learning (ML) models. However, real world datasets often suffer from errors due to sensor malfunctions, data entry mistakes, or…

机器学习 · 计算机科学 2025-03-11 Tommaso Bendinelli , Artur Dox , Christian Holz

Handling missing values in tabular datasets presents a significant challenge in training and testing artificial intelligence models, an issue usually addressed using imputation techniques. Here we introduce "Not Another Imputation Method"…

机器学习 · 计算机科学 2026-03-13 Camillo Maria Caruso , Paolo Soda , Valerio Guarrasi

Missing values in tabular data restrict the use and performance of machine learning, requiring the imputation of missing values. The most popular imputation algorithm is arguably multiple imputations using chains of equations (MICE), which…

机器学习 · 计算机科学 2022-03-01 Manar D Samad , Sakib Abrar , Norou Diawara

Industrial applications of machine learning face unique challenges due to the nature of raw industry data. Preprocessing and preparing raw industrial data for machine learning applications is a demanding task that often takes more time and…

机器学习 · 计算机科学 2021-09-09 Philipp Fleck , Manfred Kügel , Michael Kommenda