中文
相关论文

相关论文: CFMI: Flow Matching for Missing Data Imputation

200 篇论文

Missing data has a ubiquitous presence in real-life applications of machine learning techniques. Imputation methods are algorithms conceived for restoring missing values in the data, based on other entries in the database. The choice of the…

机器学习 · 计算机科学 2017-08-16 Unai Garciarena , Roberto Santana , Alexander Mendiburu

Given the prevalence of missing data in modern statistical research, a broad range of methods is available for any given imputation task. How does one choose the `best' imputation method in a given application? The standard approach is to…

应用统计 · 统计学 2022-12-01 Jeffrey Näf , Meta-Lina Spohn , Loris Michel , Nicolai Meinshausen

Handling missing values in tabular datasets presents a significant challenge in training and testing artificial intelligence models, an issue usually addressed using imputation techniques. Here we introduce "Not Another Imputation Method"…

机器学习 · 计算机科学 2026-03-13 Camillo Maria Caruso , Paolo Soda , Valerio Guarrasi

Missing value imputation is crucial for real-world data science workflows. Imputation is harder in the online setting, as it requires the imputation method itself to be able to evolve over time. For practical applications, imputation…

机器学习 · 计算机科学 2021-12-17 Yuxuan Zhao , Eric Landgrebe , Eliot Shekhtman , Madeleine Udell

Modern biomedical survival studies with high-dimensional genomic and clinical predictors are challenged by missing covariates. Existing methods conduct inference through penalization and debiasing when the number of covariates diverges with…

统计方法学 · 统计学 2026-05-22 Zhilin Zhang , Yi Li

The problem of missing data, usually absent incurated and competition-standard datasets, is an unfortunate reality for most machine learning models used in industry applications. Recent work has focused on understanding the nature and the…

Recent advancements in generative modeling, particularly diffusion models, have opened new directions for time series modeling, achieving state-of-the-art performance in forecasting and synthesis. However, the reliance of diffusion-based…

机器学习 · 计算机科学 2025-05-13 Marcel Kollovieh , Marten Lienen , David Lüdke , Leo Schwinn , Stephan Günnemann

Missing values are pervasive in real-world tabular data and can significantly impair downstream analysis. Imputing them is especially challenging in text-rich tables, where dependencies are implicit, complex, and dispersed across long…

数据库 · 计算机科学 2026-05-12 Soroush Omidvartehrani , Davood Rafiei

Case-cohort studies are conducted within cohort studies, wherein collection of exposure data is limited to a subset of the cohort, leading to a large proportion of missing data by design. Standard analysis uses inverse probability weighting…

We introduce Categorical Flow Maps, a flow-matching method for accelerated few-step generation of categorical data via self-distillation. Building on recent variational formulations of flow matching and the broader trend towards accelerated…

Gaussian Mixture models (GMMs) are a powerful tool for clustering, classification and density estimation when clustering structures are embedded in the data. The presence of missing values can largely impact the GMMs estimation process,…

机器学习 · 统计学 2020-06-05 Alessio Serafini , Thomas Brendan Murphy , Luca Scrucca

Rule models are often preferred in prediction tasks with tabular inputs as they can be easily interpreted using natural language and provide predictive performance on par with more complex models. However, most rule models' predictions are…

机器学习 · 计算机科学 2023-11-27 Lena Stempfle , Fredrik D. Johansson

There has been an increasing interest in using cell and gene therapy (CGT) to treat/cure difficult diseases. The hallmark of CGT trials are the small sample size and extremely high efficacy. Due to the innovation and novelty of such…

应用统计 · 统计学 2025-10-23 Yaoyuan Vincent Tan , Gang Xu , Chenkun Wang

The integrity of time series data in smart grids is often compromised by missing values due to sensor failures, transmission errors, or disruptions. Gaps in smart meter data can bias consumption analyses and hinder reliable predictions,…

人工智能 · 计算机科学 2025-02-21 Amir Sartipi , Joaquín Delgado Fernández , Sergio Potenciano Menci , Alessio Magitteri

Inverse problems of partial differential equations are ubiquitous across various scientific disciplines and can be formulated as statistical inference problems using Bayes' theorem. To address large-scale problems, it is crucial to develop…

数值分析 · 数学 2025-12-23 Yang Zhao , Haoyu Lu , Junxiong Jia , Tao Zhou

Modern multi-modal and multi-site data frequently suffer from blockwise missingness, where subsets of features are missing for groups of individuals, creating complex patterns that challenge standard inference methods. Existing approaches…

统计方法学 · 统计学 2025-09-18 Sarah Zhao , Emmanuel Candès

Diffusion models (DMs) have gained attention in Missing Data Imputation (MDI), but there remain two long-neglected issues to be addressed: (1). Inaccurate Imputation, which arises from inherently sample-diversification-pursuing generative…

机器学习 · 计算机科学 2024-06-25 Zhichao Chen , Haoxuan Li , Fangyikang Wang , Odin Zhang , Hu Xu , Xiaoyu Jiang , Zhihuan Song , Eric H. Wang

Missing covariates in regression or classification problems can prohibit the direct use of advanced tools for further analysis. Recent research has realized an increasing trend towards the usage of modern Machine Learning algorithms for…

机器学习 · 统计学 2022-03-23 Burim Ramosaj , Justus Tulowietzki , Markus Pauly

We study high-dimensional regression with missing entries in the covariates. A common strategy in practice is to \emph{impute} the missing entries with an appropriate substitute and then implement a standard statistical procedure acting as…

统计理论 · 数学 2020-01-28 Kabir Aladin Chandrasekher , Ahmed El Alaoui , Andrea Montanari

Iterative imputation is a popular tool to accommodate missing data. While it is widely accepted that valid inferences can be obtained with this technique, these inferences all rely on algorithmic convergence. There is no consensus on how to…

统计计算 · 统计学 2021-10-25 Hanne Ida Oberman , Stef van Buuren , Gerko Vink