English
Related papers

Related papers: A systematic approach to identify and evaluate mis…

200 papers

Missing data is common in applied data science, particularly for tabular data sets found in healthcare, social sciences, and natural sciences. Most supervised learning methods only work on complete data, thus requiring preprocessing such as…

Machine Learning · Computer Science 2023-10-25 Mike Van Ness , Tomas M. Bosschieter , Roberto Halpin-Gregorio , Madeleine Udell

We argue that the selective inclusion of data points based on latent objectives is common in practical situations, such as music sequences. Since this selection process often distorts statistical analysis, previous work primarily views it…

Machine Learning · Computer Science 2024-07-02 Yujia Zheng , Zeyu Tang , Yiwen Qiu , Bernhard Schölkopf , Kun Zhang

Missing data imputation can help improve the performance of prediction models in situations where missing data hide useful information. This paper compares methods for imputing missing categorical data for supervised classification tasks.…

Machine Learning · Statistics 2020-08-11 Jason Poulos , Rafael Valle

Algorithms and technologies are essential tools that pervade all aspects of our daily lives. In the last decades, health care research benefited from new computer-based recruiting methods, the use of federated architectures for data…

Computers and Society · Computer Science 2023-01-26 Chiara Criscuolo , Tommaso Dolci , Mattia Salnitri

When estimating causal effects using observational data, it is desirable to replicate a randomized experiment as closely as possible by obtaining treated and control groups with similar covariate distributions. This goal can often be…

Methodology · Statistics 2010-10-28 Elizabeth A. Stuart

We present a robust framework to perform linear regression with missing entries in the features. By considering an elliptical data distribution, and specifically a multivariate normal model, we are able to conditionally formulate a…

Machine Learning · Computer Science 2022-11-10 Alireza Aghasi , MohammadJavad Feizollahi , Saeed Ghadimi

Stochastic volatility models that treat the variance of a time series as a stochastic process have proven to be important tools for analyzing dynamic variability. Current methods for fitting and conducting inference on stochastic volatility…

Methodology · Statistics 2025-01-28 Gehui Zhang , Gong Tang , Lori Scott , Robert T Krafty

With increasing computing capabilities of modern supercomputers, the size of the data generated from the scientific simulations is growing rapidly. As a result, application scientists need effective data summarization techniques that can…

Human-Computer Interaction · Computer Science 2019-07-30 Soumya Dutta , Ayan Biswas , James Ahrens

Causal analysis has become an essential component in understanding the underlying causes of phenomena across various fields. Despite its significance, existing literature on causal discovery algorithms is fragmented, with inconsistent…

Artificial Intelligence · Computer Science 2024-09-05 Wenjin Niu , Zijun Gao , Liyan Song , Lingbo Li

Many scientific questions in biomedical, environmental, and psychological research involve understanding the effects of multiple factors on outcomes. While factorial experiments are ideal for this purpose, randomized controlled treatment…

Methodology · Statistics 2025-12-03 Ruoqi Yu , Peng Ding

Information theory is widely accepted as a powerful tool for analyzing complex systems and it has been applied in many disciplines. Recently, some central components of information theory - multivariate information measures - have found…

Information Theory · Computer Science 2012-08-30 Nicholas Timme , Wesley Alford , Benjamin Flecker , John M. Beggs

Conditions ensuring optimal parameter estimation in the presence of missing data are well established in inference, typically relying on the Missing-at-Random (MAR) assumption. In prediction, similar principles are often assumed to apply.…

Methodology · Statistics 2026-03-19 Pierre Catoire , Robin Genuer , Cecile Proust-Lima

Several approaches have been proposed in the literature for clustering multivariate ordinal data. These methods typically treat missing values as absent information, rather than recognizing them as valuable for profiling population…

Methodology · Statistics 2024-11-05 Alice Giampino , Antonio Canale , Bernardo Nipoti

In observational studies, the causal effect of a treatment may be confounded with variables that are related to both the treatment and the outcome of interest. In order to identify a causal effect, such studies often rely on the…

Methodology · Statistics 2017-10-17 Emma Persson , Jenny Häggström , Ingeborg Waernbaum , Xavier de Luna

Score matching is a vital tool for learning the distribution of data with applications across many areas including diffusion processes, energy based modelling, and graphical model estimation. Despite all these applications, little work…

Machine Learning · Statistics 2025-06-03 Josh Givens , Song Liu , Henry W J Reeve

In the analysis of observational data in social sciences and businesses, it is difficult to obtain a "(quasi) single-source dataset" in which the variables of interest are simultaneously observed. Instead, multiple-source datasets are…

Methodology · Statistics 2021-09-02 Masaki Mitsuhiro , Takahiro Hoshino

In many fields, and especially in the medical and social sciences and in recommender systems, data are gathered through clinical studies or targeted surveys. Participants are generally reluctant to respond to all questions in a survey or…

Statistics Theory · Mathematics 2016-11-15 Mohammad Reza Gholami , Magnus Jansson , Erik G. Ström , Ali H. Sayed

Synthetic data has been proposed as a solution to address the issue of high-quality data scarcity in the training of large language models (LLMs). Studies have shown that synthetic data can effectively improve the performance of LLMs on…

Computation and Language · Computer Science 2024-06-19 Jie Chen , Yupeng Zhang , Bingning Wang , Wayne Xin Zhao , Ji-Rong Wen , Weipeng Chen

Multivariate time series data for real-world applications typically contain a significant amount of missing values. The dominant approach for classification with such missing values is to impute them heuristically with specific values…

Machine Learning · Computer Science 2023-08-15 SeungHyun Kim , Hyunsu Kim , EungGu Yun , Hwangrae Lee , Jaehun Lee , Juho Lee

We demonstrate a simple strategy to cope with missing data in sequential inputs, addressing the task of multilabel classification of diagnoses given clinical time series. Collected from the pediatric intensive care unit (PICU) at Children's…

Machine Learning · Computer Science 2016-11-14 Zachary C. Lipton , David C. Kale , Randall Wetzel