English
Related papers

Related papers: Integrative Analysis and Imputation of Multiple Da…

200 papers

With the advancement of data science, the collection of increasingly complex datasets has become commonplace. In such datasets, the data dimension can be extremely high, and the underlying data generation process can be unknown and highly…

Machine Learning · Statistics 2024-03-29 Yaxin Fang , Faming Liang

Multifidelity models integrate data from multiple sources to produce a single approximator for the underlying process. Dense low-fidelity samples are used to reduce interpolation error, while sparse high-fidelity samples are used to…

Machine Learning · Statistics 2024-02-27 Viv Bone , Chris van der Heide , Kieran Mackle , Ingo H. J. Jahn , Peter M. Dower , Chris Manzie

When treatment policy estimands are of interest, clinical trials often attempt to collect patient data after intercurrent events (ICEs), although such data are often limited. Retrieved dropout imputation methods, which use pre-ICE and…

Methodology · Statistics 2026-03-31 Brendah Nansereko , Marcel Wolbers , James R. Carpenter , Jonathan W. Bartlett

A Gaussian process has been one of the important approaches for emulating computer simulations. However, the stationarity assumption for a Gaussian process and the intractability for large-scale dataset limit its availability in practice.…

Methodology · Statistics 2020-11-06 Chih-Li Sung , Benjamin Haaland , Youngdeok Hwang , Siyuan Lu

Electronic health records (EHRs) are increasingly used for clinical and comparative effectiveness research, but suffer from missing data. Motivated by health services research on diabetes care, we seek to increase the quality of EHRs by…

Methodology · Statistics 2020-07-14 Yajuan Si , Mari Palta , Maureen Smith

Gaussian processes (GPs) are pervasive in functional data analysis, machine learning, and spatial statistics for modeling complex dependencies. Modern scientific data sets are typically heterogeneous and often contain multiple known…

Methodology · Statistics 2021-10-19 Didong Li , Andrew Jones , Sudipto Banerjee , Barbara E. Engelhardt

Most recent network failure diagnosis systems focused on data center networks where complex measurement systems can be deployed to derive routing information and ensure network coverage in order to achieve accurate and fast fault…

Networking and Internet Architecture · Computer Science 2022-07-06 Yufeng Xin , Shih-Wen Fu , Anirban Mandal , Ryan Tanaka , Mats Rynge , Karan Vahi , Ewa Deelman

Generalized Estimation Equations (GEE) are a well-known method for the analysis of non-Gaussian longitudinal data. This method has computational simplicity and marginal parameter interpretation. However, in the presence of missing data, it…

Methodology · Statistics 2015-06-16 José Luiz P. da Silva , Enrico A. Colosimo , Fábio N. Demarqui

Gaussian processes are a powerful framework for uncertainty-aware function approximation and sequential decision-making. Unfortunately, their classical formulation does not scale gracefully to large amounts of data and modern hardware for…

Machine Learning · Computer Science 2025-07-10 Jihao Andreas Lin

Environmental, Social, and Governance (ESG) datasets are frequently plagued by significant data gaps, leading to inconsistencies in ESG ratings due to varying imputation methods. This paper explores the application of established machine…

Machine Learning · Computer Science 2024-07-30 Sergio Caprioli , Jacopo Foschi , Riccardo Crupi , Alessandro Sabatino

The episodic, irregular and asynchronous nature of medical data render them difficult substrates for standard machine learning algorithms. We would like to abstract away this difficulty for the class of time-stamped categorical variables…

Machine Learning · Statistics 2014-02-20 Thomas A. Lasko

Computer-Aided Diagnosis has shown stellar performance in providing accurate medical diagnoses across multiple testing modalities (medical images, electrophysiological signals, etc.). While this field has typically focused on fully…

Applications · Statistics 2020-10-21 Claire Donnat , Nina Miolane , Freddy Bunbury , Jack Kreindler

Commonly used methods to analyze incomplete longitudinal clinical trial data include complete case analysis (CC) and last observation carried forward (LOCF). However, such methods rest on strong assumptions, including missing completely at…

Statistics Theory · Mathematics 2007-06-13 Ivy Jansen , Caroline Beunckens , Geert Molenberghs , Geert Verbeke , Craig Mallinckrodt

Missing data are present in most real world problems and need careful handling to preserve the prediction accuracy and statistical consistency in the downstream analysis. As the gold standard of handling missing data, multiple imputation…

Machine Learning · Computer Science 2021-12-23 Zongyu Dai , Zhiqi Bu , Qi Long

Missing data is a commonly occurring problem in practice. Many imputation methods have been developed to fill in the missing entries. However, not all of them can scale to high-dimensional data, especially the multiple imputation…

Machine Learning · Computer Science 2023-03-21 Thu Nguyen , Hoang Thien Ly , Michael Alexander Riegler , Pål Halvorsen , Hugo L. Hammer

Inference on unknown quantities in dynamical systems via observational data is essential for providing meaningful insight, furnishing accurate predictions, enabling robust control, and establishing appropriate designs for future…

Methodology · Statistics 2018-02-06 M. Chung , M. Binois , R. B. Gramacy , D. J. Moquin , A. P. Smith , A. M. Smith

Incomplete data are common in real-world applications. Sensors fail, records are inconsistent, and datasets collected from different sources often differ in scale, sampling rate, and quality. These differences create missing values that…

Machine Learning · Computer Science 2025-12-08 Zalish Mahmud , Anantaa Kotal , Aritran Piplai

Missing data are frequently encountered in high-dimensional problems, but they are usually difficult to deal with using standard algorithms, such as the expectation-maximization (EM) algorithm and its variants. To tackle this difficulty,…

Methodology · Statistics 2018-02-08 Faming Liang , Bochao Jia , Jingnan Xue , Qizhai Li , Ye Luo

It is not unusual for a data analyst to encounter data sets distributed across several computers. This can happen for reasons such as privacy concerns, efficiency of likelihood evaluations, or just the sheer size of the whole data set. This…

Computation · Statistics 2018-05-22 Randy C. S. Lai , J. Hannig , Thomas C. M. Lee

Missing data are ubiquitous in the era of big data and, if inadequately handled, are known to lead to biased findings and have deleterious impact on data-driven decision makings. To mitigate its impact, many missing value imputation methods…

Machine Learning · Computer Science 2021-10-26 Yiliang Zhang , Qi Long