English
Related papers

Related papers: SHARELIFE Imputations

200 papers

The JWST has been collecting scientific data for over two years now. Scientists are now looking deeper into the data, which introduces the need to correct known systematic effects. Important limiting factors for the MIRI/MRS are the…

Instrumentation and Methods for Astrophysics · Physics 2024-06-18 Danny Gasman , Ioannis Argyriou , Jane E. Morrison , David R. Law , Alistair Glasse , Karl D. Gordon , Patrick J. Kavanagh , Craig Lage , Polychronis Patapis , G. C. Sloan

In spite of increased attention on explainable machine learning models, explaining multi-output predictions has not yet been extensively addressed. Methods that use Shapley values to attribute feature contributions to the decision making…

Machine Learning · Computer Science 2023-03-31 Célia Wafa Ayad , Thomas Bonnier , Benjamin Bosch , Jesse Read

Advancements in data collection techniques and the heterogeneity of data resources can yield high percentages of missing observations on variables, such as block-wise missing data. Under missing-data scenarios, traditional methods such as…

Methodology · Statistics 2022-05-17 Wei Lan , Xuerong Chen , Tao Zou , Chih-Ling Tsai

Datagaps are ubiquitous in real world observational data. Quantifying nonlinearity in data having gaps can be challenging. Reported research points out that interpolation can affect nonlinear quantifiers adversely, artificially introducing…

Chaotic Dynamics · Physics 2018-09-05 Sandip V. George , G. Ambika

Item non-response in surveys is usually handled by single imputation, whose main objective is to reduce the non-response bias. Imputation methods need to be adapted to the study variable. For instance, in business surveys, the interest…

Methodology · Statistics 2019-10-17 Brigitte Gelein , Guillaume Chauvet

Multi-task learning is frequently used to model a set of related response variables from the same set of features, improving predictive performance and modeling accuracy relative to methods that handle each response variable separately.…

Methodology · Statistics 2023-08-11 Snigdha Panigrahi , Natasha Stewart , Chandra Sekhar Sripada , Elizaveta Levina

Surrender poses one of the major risks to life insurance and a sound modeling of its true probability has direct implication on the risk capital demanded by the Solvency II directive. We add to the existing literature by performing…

Risk Management · Quantitative Finance 2021-08-30 Mark Kiermayer

Resampling techniques have become increasingly popular for estimation of uncertainty in data collected via surveys. Survey data are also frequently subject to missing data which are often imputed. This note addresses the issue of using…

Methodology · Statistics 2023-11-27 Michael W. Robbins , Lane Burgette , Sebastian Bauhoff

Missing data remains a very common problem in large datasets, including survey and census data containing many ordinal responses, such as political polls and opinion surveys. Multiple imputation (MI) is usually the go-to approach for…

Methodology · Statistics 2024-12-25 Chayut Wongkamthong , Olanrewaju Akande

Multivariate categorical data nested within households often include reported values that fail edit constraints---for example, a participating household reports a child's age as older than his biological parent's age---as well as missing…

Methodology · Statistics 2018-09-21 Olanrewaju Akande , Andrés Barrientos , Jerome P. Reiter

Attrition in survey and field experiments presents a challenge for social science research. Common approaches to deal with this problem -- such as complete case analysis, multiple imputation, and weighting methods -- rely on strong…

Methodology · Statistics 2026-04-13 Xiangyu Song

Mismatches between samples and their respective channel or target commonly arise in several real-world applications. For instance, whole-brain calcium imaging of freely moving organisms, multiple-target tracking or multi-person contactless…

Signal Processing · Electrical Eng. & Systems 2023-07-25 Taulant Koka , Manolis C. Tsakiris , Michael Muma , Benjamín Béjar Haro

Real-world clinical time series data sets exhibit a high prevalence of missing values. Hence, there is an increasing interest in missing data imputation. Traditional statistical approaches impose constraints on the data-generating process…

Machine Learning · Computer Science 2020-01-13 Yang Guo , Zhengyuan Liu , Pavitra Krishnswamy , Savitha Ramasamy

Pairwise comparison models have been widely used for utility evaluation and rank aggregation across various fields. The increasing scale of modern problems underscores the need to understand statistical inference in these models when the…

Statistics Theory · Mathematics 2025-12-16 Ruijian Han , Wenlu Tang , Yiming Xu

We consider functional data where an underlying smooth curve is composed not just with errors, but also with irregular spikes. We propose an approach that, combining regularized spline smoothing and an Expectation-Maximization algorithm,…

Methodology · Statistics 2023-07-18 Huy Dang , Marzia Cremona , Francesca Chiaromonte

Distributional data Shapley value (DShapley) has recently been proposed as a principled framework to quantify the contribution of individual datum in machine learning. DShapley develops the foundational game theory concept of Shapley values…

Machine Learning · Statistics 2021-02-19 Yongchan Kwon , Manuel A. Rivas , James Zou

High-quality labeled data are essential for reliable statistical inference, but are often limited by validation costs. While surrogate labels provide cost-effective alternatives, their noise can introduce non-negligible bias. To address…

Methodology · Statistics 2025-12-29 Jianmin Chen , Huiyuan Wang , Thomas Lumley , Xiaowu Dai , Yong Chen

We consider a causal inference model in which individuals interact in a social network and they may not comply with the assigned treatments. In particular, we suppose that the form of network interference is unknown to researchers. To…

Methodology · Statistics 2023-10-24 Tadao Hoshino , Takahide Yanagi

We propose a new class of univariate nonstationary time series models, using the framework of modulated time series, which is appropriate for the analysis of rapidly-evolving time series as well as time series observations with missing…

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

Methodology · Statistics 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu