English
Related papers

Related papers: Self-separated and self-connected models for media…

200 papers

Estimating long-term treatment effects has a wide range of applications in various domains. A key feature in this context is that collecting long-term outcomes typically involves a multi-stage process and is subject to monotone missing,…

Machine Learning · Computer Science 2025-04-29 Qinwei Yang , Ruocheng Guo , Shasha Han , Peng Wu

Background: Existing guidelines for handling missing data are generally not consistent with the goals of prediction modelling, where missing data can occur at any stage of the model pipeline. Multiple imputation (MI), often heralded as the…

Methodology · Statistics 2022-06-27 Rose Sisk , Matthew Sperrin , Niels Peek , Maarten van Smeden , Glen P. Martin

Missing data in multiple variables is a common issue. We investigate the applicability of the framework of graphical models for handling missing data to a complex longitudinal pharmacological study of children with HIV treated with an…

Methodology · Statistics 2025-02-12 Anastasiia Holovchak , Helen McIlleron , Paolo Denti , Michael Schomaker

Fairness-aware classification models have gained increasing attention in recent years as concerns grow on discrimination against some demographic groups. Most existing models require full knowledge of the sensitive features, which can be…

Machine Learning · Computer Science 2025-05-02 Kaiqi Jiang , Wenzhe Fan , Mao Li , Xinhua Zhang

Causal mediation analysis decomposes the total treatment effect into a portion operating through a hypothesized mediator and a residual direct portion. Identification of natural direct and indirect effects typically rests on the mediator…

Methodology · Statistics 2026-05-19 Yuki Ohnishi , Fan Li

Interventional effects have been proposed as a solution to the unidentifiability of natural (in)direct effects under mediator-outcome confounders affected by the exposure. Such confounders are an intrinsic characteristic of studies with…

Methodology · Statistics 2022-03-30 Iván Díaz , Nicholas Williams , Kara E. Rudolph

Causal discovery from data affected by unobserved variables is an important but difficult problem to solve. The effects that unobserved variables have on the relationships between observed variables are more complex in nonlinear cases than…

Machine Learning · Computer Science 2021-06-07 Takashi Nicholas Maeda , Shohei Shimizu

Missing data is a fundamental challenge in data science, significantly hindering analysis and decision-making across a wide range of disciplines, including healthcare, bioinformatics, social science, e-commerce, and industrial monitoring.…

Machine Learning · Statistics 2026-05-12 Jicong Fan

The recovery of causal effects in structural models with missing data often relies on $m$-graphs, which assume that missingness mechanisms do not directly influence substantive variables. Yet, in many real-world settings, missing data can…

Methodology · Statistics 2025-06-19 Johan de Aguas , Leonard Henckel , Johan Pensar , Guido Biele

Missing data may be disastrous for the identifiability of causal and statistical estimands. In graphical missing data models, colluders are dependence structures that have a special importance for identification considerations. It has been…

Methodology · Statistics 2024-07-04 Santtu Tikka , Juha Karvanen

In some contexts, mixture models can fit certain variables well at the expense of others in ways beyond the analyst's control. For example, when the data include some variables with non-trivial amounts of missing values, the mixture model…

Methodology · Statistics 2016-09-06 Maria DeYoreo , Jerome P. Reiter , D. Sunshine Hillygus

Intensive longitudinal data, characterized by frequent measurements across numerous time points, are increasingly common due to advances in wearable devices and mobile health technologies. We consider evaluating causal mediation pathways…

Methodology · Statistics 2025-06-26 Tianchen Qian

We consider identification of peer effects under peer group miss-specification. Two leading cases are missing data and peer group uncertainty. Missing data can take the form of some individuals being entirely absent from the data. The…

Econometrics · Economics 2022-05-03 Christiern Rose , Lizi Yu

When estimating heterogeneous treatment effects, missing outcome data can complicate treatment effect estimation, causing certain subgroups of the population to be poorly represented. In this work, we discuss this commonly overlooked…

Machine Learning · Statistics 2025-04-15 Matthew Pryce , Karla Diaz-Ordaz , Ruth H. Keogh , Stijn Vansteelandt

The process of generating data such as images is controlled by independent and unknown factors of variation. The retrieval of these variables has been studied extensively in the disentanglement, causal representation learning, and…

Machine Learning · Computer Science 2023-09-26 Gaël Gendron , Michael Witbrock , Gillian Dobbie

In the context of individual-level causal inference, we study the problem of predicting whether someone will respond or not to a treatment based on their features and past examples of features, treatment indicator (e.g., drug/no drug), and…

Machine Learning · Statistics 2019-06-04 Nathan Kallus

Causal representation learning seeks to recover latent factors that generate observational data through a mixing function. Needing assumptions on latent structures or relationships to achieve identifiability in general, prior works often…

Artificial Intelligence · Computer Science 2025-09-24 Kwonho Kim , Heejeong Nam , Inwoo Hwang , Sanghack Lee

The problem of missing data, usually absent incurated and competition-standard datasets, is an unfortunate reality for most machine learning models used in industry applications. Recent work has focused on understanding the nature and the…

Machine Learning · Computer Science 2022-01-25 Spyridon Mouselinos , Kyriakos Polymenakos , Antonis Nikitakis , Konstantinos Kyriakopoulos

Missing values in multivariate time series data can harm machine learning performance and introduce bias. These gaps arise from sensor malfunctions, blackouts, and human error and are typically addressed by data imputation. Previous work…

Machine Learning · Computer Science 2025-03-04 Mohammad Rafid Ul Islam , Prasad Tadepalli , Alan Fern

Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses on identifying features invariant across environments, improves model generalization by capturing…

Machine Learning · Statistics 2026-05-11 Yiran Jia , Jelena Bradic