English
Related papers

Related papers: Correcting for Missing Data When Evaluating Surrog…

200 papers

Semi-supervised learning is a powerful technique for leveraging unlabeled data to improve machine learning models, but it can be affected by the presence of ``informative'' labels, which occur when some classes are more likely to be labeled…

Machine Learning · Statistics 2023-02-16 Aude Sportisse , Hugo Schmutz , Olivier Humbert , Charles Bouveyron , Pierre-Alexandre Mattei

The Binary Emax model is widely employed in dose-response analysis during drug development, where missing data often pose significant challenges. Addressing nonignorable missing binary responses, where the likelihood of missing data is…

Methodology · Statistics 2026-01-01 Jiangshan Zhang , Vivek Pradhan , Yuxi Zhao

This research was motivated by studying anti-drug antibody (ADA) formation and its potential impact on long-term benefit of a biologic treatment in a randomized controlled trial, in which ADA status was not only unobserved in the control…

Methodology · Statistics 2021-01-13 Shengchun Kong , Dominik Heinzmann , Sabine Lauer , Tian Lu

An essential goal of program evaluation and scientific research is the investigation of causal mechanisms. Over the past several decades, causal mediation analysis has been used in medical and social sciences to decompose the treatment…

Methodology · Statistics 2016-01-15 K. C. G. Chan , K. Imai , S. C. P. Yam , Z. Zhang

The method of surrogate data is a tool to test whether data were generated by some class of model. Tests based on the periodogram have been proposed to decide if linear systems driven by Gaussian noise could have generated a sample time…

comp-gas · Physics 2008-02-03 J. Timmer

Nonmonotone missing data arise routinely in empirical studies of social and health sciences, and when ignored, can induce selection bias and loss of efficiency. In practice, it is common to account for nonresponse under a missing-at-random…

Methodology · Statistics 2017-07-20 Eric J. Tchetgen Tchetgen , Linbo Wang , BaoLuo Sun

Anomaly detection in time-series has a wide range of practical applications. While numerous anomaly detection methods have been proposed in the literature, a recent survey concluded that no single method is the most accurate across various…

Machine Learning · Computer Science 2023-03-14 Mononito Goswami , Cristian Challu , Laurent Callot , Lenon Minorics , Andrey Kan

Causally interpretable meta-analysis combines information from a collection of randomized controlled trials to estimate treatment effects in a target population in which experimentation may not be possible but covariate information can be…

Methodology · Statistics 2022-05-03 Jon A. Steingrimsson , David H. Barker , Ruofan Bie , Issa J. Dahabreh

Difficulties may arise when analyzing longitudinal data using mixed-effects models if there are nonparametric functions present in the linear predictor component. This study extends the use of semiparametric mixed-effects modeling in cases…

Methodology · Statistics 2024-02-05 Mozhgan Taavoni , Mohammad Arashi

Practical problems with missing data are common, and statistical methods have been developed concerning the validity and/or efficiency of statistical procedures. On a central focus, there have been longstanding interests on the mechanism…

Methodology · Statistics 2020-03-26 Rui Duan , C. Jason Liang , Pamela Shaw , Cheng Yong Tang , Yong Chen

The growing availability of observational databases like electronic health records (EHR) provides unprecedented opportunities for secondary use of such data in biomedical research. However, these data can be error-prone and need to be…

Methodology · Statistics 2024-05-28 Sarah C. Lotspeich , Gustavo G. C. Amorim , Pamela A. Shaw , Ran Tao , Bryan E. Shepherd

The win ratio (WR) is a widely used metric to compare treatments in randomized clinical trials with hierarchically ordered endpoints. Counting-based approaches, such as Pocock's algorithm, are the standard for WR estimation. However, this…

Methodology · Statistics 2026-02-17 Yi Liu , Huiman Barnhart , Sean O'Brien , Yuliya Lokhnygina , Roland A. Matsouaka

Modern datasets commonly feature both substantial missingness and many variables of mixed data types, which present significant challenges for estimation and inference. Complete case analysis, which proceeds using only the observations with…

Methodology · Statistics 2023-04-10 Joseph Feldman , Daniel R. Kowal

This paper proposes a fast and accurate method for sparse regression in the presence of missing data. The underlying statistical model encapsulates the low-dimensional structure of the incomplete data matrix and the sparsity of the…

Machine Learning · Statistics 2015-03-31 Ravi Ganti , Rebecca M. Willett

Covariate-adaptive randomization is widely used in clinical trials to balance prognostic factors, and regression adjustments are often adopted to further enhance the estimation and inference efficiency. In practice, the covariates may…

Methodology · Statistics 2025-08-15 Wanjia Fu , Yingying Ma , Hanzhong Liu

Standard tests for nonlinearity reject the null hypothesis of a Gaussian linear process whenever the data is non-stationary. Thus, they are not appropriate to distinguish nonlinearity from non-stationarity. We address the problem of…

chao-dyn · Physics 2007-05-23 Andreas Schmitz , Thomas Schreiber

Recent work has focused on nonparametric estimation of conditional treatment effects, but inference has remained relatively unexplored. We propose a class of nonparametric tests for both quantitative and qualitative treatment effect…

Methodology · Statistics 2026-04-07 Oliver Dukes , Mats J. Stensrud , Riccardo Brioschi , Aaron Hudson

When devising a course of treatment for a patient, doctors often have little quantitative evidence on which to base their decisions, beyond their medical education and published clinical trials. Stanford Health Care alone has millions of…

An endeavor central to precision medicine is predictive biomarker discovery; they define patient subpopulations which stand to benefit most, or least, from a given treatment. The identification of these biomarkers is often the byproduct of…

Methodology · Statistics 2022-08-01 Philippe Boileau , Nina Ting Qi , Mark J. van der Laan , Sandrine Dudoit , Ning Leng

Missing values challenge data analysis because many supervised and unsupervised learning methods cannot be applied directly to incomplete data. Matrix completion based on low-rank assumptions are very powerful solution for dealing with…

Machine Learning · Statistics 2020-01-30 Aude Sportisse , Claire Boyer , Julie Josse