English
Related papers

Related papers: Risk Prediction with Imperfect Survival Outcome In…

200 papers

A straightforward application of semi-supervised machine learning to the problem of treatment effect estimation would be to consider data as "unlabeled" if treatment assignment and covariates are observed but outcomes are unobserved.…

Methodology · Statistics 2020-09-15 Andrew Herren , P. Richard Hahn

COVID-19 pandemic has brought to the fore epidemiological models which, though describing a wealth of behaviors, have previously received little attention in signal processing literature. In this work, a generalized time-varying…

Methodology · Statistics 2025-08-13 Barbara Pascal , Samuel Vaiter

Increasingly, medical research is dependent on data collected for non-research purposes, such as electronic health records data (EHR). EHR data and other large databases can be prone to measurement error in key exposures, and unadjusted…

Methodology · Statistics 2020-05-13 Kyunghee Han , Thomas Lumley , Bryan E. Shepherd , Pamela A. Shaw

Typically, electronic health record data are not collected towards a specific research question. Instead, they comprise numerous observations recruited at different ages, whose medical, environmental and oftentimes also genetic data are…

Methodology · Statistics 2024-11-19 Nir Keret , Malka Gorfine

In epidemiological studies, participants' disease status is often collected through self-reported outcomes in place of formal medical tests due to budget constraints. However, self-reported outcomes are often subject to measurement errors,…

Methodology · Statistics 2023-06-21 Yujie Wu , Molin Wang

Predictive analytics is increasingly used to guide decision-making in many applications. However, in practice, we often have limited data on the true predictive task of interest, and must instead rely on more abundant data on a…

Machine Learning · Statistics 2020-05-07 Hamsa Bastani

In epidemiological studies, it is common to analyze disease risk by categorizing continuous variables, such as calorie and nutrient intake, for interpretability. When the original continuous variable is contaminated with measurement errors,…

Methodology · Statistics 2025-11-11 Huali Zhao , Tianying Wang

We study semiparametric efficiency bounds and efficient estimation of parameters defined through general moment restrictions with missing data. Identification relies on auxiliary data containing information about the distribution of the…

Statistics Theory · Mathematics 2008-04-04 Xiaohong Chen , Han Hong , Alessandro Tarozzi

In this work we provide a simple estimation procedure for a general frailty model for analysis of prospective correlated failure times. Rigorous large-sample theory for the proposed estimators of both the regression coefficient vector and…

Statistics Theory · Mathematics 2007-06-13 Malka Gorfine , David M. Zucker , Li Hsu

Electronic health records (EHR) are widely used to study clinical decisions, yet unmeasured confounding remains a persistent challenge. Proxy variables offer a potential solution. In EHR data, clinicians already record many such…

Methodology · Statistics 2026-03-23 Haley Colgate Kottler , Amy Cochran

Interval-censored competing risks data arise when each study subject may experience an event or failure from one of several causes and the failure time is not observed exactly but rather known to lie in an interval between two successive…

Methodology · Statistics 2016-03-02 Lu Mao , D. Y. Lin , Donglin Zeng

Time-to-event endpoints are frequently used as outcomes in oncology and other disease areas where the outcome of interest may not be observed within a predetermined period. Although many analytical methods address the challenges of…

Methodology · Statistics 2026-04-14 Chen-Yen Lin , Susan Halabi , Taehwa Choi

Accurate estimates of microbial species abundances are needed to advance our understanding of the role that microbiomes play in human and environmental health. However, artificially constructed microbiomes demonstrate that intuitive…

Methodology · Statistics 2025-03-17 David S Clausen , Amy D Willis

Prognostication for lung cancer, a leading cause of mortality, remains a complex task, as it needs to quantify the associations of risk factors and health events spanning a patient's entire life. One challenge is that an individual's…

Machine Learning · Statistics 2025-08-28 Stephen Salerno , Yi Li

This article considers a semi-supervised classification setting on a Gaussian mixture model, where the data is not labeled strictly as usual, but instead with uncertain labels. Our main aim is to compute the Bayes risk for this model. We…

Machine Learning · Statistics 2024-03-28 Victor Leger , Romain Couillet

In this paper, we propose a new wrapper feature selection approach with partially labeled training examples where unlabeled observations are pseudo-labeled using the predictions of an initial classifier trained on the labeled training set.…

Machine Learning · Computer Science 2020-03-11 Vasilii Feofanov , Emilie Devijver , Massih-Reza Amini

The proliferation of early diagnostic technologies, including self-monitoring systems and wearables, coupled with the application of these technologies on large segments of healthy populations may significantly aggravate the problem of…

Machine Learning · Computer Science 2021-07-23 Anna Fedyukova , Douglas Pires , Daniel Capurro

Attribute reduction is one of the most important research topics in the theory of rough sets, and many rough sets-based attribute reduction methods have thus been presented. However, most of them are specifically designed for dealing with…

Artificial Intelligence · Computer Science 2021-01-26 Can Gao , Jie Zhoua , Duoqian Miao , Xiaodong Yue , Jun Wan

While the ICD code assignment problem has been widely studied, most works have focused on post-discharge document classification. Models for early forecasting of this information could be used for identifying health risks, suggesting…

Machine Learning · Computer Science 2025-08-18 Cindy Shih-Ting Huang , Clarence Boon Liang Ng , Marek Rei

Semicontinuous outcomes commonly arise in a wide variety of fields, such as insurance claims, healthcare expenditures, rainfall amounts, and alcohol consumption. Regression models, including Tobit, Tweedie, and two-part models, are widely…

Methodology · Statistics 2024-03-26 Lu Yang