English
Related papers

Related papers: Integrating Misclassified EHR Outcomes with Valida…

200 papers

Missing data, inaccuracies in medication lists, and recording delays in electronic health records (EHR) are major limitations for target trial emulation (TTE), the process by which EHR data are used to retrospectively emulate a randomized…

Applications · Statistics 2025-01-16 Max Sunog , Colin Magdamo , Marie-Laure Charpignon , Mark Albers

Most machine learning classifiers give predictions for new examples accurately, yet without indicating how trustworthy predictions are. In the medical domain, this hampers their integration in decision support systems, which could be useful…

Machine Learning · Computer Science 2018-07-06 Telma Pereira , Sandra Cardoso , Dina Silva , Manuela Guerreiro , Alexandre de Mendonça , Sara C. Madeira

Accurate heterogeneous treatment effect (HTE) estimation is essential for personalized recommendations, making it important to evaluate and compare HTE estimators. Traditional assessment methods are inapplicable due to missing…

Methodology · Statistics 2024-12-30 Zijun Gao

In electronic health records (EHR) analysis, clustering patients according to patterns in their data is crucial for uncovering new subtypes of diseases. Existing medical literature often relies on classical hypothesis testing methods to…

Methodology · Statistics 2024-05-07 Zihan Zhu , Xin Gai , Anru R. Zhang

Healthcare datasets present many challenges to both machine learning and statistics as their data are typically heterogeneous, censored, high-dimensional and have missing information. Feature selection is often used to identify the…

Machine Learning · Computer Science 2022-07-06 Annette Spooner , Gelareh Mohammadi , Perminder S. Sachdev , Henry Brodaty , Arcot Sowmya

Kernel matching is a widely used technique for estimating treatment effects, particularly valuable in observational studies where randomized controlled trials are not feasible. While kernel-matching approaches have demonstrated practical…

Methodology · Statistics 2025-12-11 Chong Ding , Zheng Li , Hon Keung Tony Ng , Wei Gao

Electronic health records (EHRs) provide a powerful basis for predicting the onset of health outcomes. Yet EHRs primarily capture in-clinic events and miss aspects of daily behavior and lifestyle containing rich health information. Consumer…

Augmenting randomized controlled trials (RCTs) with external real-world data (RWD) has the potential to improve the finite sample efficiency of treatment effect estimators. We describe using adaptive targeted maximum likelihood estimation…

Methodology · Statistics 2025-01-30 Sky Qiu , Jens Tarp , Andrew Mertens , Mark van der Laan

Alzheimer's disease is a progressive, debilitating neurodegenerative disease that affects 50 million people globally. Despite this substantial health burden, available treatments for the disease are limited and its fundamental causes remain…

Machine Learning · Computer Science 2024-04-02 Matthew West , Colin Magdamo , Lily Cheng , Yingnan He , Sudeshna Das

In studies that rely on data from electronic health records (EHRs), unstructured text data such as clinical progress notes offer a rich source of information about patient characteristics and care that may be missing from structured data.…

Computation and Language · Computer Science 2024-05-22 Reagan Mozer , Aaron R. Kaufman , Leo A. Celi , Luke Miratrix

Outcome-dependent sampling designs are extensively utilized in various scientific disciplines, including epidemiology, ecology, and economics, with retrospective case-control studies being specific examples of such designs. Additionally, if…

Methodology · Statistics 2023-09-22 Min Zeng , Zeyang Jia , Zijian Sui , Jinfeng Xu , Hong Zhang

Using administrative patient-care data such as Electronic Health Records (EHR) and medical/ pharmaceutical claims for population-based scientific research has become increasingly common. With vast sample sizes leading to very small standard…

Methodology · Statistics 2023-08-21 Ritoban Kundu , Xu Shi , Jean Morrison , Jessica Barrett , Bhramar Mukherjee

Analyzing data from multiple sources offers valuable opportunities to improve the estimation efficiency of causal estimands. However, this analysis also poses many challenges due to population heterogeneity and data privacy constraints.…

Methodology · Statistics 2025-10-23 Rong Zhao , Jason Falvey , Xu Shi , Vernon M. Chinchilli , Chixiang Chen

Randomized controlled trials (RCTs) frequently utilize covariate-adaptive randomization (CAR) (e.g., stratified block randomization) and commonly suffer from imperfect compliance. This paper studies the identification and inference for the…

Econometrics · Economics 2025-05-02 Federico A. Bugni , Mengsi Gao , Filip Obradovic , Amilcar Velez

Randomized controlled trials are the standard method for estimating causal effects, ensuring sufficient statistical power and confidence through adequate sample sizes. However, achieving such sample sizes is often challenging. This study…

Methodology · Statistics 2025-03-28 Keisuke Hanada , Masahiro Kojima

Healthcare datasets often contain groups of highly correlated features, such as features from the same biological system. When feature selection is applied to these datasets to identify the most important features, the biases inherent in…

Machine Learning · Computer Science 2022-07-07 Annette Spooner , Gelareh Mohammadi , Perminder S. Sachdev , Henry Brodaty , Arcot Sowmya

We study how to efficiently estimate average treatment effects (ATEs) using adaptive experiments. In adaptive experiments, experimenters sequentially assign treatments to experimental units while updating treatment assignment probabilities…

Machine Learning · Statistics 2025-02-21 Masahiro Kato , Takuya Ishihara , Junya Honda , Yusuke Narita

In observational studies, covariates with substantial missing data are often omitted, despite their strong predictive capabilities. These excluded covariates are generally believed not to simultaneously affect both treatment and outcome,…

Methodology · Statistics 2024-02-23 Shanshan Luo , Mengchen Shi , Wei Li , Xueli Wang , Zhi Geng

The growing availability of observational databases like electronic health records (EHR) provides unprecedented opportunities for secondary use of such data in biomedical research. However, these data can be error-prone and need to be…

Methodology · Statistics 2024-05-28 Sarah C. Lotspeich , Gustavo G. C. Amorim , Pamela A. Shaw , Ran Tao , Bryan E. Shepherd

Data harmonization is the process by which an equivalence is developed between two variables measuring a common trait. Our problem is motivated by dementia research in which multiple tests are used in practice to measure the same underlying…

Methodology · Statistics 2021-10-13 Steven Wilkins-Reeves , Yen-Chi Chen , Kwun Chuen Gary Chan