English
Related papers

Related papers: Optimal Multi-Wave Validation of Secondary Use Dat…

200 papers

In this paper we introduce a binary search algorithm that efficiently finds initial maximum likelihood estimates for sequential experiments where a binary response is modeled by a continuous factor. The problem is motivated by switching…

Statistics Theory · Mathematics 2008-08-27 Juha Karvanen

In a sequential multiple-assignment randomized trial (SMART), a sequence of treatments is given to a patient over multiple stages. In each stage, randomization may be done to allocate patients to different treatment groups. Even though…

Methodology · Statistics 2024-01-09 Rik Ghosh , Bibhas Chakraborty , Inbal Nahum-Shani , Megan E. Patrick , Palash Ghosh

Validation studies are often used to obtain more reliable information in settings with error-prone data. Validated data on a subsample of subjects can be used together with error-prone data on all subjects to improve estimation. In…

Estimation of heterogeneous treatment effects is an essential component of precision medicine. Model and algorithm-based methods have been developed within the causal inference framework to achieve valid estimation and inference. Existing…

Methodology · Statistics 2021-05-10 Ruohong Li , Honglang Wang , Wanzhu Tu

Optimal experimental design (OED) is the general formalism of sensor placement and decisions about the data collection strategy for engineered or natural experiments. This approach is prevalent in many critical fields such as battery…

Optimization and Control · Mathematics 2022-06-28 Ahmed Attia , Emil Constantinescu

Extracting actionable insight from Electronic Health Records (EHRs) poses several challenges for traditional machine learning approaches. Patients are often missing data relative to each other; the data comes in a variety of modalities,…

Machine Learning · Computer Science 2018-11-13 Brandon Malone , Alberto Garcia-Duran , Mathias Niepert

Introduction: The discovery of causal mechanisms underlying diseases enables better diagnosis, prognosis and treatment selection. Clinical trials have been the gold standard for determining causality, but they are resource intensive,…

Machine Learning · Computer Science 2020-11-12 Xinpeng Shen , Sisi Ma , Prashanthi Vemuri , M. Regina Castro , Pedro J. Caraballo , Gyorgy J. Simon

Electronic Health Record (EHR) has emerged as a valuable source of data for translational research. To leverage EHR data for risk prediction and subsequently clinical decision support, clinical endpoints are often time to onset of a…

Methodology · Statistics 2023-11-07 Yang Wang , Qingning Zhou , Tianxi Cai , Xuan Wang

Labeling patients in electronic health records with respect to their statuses of having a disease or condition, i.e. case or control statuses, has increasingly relied on prediction models using high-dimensional variables derived from…

Methodology · Statistics 2021-10-14 Zijian Guo , Prabrisha Rakshit , Daniel S. Herman , Jinbo Chen

Detection of Out-of-Distribution (OOD) samples in real time is a crucial safety check for deployment of machine learning models in the medical field. Despite a growing number of uncertainty quantification techniques, there is a lack of…

Machine Learning · Computer Science 2022-05-09 Karina Zadorozhny , Patrick Thoral , Paul Elbers , Giovanni Cinà

This work revisits optimal response-adaptive designs from a type-I error rate perspective, highlighting when and how much these allocations exacerbate type-I error rate inflation - an issue previously undocumented. We explore a range of…

Methodology · Statistics 2025-09-09 Lukas Pin , Sofía S. Villar , William F. Rosenberger

The increasing prevalence of rich sources of data and the availability of electronic medical record databases and electronic registries opens tremendous opportunities for enhancing medical research. For example, controlled trials are…

Methodology · Statistics 2015-09-23 Liwen Ouyang , Daniel W. Apley , Sanjay Mehrotra

High-dimensional inference based on matrix-valued data has drawn increasing attention in modern statistical research, yet not much progress has been made in large-scale multiple testing specifically designed for analysing such data sets.…

Methodology · Statistics 2021-06-18 Xu Han , Sanat Sarkar , Shiyu Zhang

We study sequential testing for a binary disease outcome when risk follows an unknown logistic model. At each round, the decision maker may either pay for a test revealing the true label or predict the outcome based on patient features and…

Machine Learning · Computer Science 2026-05-05 Tavor Z. Baharav , Spyros Dragazis , Aldo Pacchiano

Risk prediction models are a crucial tool in healthcare. Risk prediction models with a binary outcome (i.e., binary classification models) are often constructed using methodology which assumes the costs of different classification errors…

Evaluating treatment effects is critical in clinical trials but sometimes involves lengthy, invasive, or costly follow-up procedures. In these cases, surrogate markers, which provide intermediate measures of the long-term treatment effect,…

Methodology · Statistics 2026-03-24 Sarah C. Lotspeich , P. D. Anh. Nguyen , Layla Parast

The robust development of Electronic Health Records (EHRs) causes a significant growth in sharing EHRs for clinical research. However, such a sharing makes it difficult to protect patient's privacy. A number of automated de-identification…

Cryptography and Security · Computer Science 2012-11-19 Jie Qian , Nafees Qamar

Randomized experiments (often known as "A/B tests") are widely used to evaluate product and service innovations. We study how to allocate limited experimentation resources across M concurrent experiments in an experiment-rich regime.…

Methodology · Statistics 2026-03-19 Fenghua Yang , Dae Woong Ham , Stefanus Jasin

Augmentation of disease diagnosis and decision-making in healthcare with machine learning algorithms is gaining much impetus in recent years. In particular, in the current epidemiological situation caused by COVID-19 pandemic, swift and…

Computers and Society · Computer Science 2021-02-23 Leopold Franz , Yash Raj Shrestha , Bibek Paudel

Traditional error detection approaches require user-defined parameters and rules. Thus, the user has to know both the error detection system and the data. However, we can also formulate error detection as a semi-supervised classification…

Machine Learning · Computer Science 2019-08-20 Felix Neutatz , Mohammad Mahdavi , Ziawasch Abedjan