English
Related papers

Related papers: A Framework for Understanding Selection Bias in Re…

200 papers

In studies that rely on data from electronic health records (EHRs), unstructured text data such as clinical progress notes offer a rich source of information about patient characteristics and care that may be missing from structured data.…

Computation and Language · Computer Science 2024-05-22 Reagan Mozer , Aaron R. Kaufman , Leo A. Celi , Luke Miratrix

The study of disparities in the liver transplantation process may focus on quantifying causal effects, particularly the average, direct, or indirect effects of various social determinants of health on being listed as a candidate for…

The use of patient-level information from previous studies, registries, and other external datasets can support the analysis of single-arm and randomized controlled trials to evaluate and test experimental treatments. However, the…

Methodology · Statistics 2025-10-23 Gopal Kotecha , Daniel E. Schwartz , Steffen Ventz , Lorenzo Trippa

The medical community believes binary medical event outcomes in EHR data contain sufficient information for making a sensible recommendation. However, there are two challenges to effectively utilizing such data: (1) modeling the…

Artificial Intelligence · Computer Science 2024-09-12 Xihao Piao , Pei Gao , Zheng Chen , Lingwei Zhu , Yasuko Matsubara , Yasushi Sakurai , Jimeng Sun

The inclusion of human sex and gender data in statistical analysis invokes multiple considerations for data collection, combination, analysis, and interpretation. These considerations are not unique to variables representing sex and gender.…

Applications · Statistics 2024-01-05 Suzanne Thornton , Rochelle E. Tractenberg

Increasingly large electronic health records (EHRs) provide an opportunity to algorithmically learn medical knowledge. In one prominent example, a causal health knowledge graph could learn relationships between diseases and symptoms and…

Applications · Statistics 2019-10-04 Irene Y. Chen , Monica Agrawal , Steven Horng , David Sontag

Machine learning models trained on real-world data may inadvertently make biased predictions that negatively impact marginalized communities. Reweighting, which assigns a weight to each data point used during model training, can mitigate…

Machine Learning · Computer Science 2026-03-20 Anil K. Saini , Jose Guadalupe Hernandez , Emily F. Wong , Debanshi Misra , Tiffani J. Bright , Jason H. Moore

Objectives: Electronic health records (EHRs) are only a first step in capturing and utilizing health-related data - the challenge is turning that data into useful information. Furthermore, EHRs are increasingly likely to include data…

Artificial Intelligence · Computer Science 2012-08-20 Casey Bennett , Tom Doub , Rebecca Selove

Machine learning models are increasingly used in critical decision-making applications. However, these models are susceptible to replicating or even amplifying bias present in real-world data. While there are various bias mitigation methods…

Machine Learning · Computer Science 2024-01-05 Shih-Chi Ma , Tatiana Ermakova , Benjamin Fabian

Motivated by two case studies using primary care records from the Clinical Practice Research Datalink, we describe statistical methods that facilitate the analysis of tall data, with very large numbers of observations. Our focus is on…

Methodology · Statistics 2018-05-14 Kirsty Rhodes , Rebecca Turner , Rupert Payne , Ian White

The increase in availability of longitudinal electronic health record (EHR) data is leading to improved understanding of diseases and discovery of novel phenotypes. The majority of clustering algorithms focus only on patient trajectories,…

Machine Learning · Computer Science 2021-11-12 Oliver Carr , Avelino Javer , Patrick Rockenschaub , Owen Parsons , Robert Dürichen

Over the years, several studies have demonstrated that there exist significant disparities in health indicators in the United States population across various groups. Healthcare expense is used as a proxy for health in algorithms that drive…

Machine Learning · Computer Science 2019-11-06 Moninder Singh , Karthikeyan Natesan Ramamurthy

Feature selection represents a measure to reduce the complexity of high-dimensional datasets and gain insights into the systematic variation in the data. This aspect is of specific importance in domains that rely on model interpretability,…

Machine Learning · Computer Science 2022-09-07 Anna Jenul , Stefan Schrunner , Jürgen Pilz , Oliver Tomic

Due to the widespread use of data-powered systems in our everyday lives, concepts like bias and fairness gained significant attention among researchers and practitioners, in both industry and academia. Such issues typically emerge from the…

Machine Learning · Computer Science 2023-05-18 Gianluca Demartini , Kevin Roitero , Stefano Mizzaro

The availability of large and deep electronic healthcare records (EHR) datasets has the potential to enable a better understanding of real-world patient journeys, and to identify novel subgroups of patients. ML-based aggregation of EHR data…

Machine Learning · Computer Science 2022-08-03 Owen Parsons , Nathan E Barlow , Janie Baxter , Karen Paraschin , Andrea Derix , Peter Hein , Robert Dürichen

Modern data is messy and high-dimensional, and it is often not clear a priori what are the right questions to ask. Instead, the analyst typically needs to use the data to search for interesting analyses to perform and hypotheses to test.…

Machine Learning · Statistics 2019-10-09 Daniel Russo , James Zou

Synthetic data is becoming increasingly integral in data-scarce fields such as medical imaging, serving as a substitute for real data. However, its inherent statistical characteristics can significantly impact downstream tasks, potentially…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Krishan Agyakari Raja Babu , Rachana Sathish , Mrunal Pattanaik , Rahul Venkataramani

The recent adoption of Electronic Health Records (EHRs) by health care providers has introduced an important source of data that provides detailed and highly specific insights into patient phenotypes over large cohorts. These datasets, in…

Feature selection is important in data representation and intelligent diagnosis. Elastic net is one of the most widely used feature selectors. However, the features selected are dependant on the training data, and their weights dedicated…

Machine Learning · Computer Science 2021-01-01 Shaode Yu , Haobo Chen , Hang Yu , Zhicheng Zhang , Xiaokun Liang , Wenjian Qin , Yaoqin Xie , Ping Shi

When evaluating the performance of clinical machine learning models, one must consider the deployment population. When the population of patients with observed labels is only a subset of the deployment population (label selection), standard…

Machine Learning · Computer Science 2022-09-20 Conor K. Corbin , Michael Baiocchi , Jonathan H. Chen