English
Related papers

Related papers: Data Integration with High Dimensionality

200 papers

Machine learning models are often used to inform real world risk assessment tasks: predicting consumer default risk, predicting whether a person suffers from a serious illness, or predicting a person's risk to appear in court. Given…

Machine Learning · Computer Science 2023-06-27 Jamelle Watson-Daniels , David C. Parkes , Berk Ustun

Complex diseases are multifactorial traits caused by both genetic and environmental factors. They represent the most part of human diseases and include those with largest prevalence and mortality (cancer, heart disease, obesity, etc.).…

Quantitative Methods · Quantitative Biology 2009-05-08 Roberto Amato , Michele Pinelli , Daniel D'Andrea , Gennaro Miele , Mario Nicodemi , Giancarlo Raiconi , Sergio Cocozza

Information integration plays a pivotal role in biomedical studies by facilitating the combination and analysis of independent datasets from multiple studies, thereby uncovering valuable insights that might otherwise remain obscured due to…

Methodology · Statistics 2024-07-02 Chixiang Chen , Jia Liang , Elynn Chen , Ming Wang

Integrating data from multiple heterogeneous sources has become increasingly popular to achieve a large sample size and diverse study population. This paper reviews development in causal inference methods that combines multiple datasets…

Methodology · Statistics 2021-10-05 Xu Shi , Ziyang Pan , Wang Miao

A common problem in health research is that we have a large database with many variables measured on a large number of individuals. We are interested in measuring additional variables on a subsample; these measurements may be newly…

Methodology · Statistics 2022-03-22 Thomas Lumley , Tong Chen

In this paper we outline some mathematical questions that emerge from trying to "turn the scientific method into math". Specifically, we consider the problem of experiment planning (choosing the best experiment to do next) in explicit…

Information Theory · Computer Science 2012-10-18 Christopher J. Lee , Marc Harper

Valid statistical inference is challenging when the sample is subject to unknown selection bias. Data integration can be used to correct for selection bias when we have a parallel probability sample from the same population with some common…

Methodology · Statistics 2023-07-24 Zhonglei Wang , Shu Yang , Jae Kwang Kim

A common goal in modern biostatistics is to form a biomarker signature from high dimensional gene expression data that is predictive of some outcome of interest. After learning this biomarker signature, an important question to answer is…

Statistics Theory · Mathematics 2015-10-05 Samuel M. Gross , Jonathan Taylor , Robert Tibshirani

As we gain access to a greater depth and range of health-related information about individuals, three questions arise: (1) Can we build better models to predict individual-level risk of ill health? (2) How much data do we need to…

Machine Learning · Statistics 2021-04-27 Mark Green

Prediction problems often admit competing models that perform almost equally well. This effect challenges key assumptions in machine learning when competing models assign conflicting predictions. In this paper, we define predictive…

Machine Learning · Computer Science 2020-09-17 Charles T. Marx , Flavio du Pin Calmon , Berk Ustun

Genetic risk prediction is an important component of individualized medicine, but prediction accuracies remain low for many complex diseases. A fundamental limitation is the sample sizes of the studies on which the prediction algorithms are…

Methodology · Statistics 2017-06-20 Sihai Dave Zhao

We study high-dimensional two-sample mean comparison and address the curse of dimensionality through data-adaptive projections. Leveraging the low-dimensional and localized signal structures commonly seen in single-cell genomics data, our…

Methodology · Statistics 2025-06-12 Tianyu Zhang , Jing Lei , Kathryn Roeder

Missing data are often dealt with multiple imputation. A crucial part of the multiple imputation process is selecting sensible models to generate plausible values for incomplete data. A method based on posterior predictive checking is…

Computation · Statistics 2026-05-14 Mingyang Cai , Stef van Buuren , Gerko Vink

We consider a distributed learning setting where each agent/learner holds a specific parametric model and data source. The goal is to integrate information across a set of learners to enhance the prediction accuracy of a given learner. A…

Methodology · Statistics 2021-09-21 Jiaying Zhou , Jie Ding , Kean Ming Tan , Vahid Tarokh

A critical step in data analysis for many different types of experiments is the identification of features with theoretically defined shapes in N-dimensional datasets; examples of this process include finding peaks in multi-dimensional…

Data Analysis, Statistics and Probability · Physics 2022-08-25 Korak Kumar Ray , Anjali R. Verma , Ruben L. Gonzalez , Colin D. Kinz-Thompson

As clinical decision-making increasingly moves toward individualized and context-specific treatment recommendations, reliance on any single evidence source, randomized or observational, may be insufficient. Principled integration of…

As artificial intelligence and machine learning tools become more accessible, and scientists face new obstacles to data collection (e.g., rising costs, declining survey response rates), researchers increasingly use predictions from…

Machine Learning · Statistics 2025-12-08 Stephen Salerno , Kentaro Hoffman , Awan Afiaz , Anna Neufeld , Tyler H. McCormick , Jeffrey T. Leek

Symmetries are key properties of physical models and of experimental designs, but any proposed symmetry may or may not be realized in nature. In this paper, we introduce a practical and general method to test such suspected symmetries in…

High Energy Physics - Phenomenology · Physics 2022-08-25 Rupert Tombs , Christopher G. Lester

A system level view of cellular processes for human and several organisms can be cap- tured by analyzing molecular interaction networks. A molecular interaction network formed of differentially expressed genes and their interactions helps…

Molecular Networks · Quantitative Biology 2016-11-09 Jeethu V. Devasia , Priya Chandran

Machine learning models are often personalized with information that is protected, sensitive, self-reported, or costly to acquire. These models use information about people but do not facilitate nor inform their consent. Individuals cannot…

Machine Learning · Computer Science 2023-10-13 Hailey Joren , Chirag Nagpal , Katherine Heller , Berk Ustun