English
Related papers

Related papers: Polygenic Risk Score in Africa Population: Progres…

200 papers

In comparative studies, such as in causal inference and clinical trials, balancing important covariates is often one of the most important concerns for both efficient and credible comparison. However, chance imbalance still exists in many…

Methodology · Statistics 2018-07-30 Yichen Qin , Yang Li , Wei Ma , Feifang Hu

The distribution of genetic polymorphisms in a population contains information about the mutation rate and the strength of natural selection at a locus. Here, we show that the Poisson Random Field (PRF) method of population-genetic…

Populations and Evolution · Quantitative Biology 2007-07-18 Michael M Desai , Joshua B. Plotkin

According to the WHO, in 2021, about 32% of pregnant women in sub-Saharan Africa were infected with malaria during pregnancy. Malaria infection during pregnancy can cause various adverse birth outcomes such as low birthweight. Over the past…

Propensity score methods are an important tool to help reduce confounding in non-experimental studies. Most propensity score methods assume that covariates are measured without error. However, covariates are often measured with error, which…

Methodology · Statistics 2017-06-08 Hwanhee Hong , David A. Aaby , Juned Siddique , Elizabeth A. Stuart

Sleep traits are shaped by genetic and environmental factors and may influence many health conditions. The All of Us Research Program, which includes EHR, physical measurements, genomic data, and wearable data across ancestry groups,…

Genomics · Quantitative Biology 2026-05-25 Jiheum Park , Stephanie Y. Shue , Rocio Barragan , Jeong Yun Yang , Tian Gu , Chin Hur , Marie-Pierre St-Onge

Background: Subgroup analyses are frequently conducted in randomized clinical trials to assess evidence of heterogeneous treatment effect across patient subpopulations. Although randomization balances covariates within subgroups in…

Methodology · Statistics 2021-05-27 Siyun Yang , Fan Li , Laine E. Thomas , Fan Li

This paper focuses on the Bayesian Network Propensity Score (BNPS), a novel approach for estimating treatment effects in observational studies characterized by unknown (and likely unbalanced) designs and complex dependency structures among…

Genetic programming (GP) is a commonly used approach to solve symbolic regression (SR) problems. Compared with the machine learning or deep learning methods that depend on the pre-defined model and the training dataset for solving SR…

Neural and Evolutionary Computing · Computer Science 2022-05-23 Baihe He , Qiang Lu , Qingyun Yang , Jake Luo , Zhiguang Wang

In low-resource settings, prevalence mapping relies on empirical prevalence data from a finite, often spatially sparse, set of surveys of communities within the region of interest, possibly supplemented by remotely sensed images that can…

Applications · Statistics 2015-05-27 Peter J. Diggle , Emanuele Giorgi

Randomized controlled trials (RCTs) provide strong internal validity compared with observational studies. However, selection bias threatens the external validity of randomized trials. Thus, RCT results may not apply to either broad public…

Methodology · Statistics 2017-04-26 Ziyue Chen , Eloise Kaizar

High-throughput shotgun sequence data makes it possible in principle to accurately estimate population genetic parameters without confounding by SNP ascertainment bias. One such statistic of interest is the proportion of heterozygous sites…

Populations and Evolution · Quantitative Biology 2012-12-18 Katarzyna Bryc , Nick Patterson , David Reich

To deal with non-linear relations between the predictors and the response, we can use transformations to make the data look linear or approximately linear. In practice, however, transformation methods may be ineffective, and it may be more…

Methodology · Statistics 2023-01-02 Mithun Kumar Acharjee , Kumer Pial Das

When assessing the causal effect of a binary exposure using observational data, confounder imbalance across exposure arms must be addressed. Matching methods, including propensity score-based matching, can be used to deconfound the causal…

Methodology · Statistics 2024-10-01 Ernesto Ulloa-Pérez , Marco Carone , Alex Luedtke

Iterative data generation and model re-training can effectively align large language models(LLMs) to human preferences. The process of data sampling is crucial, as it significantly influences the success of policy improvement. Repeated…

Computation and Language · Computer Science 2024-10-07 Hai Ye , Hwee Tou Ng

In this paper, we propose Forest-PLS, a feature selection method for analyzing policy effect heterogeneity in a more flexible and comprehensive manner than is typically available with conventional methods. In particular, our method is able…

Econometrics · Economics 2023-07-25 Maria Nareklishvili , Nicholas Polson , Vadim Sokolov

Disease-gene association through Genome-wide association study (GWAS) is an arduous task for researchers. Investigating single nucleotide polymorphisms (SNPs) that correlate with specific diseases needs statistical analysis of associations.…

Quantitative Methods · Quantitative Biology 2020-12-21 Sezin Kircali Ata , Min Wu , Yuan Fang , Le Ou-Yang , Chee Keong Kwoh , Xiao-Li Li

The majority of common diseases are influenced by multiple genetic and environmental factors such as Cancer. Even though uncovering the main causes of disease is deemed difficult due to the complexity of gene-gene and gene-environment…

Other Computer Science · Computer Science 2017-05-10 Layan Nahlawi

Mendelian randomization is the use of genetic variants to assess the existence of a causal relationship between a risk factor and an outcome of interest. Here, we focus on two-sample summary-data Mendelian randomization analyses with many…

Quantitative Methods · Quantitative Biology 2022-09-16 Apostolos Gkatzionis , Stephen Burgess , Paul J. Newcombe

Estimating racial disparities in loan-approval probabilities when race is unobserved is routinely required for fair lending compliance. In such cases, race probabilities-typically from Bayesian Improved Surname Geocoding (BISG)-stand in for…

Applications · Statistics 2025-11-20 Sam Fisher , Dmitry Lesnik , Tobias Schäfer

Sharing diverse genomic and other biomedical datasets is critical to advance scientific discoveries and their equitable translation to improve human health. However, data sharing remains challenging in the context of legacy datasets,…