English
Related papers

Related papers: DOT: Gene-set analysis by combining decorrelated a…

200 papers

In many scientific problems, researchers try to relate a response variable $Y$ to a set of potential explanatory variables $X = (X_1,\dots,X_p)$, and start by trying to identify variables that contribute to this relationship. In statistical…

Statistics Theory · Mathematics 2020-10-07 Wenshuo Wang , Lucas Janson

Sequential multiple assignment randomized trials (SMARTs) are used to construct data-driven optimal intervention strategies for subjects based on their intervention and covariate histories in different branches of health and behavioral…

Methodology · Statistics 2022-04-28 Palash Ghosh , Xiaoxi Yan , Bibhas Chakraborty

This paper presents and compares alternative transfer learning methods that can increase the power of conditional testing via knockoffs by leveraging prior information in external data sets collected from different populations or measuring…

Applications · Statistics 2021-08-20 Shuangning Li , Zhimei Ren , Chiara Sabatti , Matteo Sesia

An important aspect of AI design and ethics is to create systems that reflect aggregate preferences of the society. To this end, the techniques of social choice theory are often utilized. We propose a new social choice function motivated by…

Multiagent Systems · Computer Science 2021-03-02 Gergei Bana , Wojciech Jamroga , David Naccache , Peter Y. A. Ryan

Consider a genetic locus carrying a strongly beneficial allele which has recently fixed in a large population. As strongly beneficial alleles fix quickly, sequence diversity at partially linked neutral loci is reduced. This phenomenon is…

Populations and Evolution · Quantitative Biology 2007-05-23 P. Pfaffelhuber , A. Studeny

As a living information and communications system, the genome encodes patterns in single nucleotide polymorphisms (SNPs) reflecting human adaption that optimizes population survival in differing environments. This paper mathematically…

Populations and Evolution · Quantitative Biology 2018-03-22 James Lindesay , Tshela E. Mason , William Hercules , Georgia M. Dunston

The aetiology of polygenic obesity is multifactorial, which indicates that life-style and environmental factors may influence multiples genes to aggravate this disorder. Several low-risk single nucleotide polymorphisms (SNPs) have been…

Genomics · Quantitative Biology 2018-08-27 Casimiro A. Curbelo Montañez , Paul Fergus , Carl Chalmers , Jade Hind

Online Multiple Testing (OMT), a fundamental pillar of sequential statistical inference, traditionally evaluates the False Discovery Rate (FDR) and statistical power in isolation, obscuring the highly asymmetric costs of false positives and…

Machine Learning · Statistics 2026-05-15 Qingyang Hao , Kongchang Zhou , Fang Kong , Hongxin Wei

The likelihood function represents statistical evidence in the context of data and a probability model. Considerable theory has demonstrated that evidence strength for different parameter values can be interpreted from the ratio of…

Applications · Statistics 2016-11-17 Zeynep Baskurt , Lisa Strug

The growing availability of large health databases has expanded the use of observational studies for comparative effectiveness research. Unlike randomized trials, observational studies must adjust for systematic differences in patient…

Methodology · Statistics 2026-01-21 Haidong Lu , Fan Li , Laine E. Thomas , Fan Li

Causal inference plays an important role in under standing the underlying mechanisation of the data generation process across various domains. It is challenging to estimate the average causal effect and individual causal effects from…

Data Structures and Algorithms · Computer Science 2023-01-05 Haoran Zhao , Yinghao Zhang , Debo Cheng , Chen Li , Zaiwen Feng

Quantitative analysis of large-scale data is often complicated by the presence of diverse subgroups, which reduce the accuracy of inferences they make on held-out data. To address the challenge of heterogeneous data analysis, we introduce…

Machine Learning · Computer Science 2021-09-01 Nazanin Alipourfard , Keith Burghardt , Kristina Lerman

Association testing aims to discover the underlying relationship between genotypes (usually Single Nucleotide Polymorphisms, or SNPs) and phenotypes (attributes, or traits). The typically large data sets used in association testing often…

Applications · Statistics 2012-07-04 Zhen Li , Vikneswaran Gopal , Xiaobo Li , John M. Davis , George Casella

We seek to identify genes involved in Parkinson's Disease (PD) by combining information across different experiment types. Each experiment, taken individually, may contain too little information to distinguish some important genes from…

In clinical trials, there is potential to improve precision and reduce the required sample size by appropriately adjusting for baseline variables in the statistical analysis. This is called covariate adjustment. Despite recommendations by…

Methodology · Statistics 2022-06-20 Kelly Van Lancker , Joshua Betz , Michael Rosenblum

Data depth has emerged as an invaluable nonparametric measure for the ranking of multivariate samples. The main contribution of depth-based two-sample comparisons is the introduction of the Q statistic (Liu and Singh, 1993), a quality…

Methodology · Statistics 2024-08-21 Yiting Chen , Min Gao , Wei Lin , Andrew Jirasek , Kirsty Milligan , Xiaoping Shi

The development of next generation sequencing (NGS) technology and genotype imputation methods enabled researchers to measure both common and rare variants in genome-wide association studies (GWAS). Statistical methods have been proposed to…

Methodology · Statistics 2018-12-14 XIaoyu Cai , Lo-Bin Chang , Chi Song

Standard approaches to tackle high-dimensional supervised classification problem often include variable selection and dimension reduction procedures. The novel methodology proposed in this paper combines clustering of variables and feature…

Statistics Theory · Mathematics 2018-11-07 Marie Chavent , Robin Genuer , Jerome Saracco

Predicting causal structure from time series data is crucial for understanding complex phenomena in physiology, brain connectivity, climate dynamics, and socio-economic behaviour. Causal discovery in time series is hindered by the…

Machine Learning · Computer Science 2026-01-06 Pedro P. Sanchez , Damian Machlanski , Steven McDonagh , Sotirios A. Tsaftaris

Traditional perturbative statistical disclosure control (SDC) approaches such as microaggregation, noise addition, rank swapping, etc, perturb the data in an ``ad-hoc" way in the sense that while they manage to preserve some particular…

Applications · Statistics 2023-11-14 Elias Chaibub Neto