English
Related papers

Related papers: The Mahalanobis kernel for heritability estimation…

200 papers

Multi-trait genome-wide association studies (GWAS) use multi-variate statistical methods to identify associations between genetic variants and multiple correlated traits simultaneously, and have higher statistical power than independent…

Genomics · Quantitative Biology 2022-02-10 Muhammad Ammar Malik , Adriaan-Alexander Ludl , Tom Michoel

Gaussian mixture models (GMM) are powerful parametric tools with many applications in machine learning and computer vision. Expectation maximization (EM) is the most popular algorithm for estimating the GMM parameters. However, EM…

Computer Vision and Pattern Recognition · Computer Science 2017-11-17 Soheil Kolouri , Gustavo K. Rohde , Heiko Hoffmann

Genome-Wide Association Studies (GWAS) face unique challenges in the era of big genomics data, particularly when dealing with ultra-high-dimensional datasets where the number of genetic features significantly exceeds the available samples.…

Genomics · Quantitative Biology 2023-12-27 Kexuan Li

Representing, comparing, and measuring the distance between probability distributions is a key task in computational statistics and machine learning. The choice of representation and the associated distance determine properties of the…

Machine Learning · Statistics 2026-02-26 Masha Naslidnyk

Linear mixed models (LMMs) are a powerful and established tool for studying genotype-phenotype relationships. A limiting assumption of LMMs is that the residuals are Gaussian distributed, a requirement that rarely holds in practice.…

Genomics · Quantitative Biology 2014-08-10 Nicolo Fusi , Christoph Lippert , Neil D. Lawrence , Oliver Stegle

The broad sense genetic heritability, which quantifies the total proportion of phenotypic variation in a population due to genetic factors, is crucial for understanding trait inheritance. While many existing methods focus on estimating…

Methodology · Statistics 2024-11-04 Olivia Bley , Elizabeth Lei , Andy Zhou , Xiaoxi Shen

Imbalanced regression arises when the target distribution is skewed, causing models to focus on dense regions and struggle with underrepresented (minority) samples. Despite its relevance across many applications, few methods have been…

Machine Learning · Computer Science 2025-08-05 Shayan Alahyari , Shiva Mehdipour Ghobadlou , Mike Domaratzki

Motivation: Genome-wide association studies (GWASs), which assay more than a million single nucleotide polymorphisms (SNPs) in thousands of individuals, have been widely used to identify genetic risk variants for complex diseases. However,…

Computational Engineering, Finance, and Science · Computer Science 2015-01-27 Ben Teng , Can Yang , Jiming Liu , Zhipeng Cai , Xiang Wan

Motivation: Genome-Wide Association Studies (GWAS) seek to identify causal genomic variants associated with rare human diseases. The classical statistical approach for detecting these variants is based on univariate hypothesis testing, with…

Methodology · Statistics 2018-10-22 Florent Guinot , Marie Szafranski , Christophe Ambroise , Franck Samson

While likelihood-based inference and its variants provide a statistically efficient and widely applicable approach to parametric inference, their application to models involving intractable likelihoods poses challenges. In this work, we…

Methodology · Statistics 2019-06-17 Francois-Xavier Briol , Alessandro Barp , Andrew B. Duncan , Mark Girolami

Despite successful use in a wide variety of disciplines for data analysis and prediction, machine learning (ML) methods suffer from a lack of understanding of the reliability of predictions due to the lack of transparency and black-box…

Materials Science · Physics 2023-04-04 Evan Askanazi , Ilya Grinberg

Genome-Wide Association Studies are typically conducted using linear models to find genetic variants associated with common diseases. In these studies, association testing is done on a variant-by-variant basis, possibly missing out on…

Genome-wide association studies (GWA studies or GWAS) investigate the relationships between genetic variants such as single-nucleotide polymorphisms (SNPs) and individual traits. Recently, incorporating biological priors together with…

Machine Learning · Statistics 2017-09-13 Tao Yang , Paul Thompson , Sihai Zhao , Jieping Ye

The generalized linear mixed model (GLMM) is widely used for analyzing correlated data, particularly in large-scale biomedical and social science applications. Scalable Bayesian inference for GLMMs is challenging because the marginal…

Computation · Statistics 2026-01-07 Samuel I. Berchuck , Youngsoo Baek , Felipe A. Medeiros , Andrea Agazzi

The problems of large-scale multiple testing are often encountered in modern scientific researches. Conventional multiple testing procedures usually suffer considerable loss of testing efficiency due to the lack of consideration of…

Methodology · Statistics 2022-12-21 Pengfei Wang , Zhaofeng Tian

In this paper, we consider unsupervised partitioning problems, such as clustering, image segmentation, video segmentation and other change-point detection problems. We focus on partitioning problems based explicitly or implicitly on the…

Machine Learning · Computer Science 2013-03-07 Rémi Lajugie , Sylvain Arlot , Francis Bach

While linear mixed model (LMM) has shown a competitive performance in correcting spurious associations raised by population stratification, family structures, and cryptic relatedness, more challenges are still to be addressed regarding the…

Machine Learning · Computer Science 2023-02-15 Wenting Ye , Xiang Liu , Tianwei Yue , Wenping Wang

Mixture regression models are powerful tools for capturing heterogeneous covariate-response relationships, yet classical finite mixtures and Bayesian nonparametric alternatives often suffer from instability or overestimation of clusters…

Methodology · Statistics 2025-12-19 Yuta Hayashida , Shonosuke Sugasawa

Datasets with extreme observations and/or heavy-tailed error distributions are commonly encountered and should be analyzed with careful consideration of these features from a statistical perspective. Small deviations from an assumed model,…

Methodology · Statistics 2023-01-12 Meadhbh O'Neill , Kevin Burke

Mahalanobis distance (MD) is a simple and popular post-processing method for detecting out-of-distribution (OOD) inputs in neural networks. We analyze its failure modes for near-OOD detection and propose a simple fix called relative…

Machine Learning · Computer Science 2021-06-18 Jie Ren , Stanislav Fort , Jeremiah Liu , Abhijit Guha Roy , Shreyas Padhy , Balaji Lakshminarayanan