English
Related papers

Related papers: The Mahalanobis kernel for heritability estimation…

200 papers

Genome-wide association studies (GWAS) require accurate cohort phenotyping, but expert labeling can be costly, time-intensive, and variable. Here we develop a machine learning (ML) model to predict glaucomatous optic nerve head features…

The delimitation of biological species, i.e., deciding which individuals belong to the same species and whether and how many different species are represented in a data set, is key to the conservation of biodiversity. Much existing work…

Populations and Evolution · Quantitative Biology 2025-12-15 Gabriele d'Angella , Christian Hennig

The kernel two-sample test based on the maximum mean discrepancy (MMD) is one of the most popular methods for detecting differences between two distributions over general metric spaces. In this paper we propose a method to boost the power…

Methodology · Statistics 2024-09-06 Anirban Chatterjee , Bhaswar B. Bhattacharya

Linear Mixed Model (LMM) is a common statistical approach to model the relation between exposure and outcome while capturing individual variability through random effects. However, this model assumes the homogeneity of the error term's…

Methodology · Statistics 2026-01-27 Vincent Jeanselme , Marco Palma , Jessica K Barrett

Genetic Gaussian network of multiple phenotypes constructed through the genetic correlation matrix is informative for understanding their biological dependencies. However, its interpretation may be challenging because the estimated genetic…

Methodology · Statistics 2024-12-31 Yihe Yang , Noah Lorincz-Comi , Xiaofeng Zhu

Mass-spectrometry technologies are widely used in the fields of ionomics and metabolomics to simultaneously profile at the genome scale intracellular concentrations of e.g. amino acids or elements. Short profiles of molecular or…

Molecular Networks · Quantitative Biology 2020-11-12 Jacopo Iacovacci , Alina Peluso , Timothy Ebbels , Markus Ralser , Robert Charles Glen

In many applications, data can be heterogeneous in the sense of spanning latent groups with different underlying distributions. When predictive models are applied to such data the heterogeneity can affect both predictive performance and…

Machine Learning · Statistics 2022-05-04 Thomas Lartigue , Sach Mukherjee

Mediation analysis is a powerful tool for studying causal pathways between exposure, mediator, and outcome variables of interest. While classical mediation analysis using observational data often requires strong and sometimes unrealistic…

Methodology · Statistics 2024-05-20 Rita Qiuran Lyu , Chong Wu , Xinwei Ma , Jingshen Wang

This report provides an exploration of different distance measures that can be used with the $K$-means algorithm for cluster analysis. Specifically, we investigate the Mahalanobis distance, and critically assess any benefits it may have…

Other Statistics · Statistics 2024-04-23 Zoe Shapcott

The Hirschfeld-Gebelein-R\'enyi (HGR) correlation coefficient is an extension of Pearson's correlation that is not limited to linear correlations, with potential applications in algorithmic fairness, scientific analysis, and causal…

Machine Learning · Computer Science 2025-09-12 Luca Giuliani , Michele Lombardi

Representational similarity analysis (RSA) tests models of brain computation by investigating how neural activity patterns reflect experimental conditions. Instead of predicting activity patterns directly, the models predict the geometry of…

Gaussian mixture models (GMMs) are widely used in machine learning for tasks such as clustering, classification, image reconstruction, and generative modeling. A key challenge in working with GMMs is defining a computationally efficient and…

Machine Learning · Computer Science 2025-08-05 Moritz Piening , Robert Beinert

Standard approaches to analysing data in genome-wide association studies (GWAS) ignore any potential functional relationships between genetic markers. In contrast gene pathways analysis uses prior information on functional structure within…

Methodology · Statistics 2013-02-26 M. Silver , P. Chen , L. Ruoying , C. Y. Cheng , T. Y. Wong , E. Tai , Y. Y. Teo , G. Montana

Recent technological advances coupled with large sample sets have uncovered many factors underlying the genetic basis of traits and the predisposition to complex disease, but much is left to discover. A common thread to most genetic…

Applications · Statistics 2013-12-11 Andrew Crossett , Ann B. Lee , Lambertus Klei , Bernie Devlin , Kathryn Roeder

In the linear mixed model (LMM), the simultaneous assessment and comparison of dispersion relevance of explanatory variables associated with fixed and random effects remains an important open practical problem. Based on the restricted…

Methodology · Statistics 2023-05-31 Nicholas Schreck , Manuel Wiesenfarth

Although Gaussian processes (GPs) with deep kernels have been successfully used for meta-learning in regression tasks, its uncertainty estimation performance can be poor. We propose a meta-learning method for calibrating deep kernel GPs for…

Machine Learning · Statistics 2023-12-14 Tomoharu Iwata , Atsutoshi Kumagai

High-cardinality categorical features are pervasive in actuarial data (e.g. occupation in commercial property insurance). Standard categorical encoding methods like one-hot encoding are inadequate in these settings. In this work, we present…

Machine Learning · Statistics 2024-11-20 Benjamin Avanzi , Greg Taylor , Melantha Wang , Bernard Wong

Cross-entropy loss has long been the standard choice for training deep neural networks, yet it suffers from interpretability limitations, unbounded weight growth, and inefficiencies that can contribute to costly training dynamics. The…

Machine Learning · Computer Science 2026-04-30 Maxwell Miller-Golub , Collin Coil , Kamil Faber , Marcin Pietron , Panpan Zheng , Pasquale Minervini , Roberto Corizzo

In the field of machine learning, model performance is usually assessed by randomly splitting data into training and test sets. Different random splits, however, can yield markedly different performance estimates, so a genuinely good model…

Uncertainty estimations for machine learning interatomic potentials (MLIPs) are crucial for quantifying model error and identifying informative training samples in active learning strategies. In this study, we evaluate uncertainty…

Machine Learning · Computer Science 2025-01-10 Matthias Holzenkamp , Dongyu Lyu , Ulrich Kleinekathöfer , Peter Zaspel