English
Related papers

Related papers: Powerful extreme phenotype sampling designs and sc…

200 papers

We introduce a statistical method that can reconstruct nonlinear genetic models (i.e., including epistasis, or gene-gene interactions) from phenotype-genotype (GWAS) data. The computational and data resource requirements are similar to…

Genomics · Quantitative Biology 2015-09-29 Chiu Man Ho , Stephen D. H. Hsu

We introduce a new formulation of structural causal models for extremes, called the extremal structural causal model (eSCM). Unlike conventional structural causal models, where randomness is governed by a probability distribution, eSCMs use…

Statistics Theory · Mathematics 2026-05-27 Shuyang Bai , Fei Fang , Tiandong Wang

In the past two decades, psychological science has experienced an unprecedented replicability crisis which uncovered several issues. Among others, statistical inference is too often viewed as an isolated procedure limited to the analysis of…

We introduce a novel class of graphical models, termed profile graphical models, that represent, within a single graph, how an external factor influences the dependence structure of a multivariate set of variables. This class is quite…

Methodology · Statistics 2026-03-31 Alejandra Avalos-Pacheco , Monia Lupparelli , Francesco C. Stingo

Mendelian randomization (MR) is a method of exploiting genetic variation to unbiasedly estimate a causal effect in presence of unmeasured confounding. MR is being widely used in epidemiology and other related areas of population science. In…

Applications · Statistics 2019-01-03 Qingyuan Zhao , Jingshu Wang , Gibran Hemani , Jack Bowden , Dylan S. Small

Genotype-to-phenotype mappings translate genotypic variations such as mutations into phenotypic changes. Neutrality is the observation that some mutations do not lead to phenotypic changes. Studying the search trajectories in genotypic and…

Populations and Evolution · Quantitative Biology 2023-06-26 Ting Hu , Gabriela Ochoa , Wolfgang Banzhaf

Archetypal analysis serves as an exploratory tool that interprets a collection of observations as convex combinations of pure (extreme) patterns. When these patterns correspond to actual observations within the sample, they are termed…

Methodology · Statistics 2026-01-12 Aleix Alcacer , Irene Epifanio

The delimitation of biological species, i.e., deciding which individuals belong to the same species and whether and how many different species are represented in a data set, is key to the conservation of biodiversity. Much existing work…

Populations and Evolution · Quantitative Biology 2025-12-15 Gabriele d'Angella , Christian Hennig

Variations in complex traits are influenced by multiple genetic variants, environmental risk factors, and their interactions. Though substantial progress has been made in identifying single genetic variants associated with complex traits,…

Genomics · Quantitative Biology 2025-08-22 Ming Li , Ruo-Sin Peng , Changshuai Wei , Qing Lu

A computationally simple genome-wide association study (GWAS) algorithm for estimating the main and epistatic effects of markers or single nucleotide polymorphisms (SNPs) is proposed. It is based on the intuitive assumption that changes of…

Quantitative Methods · Quantitative Biology 2017-08-08 Lev V. Utkin , Irina L. Utkina

The objective of a genome-wide association study (GWAS) is to associate subsequences of individuals' genomes to the observable characteristics called phenotypes (e.g., high blood pressure). Motivated by the GWAS problem, in this paper we…

Information Theory · Computer Science 2020-10-15 Behrooz Tahmasebi , Mohammad Ali Maddah-Ali , Seyed Abolfazl Motahari

In many transcriptomic studies, the correlation of genes might fluctuate with quantitative factors such as genetic ancestry. We propose a method that models the covariance between two variables to vary against a continuous covariate. For…

Methodology · Statistics 2021-05-03 Tae Hyun Kim , Dan Nicolae

Overlap, also known as positivity, is a key condition for causal treatment effect estimation. Many popular estimators suffer from high variance and become brittle when features differ strongly across treatment groups. This is especially…

Machine Learning · Statistics 2026-04-02 Oscar Clivio , Alexander D'Amour , Alexander Franks , David Bruns-Smith , Chris Holmes , Avi Feller

This issue includes six articles that develop and apply statistical methods for the analysis of gene sequencing data of different types. The methods are tailored to the different data types and, in each case, lead to biological insights not…

Applications · Statistics 2012-06-29 Karen Kafadar

We present two results about using allele-count (AC) burdens of rare SNPs discovered in a case-control sequencing study for prediction or validation in an external prospective study. When genotyping only the SNPs polymorphic in the sequence…

Applications · Statistics 2015-10-19 C. Ryan King , Paul J. Rathouz , Dan L. Nicolae

When primed with only a handful of training samples, very large, pretrained language models such as GPT-3 have shown competitive results when compared to fully-supervised, fine-tuned, large, pretrained language models. We demonstrate that…

Computation and Language · Computer Science 2022-03-04 Yao Lu , Max Bartolo , Alastair Moore , Sebastian Riedel , Pontus Stenetorp

Observed differences in mean phenotypic values across human groups have attracted renewed interest with the rise of large-scale genomic studies and polygenic risk prediction. However, the genetic basis of these differences is far more…

Populations and Evolution · Quantitative Biology 2026-05-25 Nicole Kleman , Meng Lin , Christopher R. Gignoux , Arslan A. Zaidi

In extreme value analysis, sensitivity of inference to the definition of extreme event is a paramount issue. Under the peaks-over-threshold (POT) approach, this translates directly into the need of fitting a Generalized Pareto distribution…

Methodology · Statistics 2020-09-01 Jessica Silva Lomba , Maria Isabel Fraga Alves

Mendelian randomization is a powerful tool for causal inference in observational studies. The two-sample summary-data design, which estimates genetic associations with exposures and outcomes in separate cohorts, is the most widely used…

Methodology · Statistics 2026-04-29 Dingke Tang , Xuming He , Shu Yang

The additive genetic effect is arguably the most important quantity inferred in animal and plant breeding analyses. The term effect indicates that it represents causal information, which is different from standard statistical concepts as…

Quantitative Methods · Quantitative Biology 2015-04-27 Bruno Dourado Valente , Gota Morota , Guilherme Jordao Magalhaes Rosa , Daniel Gianola , Kent Weigel