English
Related papers

Related papers: Using Volcano Plots and Regularized-Chi Statistics…

200 papers

For testing goodness of fit it is very popular to use either the chi square statistic or G statistics (information divergence). Asymptotically both are chi square distributed so an obvious question is which of the two statistics that has a…

Statistics Theory · Mathematics 2012-06-19 Peter Harremoës , Gábor Tusnády

We consider a single genetic locus which carries two alleles, labelled P and Q. This locus experiences selection and mutation. It is linked to a second neutral locus with recombination rate r. If r=0, this reduces to the study of a single…

Probability · Mathematics 2007-05-23 N. H. Barton , A. M. Etheridge , A. K. Sturm

Count data are collected in many scientific and engineering tasks including image processing, single-cell RNA sequencing and ecological studies. Such data sets often contain missing values, for example because some ecological sites cannot…

Methodology · Statistics 2018-10-25 Geneviève Robin , Julie Josse , Eric Moulines , Sylvain Sardy

Pearson's Chi-square test is a widely used tool for analyzing categorical data, yet its statistical power has remained theoretically underexplored. Due to the difficulties in obtaining its power function in the usual manner, Cochran (1952)…

Methodology · Statistics 2024-09-24 Qingyang Zhang

Meta-analysis seeks to combine the results of several experiments in order to improve the accuracy of decisions. It is common to use a test for homogeneity to determine if the results of the several experiments are sufficiently similar to…

Methodology · Statistics 2009-08-01 Elena Kulinskaya , Michael B. Dollinger , Kirsten Bjørkestøl

Parameter estimates for associated genetic variants, report ed in the initial discovery samples, are often grossly inflated compared to the values observed in the follow-up replication samples. This type of bias is a consequence of the…

Applications · Statistics 2011-04-15 Lizhen Xu , Radu V. Craiu , Lei Sun

Gene expression analysis aims at identifying the genes able to accurately predict biological parameters like, for example, disease subtyping or progression. While accurate prediction can be achieved by means of many different techniques,…

Methodology · Statistics 2008-09-11 Christine De Mol , Sofia Mosci , Magali Traskine , Alessandro Verri

The magnitude of Pearson correlation between two scalar random variables can be visually judged from the two-dimensional scatter plot of an independent and identically distributed sample drawn from the joint distribution of the two…

Methodology · Statistics 2023-03-14 Shanjun Mao , Xiaodan Fan , Jie Hu

Log-linear models are widely used to express the association in multivariate frequency data on contingency tables. The paper focuses on the power analysis for testing the goodness-of-fit hypothesis for this model type. Conventionally, for…

Methodology · Statistics 2024-05-06 Anna Klimova

We propose and study a fully efficient method to estimate associations of an exposure with disease incidence when both, incident cases and prevalent cases, i.e. individuals who were diagnosed with the disease at some prior time point and…

Methodology · Statistics 2018-03-20 Marlena Maziarz , Yukun Liu , Jing Qin , Ruth Pfeiffer

The vast majority of connections between complex disease and common genetic variants were identified through meta-analysis, a powerful approach that enables large samples sizes while protecting against common artifacts due to population…

High dimensional case control studies are ubiquitous in the biological sciences, particularly genomics. To maximise power while constraining cost and to minimise type-1 error rates, researchers typically seek to replicate findings in a…

Methodology · Statistics 2017-07-11 James Liley

Genome-wide association studies (GWAS) have identified thousands of genetic variants associated with complex traits, and some variants are shown to be associated with multiple complex traits. Genetic covariance between two traits is defined…

Methodology · Statistics 2023-10-06 Jianqiao Wang , Sai Li , Hongzhe Li

Prediction in high dimensional settings is difficult due to large by number of variables relative to the sample size. We demonstrate how auxiliary "co-data" can be used to improve the performance of a Random Forest in such a setting.…

Applications · Statistics 2017-06-05 Dennis E. te Beest , Steven W. Mes , Ruud H. Brakenhoff , Mark A. van de Wiel

Because biological processes can make different loci have different evolutionary histories, species tree estimation requires multiple loci from across the genome. While many processes can result in discord between gene trees and species…

Quantitative Methods · Quantitative Biology 2018-03-13 Md. Shamsuzzoha Bayzid , Siavash Mirarab , Bastien Boussau , Tandy Warnow

Additive genetic variance in natural populations is commonly estimated using mixed models, in which the covariance of the genetic effects is modeled by a genetic similarity matrix derived from a dense set of markers. An important but…

Applications · Statistics 2015-09-09 Willem Kruijer

The two-point correlation function has been the standard statistic for quantifying how galaxies are clustered. The statistic uses the positions of galaxies, but not their properties. Clustering as a function of galaxy property, be it type,…

Astrophysics · Physics 2007-05-23 Ravi K. Sheth , Andrew J. Connolly , Ramin Skibba

Software packages usually report the results of statistical tests using p-values. Users often interpret these by comparing them to standard thresholds, e.g. 0.1%, 1% and 5%, which is sometimes reinforced by a star rating (***, **, *). We…

Methodology · Statistics 2019-11-05 Axel Gandy , Georg Hahn , Dong Ding

In a case-control study aimed at localizing disease variants, association between a marker and the disease status is often tested by comparing the marker allele frequencies among cases and controls. These marker allele frequencies are…

Methodology · Statistics 2015-09-22 M. A. Jonker , M. W. T. Tanck

An important objective of experimental biology is the quantification of the relationship between predictor and response variables, a statistical analysis often termed variance partitioning (VP). In this paper, a series of simulations is…

Applications · Statistics 2019-11-27 Matthias M. Fischer