English

Accurate $p$-Value Calculation for Generalized Fisher's Combination Tests Under Dependence

Methodology 2020-03-04 v1 Computation

Abstract

Combining dependent tests of significance has broad applications but the pp-value calculation is challenging. Current moment-matching methods (e.g., Brown's approximation) for Fisher's combination test tend to significantly inflate the type I error rate at the level less than 0.05. It could lead to significant false discoveries in big data analyses. This paper provides several more accurate and computationally efficient pp-value calculation methods for a general family of Fisher type statistics, referred as the GFisher. The GFisher covers Fisher's combination, Good's statistic, Lancaster's statistic, weighted Z-score combination, etc. It allows a flexible weighting scheme, as well as an omnibus procedure that automatically adapts proper weights and degrees of freedom to a given data. The new pp-value calculation methods are based on novel ideas of moment-ratio matching and joint-distribution surrogating. Systematic simulations show that they are accurate under multivariate Gaussian, and robust under the generalized linear model and the multivariate tt-distribution, down to at least 10610^{-6} level. We illustrate the usefulness of the GFisher and the new pp-value calculation methods in analyzing both simulated and real data of gene-based SNP-set association studies in genetics. Relevant computation has been implemented into R package GFisherGFisher.

Cite

@article{arxiv.2003.01286,
  title  = {Accurate $p$-Value Calculation for Generalized Fisher's Combination Tests Under Dependence},
  author = {Hong Zhang and Zheyang Wu},
  journal= {arXiv preprint arXiv:2003.01286},
  year   = {2020}
}

Comments

53 pages, 17 figures

R2 v1 2026-06-23T14:01:26.789Z