English
Related papers

Related papers: Effective Positive Cauchy Combination Test

200 papers

We introduce the notion of p*-values (p*-variables), which generalizes p-values (p-variables) in several senses. The new notion has four natural interpretations: operational, probabilistic, Bayesian, and frequentist. A main example of a…

Statistics Theory · Mathematics 2022-02-24 Ruodu Wang

Batch effects represent a major confounder in genomic diagnostics. In copy number variant (CNV) detection from NGS, many algorithms compare read depth between test samples and a reference sample, assuming they are process-matched. When this…

Genomics · Quantitative Biology 2026-01-16 Austin Talbot , Yue Ke

Conjoint analysis is a popular experimental design used to measure multidimensional preferences. Researchers examine how varying a factor of interest, while controlling for other relevant factors, influences decision-making. Currently,…

Methodology · Statistics 2024-11-20 Dae Woong Ham , Kosuke Imai , Lucas Janson

The paper discusses a test for the hypothesis that a random sample comes from the Cauchy distribution. The test statistics is derived from a characterization and is based on the characteristic function. Properties of the test are discussed…

Statistics Theory · Mathematics 2016-11-21 Emanuele Taufer

We formalize a problem we call combinatorial pair testing (CPT), which has applications to the identification of uncooperative or unproductive participants in pair programming, massively distributed computing, and crowdsourcing…

Data Structures and Algorithms · Computer Science 2013-05-02 David Eppstein , Michael T. Goodrich , Daniel S. Hirschberg

Null hypothesis statistical significance testing (NHST) is the dominant approach for evaluating results from randomized controlled trials. Whereas NHST comes with long-run error rate guarantees, its main inferential tool -- the $p$-value --…

Methodology · Statistics 2022-06-10 František Bartoš , Samuel Pawel , Eric-Jan Wagenmakers

In this paper, we present a novel approach for conformal prediction (CP), in which we aim to identify a set of promising prediction candidates -- in place of a single prediction. This set is guaranteed to contain a correct answer with high…

Machine Learning · Computer Science 2021-02-03 Adam Fisch , Tal Schuster , Tommi Jaakkola , Regina Barzilay

We develop a novel algorithm, Predictive Hierarchical Clustering (PHC), for agglomerative hierarchical clustering of current procedural terminology (CPT) codes. Our predictive hierarchical clustering aims to cluster subgroups, not…

Methodology · Statistics 2017-08-03 Elizabeth C. Lorenzi , Stephanie L. Brown , Zhifei Sun , Katherine Heller

Cross-validation plays a fundamental role in Machine Learning, enabling robust evaluation of model performance and preventing overestimation on training and validation data. However, one of its drawbacks is the potential to create data…

Machine Learning · Computer Science 2025-08-28 Afonso Martini Spezia , Thomas Fontanari , Mariana Recamonde-Mendoza

This paper considers testing linear hypotheses of a set of mean vectors with unequal covariance matrices in large dimensional setting. The problem of testing the hypothesis $H_0 : \sum_{i=1}^q \beta_i \bmu_i =\bmu_0 $ for a given vector…

Methodology · Statistics 2015-12-22 Dandan Jiang

$\textbf{Motivation:}$ Small $p$-values are often required to be accurately estimated in large-scale genomic studies for the adjustment of multiple hypothesis tests and the ranking of genomic features based on their statistical…

Applications · Statistics 2023-08-29 Yang Shi , Mengqiao Wang , Weiping Shi , Ji-Hyun Lee , Huining Kang , Hui Jiang

The higher criticism of a family of tests starts with the individual uncorrected p-values of each test. It then requires a procedure for deciding whether the collection of p-values indicates the presence of a real effect and if possible…

The increasing interest in subpopulation analysis has led to the development of various new trial designs and analysis methods in the fields of personalized medicine and targeted therapies. In this paper, subpopulations are defined in terms…

Methodology · Statistics 2020-12-01 Roland Gerard Gera , Tim Friede

In this paper, we compared the general forms of CCA and PLS on three simulated and two empirical datasets, all having large sample sizes. We took successively smaller subsamples of these data to evaluate sensitivity, reliability, and…

Methodology · Statistics 2022-06-28 Anthony R McIntosh

There is a vast body of literature related to methods for detecting changepoints (CP). However, less attention has been paid to assessing the statistical reliability of the detected CPs. In this paper, we introduce a novel method to perform…

Machine Learning · Statistics 2021-02-23 Vo Nguyen Le Duy , Hiroki Toda , Ryota Sugiyama , Ichiro Takeuchi

Higher criticism, or second-level significance testing, is a multiple-comparisons concept mentioned in passing by Tukey. It concerns a situation where there are many independent tests of significance and one is interested in rejecting the…

Statistics Theory · Mathematics 2007-06-13 David Donoho , Jiashun Jin

Researchers in genetics and other life sciences commonly use permutation tests to evaluate differences between groups. Permutation tests have desirable properties, including exactness if data are exchangeable, and are applicable even when…

Computation · Statistics 2018-11-01 Brian Segal , Thomas Braun , Michael Elliott , Hui Jiang

In this article, we propose a new class of consistent tests for $p$-variate normality. These tests are based on the characterization of the standard multivariate normal distribution, that the Hessian of the corresponding cumulant generating…

Methodology · Statistics 2023-03-22 Kwun Chuen Gary Chan , Hok Kan Ling , Chuan-Fa Tang , Sheung Chi Phillip Yam

Randomized clinical trials (RCTs) are widely considered the gold standard for evaluating the effectiveness of new treatments or interventions in drug development. Still, they may not be feasible in certain cases, such as with rare diseases…

Methodology · Statistics 2025-08-05 Di Ran , Fanni Zhang , Sima Shahsavari , Kristine Broglio , Alasdair Henderson , Binbing Yu

Conditional independence tests (CIT) are widely used for causal discovery and feature selection. Even with false discovery rate (FDR) control procedures, they often fail to provide frequentist guarantees in practice. We highlight two common…

Methodology · Statistics 2026-02-25 Milleno Pan , Antoine de Mathelin , Wesley Tansey