English

Calculating $p$-values and their significances with the Energy Test for large datasets

Data Analysis, Statistics and Probability 2018-04-19 v2 High Energy Physics - Experiment

Abstract

The energy test method is a multi-dimensional test of whether two samples are consistent with arising from the same underlying population, through the calculation of a single test statistic (called the TT-value). The method has recently been used in particle physics to search for differences between samples that arise from CP violation. The generalised extreme value function has previously been used to describe the distribution of TT-values under the null hypothesis that the two samples are drawn from the same underlying population. We show that, in a simple test case, the distribution is not sufficiently well described by the generalised extreme value function. We present a new method, where the distribution of TT-values under the null hypothesis when comparing two large samples can be found by scaling the distribution found when comparing small samples drawn from the same population. This method can then be used to quickly calculate the pp-values associated with the results of the test.

Keywords

Cite

@article{arxiv.1801.05222,
  title  = {Calculating $p$-values and their significances with the Energy Test for large datasets},
  author = {W. Barter and C. Burr and C. Parkes},
  journal= {arXiv preprint arXiv:1801.05222},
  year   = {2018}
}

Comments

9 pages (including title page); 4 figures