English

Significance analysis and statistical mechanics: an application to clustering

Molecular Networks 2015-05-19 v1 Statistical Mechanics Quantitative Methods

Abstract

This paper addresses the statistical significance of structures in random data: Given a set of vectors and a measure of mutual similarity, how likely does a subset of these vectors form a cluster with enhanced similarity among its elements? The computation of this cluster p-value for randomly distributed vectors is mapped onto a well-defined problem of statistical mechanics. We solve this problem analytically, establishing a connection between the physics of quenched disorder and multiple testing statistics in clustering and related problems. In an application to gene expression data, we find a remarkable link between the statistical significance of a cluster and the functional relationships between its genes.

Keywords

Cite

@article{arxiv.1009.2470,
  title  = {Significance analysis and statistical mechanics: an application to clustering},
  author = {Marta Łuksza and Michael Lässig and Johannes Berg},
  journal= {arXiv preprint arXiv:1009.2470},
  year   = {2015}
}

Comments

to appear in Phys. Rev. Lett