English

Using Ramsey theory to measure unavoidable spurious correlations in Big Data

Combinatorics 2018-03-06 v2

Abstract

Given a dataset we quantify how many patterns must always exist in the dataset. Formally this is done through the lens of Ramsey theory of graphs, and a quantitative bound known as Goodman's theorem. Combining statistical tools with Ramsey theory of graphs gives a nuanced understanding of how far away a dataset is from random, and what qualifies as a meaningful pattern. This method is applied to a dataset of repeated voters in the 1984 US congress, to quantify how homogeneous a subset of congressional voters is. We also measure how transitive a subset of voters is. Statistical Ramsey theory is also used with global economic trading data to provide evidence that global markets are quite transitive.

Keywords

Cite

@article{arxiv.1712.09471,
  title  = {Using Ramsey theory to measure unavoidable spurious correlations in Big Data},
  author = {Micheal Pawliuk and Michael Alexander Waddell},
  journal= {arXiv preprint arXiv:1712.09471},
  year   = {2018}
}

Comments

21 pages

R2 v1 2026-06-22T23:29:52.299Z