English

Necessary and sufficient conditions for identifiability in the admixture model

Statistics Theory 2022-03-16 v2 Statistics Theory

Abstract

We consider M SNP data from N individuals who are an admixture of K unknown ancient populations. Let Πsi\Pi_{si} be the frequency of the reference allele of individual i at SNP s. So the number of reference alleles at SNP s for a diploid individual is binomially distributed with parameters 2 and Πsi\Pi_{si}. We suppose Πsi=k=1KFskQki\Pi_{si}=\sum_{k=1}^KF_{sk}Q_{ki}, where FskF_{sk} is the allele frequency of SNP s in population k and QkiQ_{ki} is the proportion of population k in the ancestry of individual i. I am interested in the identifiability of F and Q, up to a relabelling of the ancient populations. Under what conditions, when Π=F1Q1=F2Q2\Pi =F^1Q^1=F^2Q^2 are F1F^1 and F2F^2 and Q1Q^1 and Q2Q^2 equal? I show that the anchor condition (Cabreros and Storey, 2019) on one matrix together with an independence condition on the other matrix is sufficient for identifiability. I will argue that the proof of the necessary condition in Cabreros and Storey, 2019 is incorrect, and I will provide a correct proof, which in addition does not require knowledge of the number of ancestral populations. I will also provide abstract necessary and sufficient conditions for identifiability. I will show that one cannot deviate substantially from the anchor condition without losing identifiability. Finally, I show necessary and sufficient conditions for identifiability for the non-admixed case.

Keywords

Cite

@article{arxiv.2202.05540,
  title  = {Necessary and sufficient conditions for identifiability in the admixture model},
  author = {Jan van Waaij},
  journal= {arXiv preprint arXiv:2202.05540},
  year   = {2022}
}