English

Statistical comparison of Hidden Markov Models via Fragment Analysis

Methodology 2025-05-01 v1 Statistics Theory Statistics Theory

Abstract

Standard practice in Hidden Markov Model (HMM) selection favors the candidate with the highest full-sequence likelihood, although this is equivalent to making a decision based on a single realization. We introduce a \emph{fragment-based} framework that redefines model selection as a formal statistical comparison. For an unknown true model HMM0\mathrm{HMM}_0 and a candidate HMMj\mathrm{HMM}_j, let μj(r)\mu_j(r) denote the probability that HMMj\mathrm{HMM}_j and HMM0\mathrm{HMM}_0 generate the same sequence of length~rr. We show that if HMMi\mathrm{HMM}_i is closer to HMM0\mathrm{HMM}_0 than HMMj\mathrm{HMM}_j, there exists a threshold rr^{*} -- often small -- such that μi(r)>μj(r)\mu_i(r)>\mu_j(r) for all rrr\geq r^{*}. Sampling kk independent fragments yields unbiased estimators μ^j(r)\hat{\mu}_j(r) whose differences are asymptotically normal, enabling a straightforward ZZ-test for the hypothesis H0 ⁣:μi(r)=μj(r)H_0\!:\,\mu_i(r)=\mu_j(r). By evaluating only short subsequences, the procedure circumvents full-sequence likelihood computation and provides valid pp-values for model comparison.

Keywords

Cite

@article{arxiv.2504.21046,
  title  = {Statistical comparison of Hidden Markov Models via Fragment Analysis},
  author = {Carlos M. Hernandez-Suarez and Osval A. Montesinos-López},
  journal= {arXiv preprint arXiv:2504.21046},
  year   = {2025}
}

Comments

Eight pages, 1 figure