English
Related papers

Related papers: A nonparametric HMM for genetic imputation and coa…

200 papers

It is very challenging to select informative features from tens of thousands of measured features in high-throughput data analysis. Recently, several parametric/regression models have been developed utilizing the gene network information to…

Applications · Statistics 2014-08-01 Yize Zhao , Jian Kang , Tianwei Yu

Dropout represents a typical issue to be addressed when dealing with longitudinal studies. If the mechanism leading to missing information is non-ignorable, inference based on the observed data only may be severely biased. A frequent…

Methodology · Statistics 2018-03-23 Maria Francesca Marino , Marco Alfo'

State-space models (SSMs) are commonly used to model time series data where the observations depend on an unobserved latent process. However, inference on the model parameters of an SSM can be challenging, especially when the likelihood of…

Computation · Statistics 2023-08-08 Mary Llewellyn , Ruth King , Víctor Elvira , Gordon Ross

Hidden Markov models (HMMs) are commonly used for disease progression modeling when the true patient health state is not fully known. Since HMMs typically have multiple local optima, incorporating additional patient covariates can improve…

Machine Learning · Statistics 2021-10-05 Matt Baucum , Anahita Khojandi , Theodore Papamarkou

Probabilistic programming uses programs to express generative models whose posterior probability is then computed by built-in inference engines. A challenging goal is to develop general purpose inference algorithms that work out-of-the-box…

Machine Learning · Computer Science 2022-11-03 Carol Mak , Fabian Zaiser , Luke Ong

Nonmonotone missing data arise routinely in empirical studies of social and health sciences, and when ignored, can induce selection bias and loss of efficiency. In practice, it is common to account for nonresponse under a missing-at-random…

Methodology · Statistics 2017-07-20 Eric J. Tchetgen Tchetgen , Linbo Wang , BaoLuo Sun

Stochastic gradient MCMC (SG-MCMC) algorithms have proven useful in scaling Bayesian inference to large datasets under an assumption of i.i.d data. We instead develop an SG-MCMC algorithm to learn the parameters of hidden Markov models…

Machine Learning · Statistics 2017-06-16 Yi-An Ma , Nicholas J. Foti , Emily B. Fox

We propose the segmented iHMM (siHMM), a hierarchical infinite hidden Markov model (iHMM) that supports a simple, efficient inference scheme. The siHMM is well suited to segmentation problems, where the goal is to identify points at which a…

Machine Learning · Statistics 2016-02-23 Ardavan Saeedi , Matthew Hoffman , Matthew Johnson , Ryan Adams

In this article, discrete and stochastic changes in (effective) population size are incorporated into the spectral representation of a biallelic diffusion process for drift and small mutation rates. A forward algorithm inspired by…

Populations and Evolution · Quantitative Biology 2024-03-11 Lynette Caitlin Mikula , Claus Vogl

We propose a probabilistic modeling framework for learning the dynamic patterns in the collective behaviors of social agents and developing profiles for different behavioral groups, using data collected from multiple information sources.…

Machine Learning · Statistics 2016-06-28 Lin Li , Ananthram Swami , Anna Scaglione

Linear Mixed Models (LMMs) are important tools in statistical genetics. When used for feature selection, they allow to find a sparse set of genetic traits that best predict a continuous phenotype of interest, while simultaneously correcting…

We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new method of nonparametric regression that accommodates continuous and categorical inputs, and responses that can be modeled by a generalized linear model. We…

Machine Learning · Statistics 2010-07-16 Lauren A. Hannah , David M. Blei , Warren B. Powell

Non-linear state space models are a widely-used class of models for biological, economic, and physical processes. Fitting these models to observed data is a difficult inference problem that has no straightforward solution. We take a…

Computation · Statistics 2013-05-03 Alexander Y. Shestopaloff , Radford M. Neal

We develop a structural framework for modeling and inferring unobserved heterogeneity in dynamic panel-data models. Unlike methods treating clustering as a descriptive device, we model heterogeneity as arising from a latent clustering…

Econometrics · Economics 2025-10-29 Jean-Pierre Florens , Anna Simoni

In the genomic era, the identification of gene signatures associated with disease is of significant interest. Such signatures are often used to predict clinical outcomes in new patients and aid clinical decision-making. However, recent…

Methodology · Statistics 2019-03-27 Naim U. Rashid , Quefeng Li , Jen Jen Yeh , Joseph G. Ibrahim

In public health management there is a need to produce subnational estimates of health outcomes. Often, however, funds are not available to collect samples large enough to produce traditional survey sample estimates for each subnational…

Applications · Statistics 2008-12-18 Donald Malec , Peter Müller

Hidden Markov models (HMMs) and their extensions have proven to be powerful tools for classification of observations that stem from systems with temporal dependence as they take into account that observations close in time are likely…

Applications · Statistics 2021-11-22 Sofia Ruiz-Suarez , Vianey Leos-Barajas , Juan Manuel Morales

In the field of population health research, understanding the similarities between geographical areas and quantifying their shared effects on health outcomes is crucial. In this paper, we synthesise a number of existing methods to create a…

Applications · Statistics 2023-11-27 Wala Draidi Areed , Aiden Price , Helen Thompson , Reid Malseed , Kerrie Mengersen

In genetic studies, haplotype data provide more refined information than data about separate genetic markers. However, large-scale studies that genotype hundreds to thousands of individuals may only provide results of pooled data, where…

Methodology · Statistics 2023-09-01 Yong See Foo , Jennifer A. Flegg

Handling incomplete and heterogeneous data remains a central challenge in real-world machine learning, where missing values may follow complex mechanisms (MCAR, MAR, MNAR) and features can be of mixed types (numerical and categorical).…

Machine Learning · Computer Science 2025-07-30 Youran Zhou , Mohamed Reda Bouadjenek , Jonathan Wells , Sunil Aryal