English
Related papers

Related papers: Variational Bayes latent class approach for EHR-ba…

200 papers

We propose a categorical matrix factorization method to infer latent diseases from electronic health records (EHR) data in an unsupervised manner. A latent disease is defined as an unknown biological aberration that causes a set of common…

Applications · Statistics 2019-02-15 Yang Ni , Peter Mueller , Yuan Ji

Bayesian methods have proved powerful in many applications for the inference of model parameters from data. These methods are based on Bayes' theorem, which itself is deceptively simple. However, in practice the computations required are…

Methodology · Statistics 2020-07-10 Michael A. Chappell , Mark W. Woolrich

In the following short article we adapt a new and popular machine learning model for inference on medical data sets. Our method is based on the Variational AutoEncoder (VAE) framework that we adapt to survival analysis on small data sets…

Machine Learning · Statistics 2018-12-06 Cédric Beaulac , Jeffrey S. Rosenthal , David Hodgson

Objectives: In the United States, 25% of people with type 2 diabetes are undiagnosed. Conventional screening models use limited demographic information to assess risk. We evaluated whether electronic health record (EHR) phenotyping could…

Quantitative Methods · Quantitative Biology 2015-01-13 Ariana E. Anderson , Wesley T. Kerr , April Thames , Tong Li , Jiayang Xiao , Mark S. Cohen

Multi-state models of cancer natural history are widely used for designing and evaluating cancer early detection strategies. Calibrating such models against longitudinal data from screened cohorts is challenging, especially when fitting…

Computation · Statistics 2025-08-14 Raphael Morsomme , Shannon Holloway , Marc Ryser , Jason Xu

Electronic health records (EHRs) form an invaluable resource for training clinical decision support systems. To leverage the potential of such systems in high-risk applications, we need large, structured tabular datasets on which we can…

Artificial Intelligence · Computer Science 2025-11-24 Paloma Rabaey , Adrick Tench , Stefan Heytens , Thomas Demeester

Due to the increasing adoption of electronic health records (EHR), large scale EHRs have become another rich data source for translational clinical research. Despite its potential, deriving generalizable knowledge from EHR data remains…

Machine Learning · Statistics 2023-06-01 Junwei Lu , Jin Yin , Tianxi Cai

The analysis of mixed data has been raising challenges in statistics and machine learning. One of two most prominent challenges is to develop new statistical techniques and methodologies to effectively handle mixed data by making the data…

Machine Learning · Computer Science 2017-08-21 Tu Dinh Nguyen , Truyen Tran , Dinh Phung , Svetha Venkatesh

Hierarchical parametric models consisting of observable and latent variables are widely used for unsupervised learning tasks. For example, a mixture model is a representative hierarchical model for clustering. From the statistical point of…

Machine Learning · Statistics 2014-01-24 Keisuke Yamazaki

Modern datasets are becoming heterogeneous. To this end, we present in this paper Mixed-Variate Restricted Boltzmann Machines for simultaneously modelling variables of multiple types and modalities, including binary and continuous…

Machine Learning · Statistics 2014-08-07 Truyen Tran , Dinh Phung , Svetha Venkatesh

In Bayesian analysis, the posterior follows from the data and a choice of a prior and a likelihood. One hopes that the posterior is robust to reasonable variation in the choice of prior and likelihood, since this choice is made by the…

Methodology · Statistics 2015-12-09 Ryan Giordano , Tamara Broderick , Michael Jordan

Collaborative filtering (CF) has been successfully employed by many modern recommender systems. Conventional CF-based methods use the user-item interaction data as the sole information source to recommend items to users. However, CF-based…

Machine Learning · Computer Science 2018-09-25 Kenan Cui , Xu Chen , Jiangchao Yao , Ya Zhang

Hierarchical learning models, such as mixture models and Bayesian networks, are widely employed for unsupervised learning tasks, such as clustering analysis. They consist of observable and hidden variables, which represent the given data…

Machine Learning · Statistics 2018-01-08 Keisuke Yamazaki

Approximate Bayesian inference for models with computationally expensive, black-box likelihoods poses a significant challenge, especially when the posterior distribution is complex. Many inference methods struggle to explore the parameter…

Machine Learning · Statistics 2025-11-11 Francesco Silvestrin , Chengkun Li , Luigi Acerbi

Mendelian Randomization (MR) is a popular method in epidemiology and genetics that uses genetic variation as instrumental variables for causal inference. Existing MR methods usually assume most genetic variants are valid instrumental…

Applications · Statistics 2022-06-15 Daniel Iong , Qingyuan Zhao , Yang Chen

Electronic health records (EHRs) are increasingly used for clinical and comparative effectiveness research, but suffer from missing data. Motivated by health services research on diabetes care, we seek to increase the quality of EHRs by…

Methodology · Statistics 2020-07-14 Yajuan Si , Mari Palta , Maureen Smith

Diabetes remains a significant health challenge globally, contributing to severe complications like kidney disease, vision loss, and heart issues. The application of machine learning (ML) in healthcare enables efficient and accurate disease…

Machine Learning · Computer Science 2025-05-13 Mahade Hasan , Farhana Yasmin

Model comparison is the cornerstone of theoretical progress in psychological research. Common practice overwhelmingly relies on tools that evaluate competing models by balancing in-sample descriptive adequacy against model flexibility, with…

Applications · Statistics 2021-10-11 Viet-Hung Dao , David Gunawan , Minh-Ngoc Tran , Robert Kohn , Guy E. Hawkins , Scott D. Brown

Clustering mixed-type data remains a major challenge in biomedical research to uncover clinically meaningful subgroups within heterogeneous patient populations. Most existing clustering methods impose restrictive assumptions like local…

Applications · Statistics 2026-04-23 Yueting Wang , Shu Wang , Jonathan G. Yabes , Chung-Chou H. Chang

Diabetes is a chronic disease with a significant global health burden, requiring multi-stakeholder collaboration for optimal management. Large language models (LLMs) have shown promise in various healthcare scenarios, but their…

Computation and Language · Computer Science 2025-03-14 Lai Wei , Zhen Ying , Muyang He , Yutong Chen , Qian Yang , Yanzhe Hong , Jiaping Lu , Kaipeng Zheng , Shaoting Zhang , Xiaoying Li , Weiran Huang , Ying Chen