中文
相关论文

相关论文: Rank-based Bayesian clustering via covariate-infor…

200 篇论文

We consider the Bayesian mixture of finite mixtures (MFMs) and Dirichlet process mixture (DPM) models for clustering. Recent asymptotic theory has established that DPMs overestimate the number of clusters for large samples and that…

统计方法学 · 统计学 2022-08-01 Yannis Chaumeny , Johan van der Molen Moris , Anthony C. Davison , Paul D. W. Kirk

Generalized linear mixed models (GLMM) are used for inference and prediction in a wide range of different applications providing a powerful scientific tool. An increasing number of sources of data are becoming available, introducing a…

统计计算 · 统计学 2019-03-19 Aliaksandr Hubin , Geir Storvik

In social sciences, studies are often based on questionnaires asking participants to express ordered responses several times over a study period. We present a model-based clustering algorithm for such longitudinal ordinal data. Assuming…

统计方法学 · 统计学 2024-01-29 Francesco Amato , Julien Jacques , Isabelle Prim-Allaz

Data in the form of ranking lists are frequently encountered, and combining ranking results from different sources can potentially generate a better ranking list and help understand behaviors of the rankers. Of interest here are the rank…

统计方法学 · 统计学 2020-07-20 Xinran Li , Dingdong Yi , Jun S. Liu

The FBMS R package facilitates Bayesian model selection and model averaging in complex regression settings by employing a variety of Monte Carlo model exploration methods. At its core, the package implements an efficient Mode Jumping Markov…

统计方法学 · 统计学 2025-09-03 Florian Frommlet , Jon Lachmann , Geir Storvik , Aliaksandr Hubin

Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying…

机器学习 · 统计学 2008-03-26 Benhuai Xie , Wei Pan , Xiaotong Shen

The Gaussian mixture model (GMM) provides a simple yet principled framework for clustering, with properties suitable for statistical inference. In this paper, we propose a new model-based clustering algorithm, called EGMM (evidential GMM),…

机器学习 · 计算机科学 2022-11-29 Lianmeng Jiao , Thierry Denoeux , Zhun-ga Liu , Quan Pan

Given a set of pairwise comparisons, the classical ranking problem computes a single ranking that best represents the preferences of all users. In this paper, we study the problem of inferring individual preferences, arising in the context…

机器学习 · 统计学 2015-12-18 Rui Wu , Jiaming Xu , R. Srikant , Laurent Massoulié , Marc Lelarge , Bruce Hajek

In this paper, we propose a general framework for combining evidence of varying quality to estimate underlying binary latent variables in the presence of restrictions imposed to respect the scientific context. The resulting algorithms…

统计方法学 · 统计学 2018-08-28 Zhenke Wu , Livia Casciola-Rosen , Antony Rosen , Scott L. Zeger

We analyze the generalized Mallows model, a popular exponential model over rankings. Estimating the central (or consensus) ranking from data is NP-hard. We obtain the following new results: (1) We show that search methods can estimate both…

机器学习 · 计算机科学 2012-06-26 Marina Meila , Kapil Phadnis , Arthur Patterson , Jeff A. Bilmes

Healthcare cost prediction is a challenging task due to the high-dimensionality and high correlation among covariates. Additionally, the skewed, heavy-tailed, and often multi-modal nature of cost data can complicate matters further due to…

统计方法学 · 统计学 2023-03-13 Zhengxiao Li , Yifan Huang , Yang Cao

High-dimensional variable selection, with many more covariates than observations, is widely documented in standard regression models, but there are still few tools to address it in non-linear mixed-effects models where data are collected…

Boosting techniques from the field of statistical learning have grown to be a popular tool for estimating and selecting predictor effects in various regression models and can roughly be separated in two general approaches, namely gradient…

统计方法学 · 统计学 2019-12-16 Colin Griesbach , Andreas Groll , Elisabeth Waldmann

Personalized product search aims to retrieve and rank items that match users' preferences and search intent. Despite their effectiveness, existing approaches typically assume that users' query fully captures their real motivation. However,…

信息检索 · 计算机科学 2025-05-20 Weicong Qin , Yi Xu , Weijie Yu , Chenglei Shen , Ming He , Jianping Fan , Xiao Zhang , Jun Xu

The parsimonious Gaussian mixture models, which exploit an eigenvalue decomposition of the group covariance matrices of the Gaussian mixture, have shown their success in particular in cluster analysis. Their estimation is in general…

机器学习 · 统计学 2018-10-18 Faicel Chamroukhi , Marius Bartcus , Hervé Glotin

We present the Bayesian Case Model (BCM), a general framework for Bayesian case-based reasoning (CBR) and prototype classification and clustering. BCM brings the intuitive power of CBR to a Bayesian generative framework. The BCM learns…

机器学习 · 统计学 2019-04-05 Been Kim , Cynthia Rudin , Julie Shah

We study the secretary problem in which rank-ordered lists are generated by the Mallows model and the goal is to identify the highest-ranked candidate through a sequential interview process which does not allow rejected candidates to be…

统计方法学 · 统计学 2023-03-03 Xujun Liu , Olgica Milenkovic , George V. Moustakides

The results from Genome-Wide Association Studies (GWAS) on thousands of phenotypes provide an unprecedented opportunity to infer the causal effect of one phenotype (exposure) on another (outcome). Mendelian randomization (MR), an…

统计方法学 · 统计学 2019-04-30 Jia Zhao , Jingsi Ming , Xianghong Hu , Gang Chen , Jin Liu , Can Yang

Survey data are often collected under multistage sampling designs where units are binned to clusters that are sampled in a first stage. The unit-indexed population variables of interest are typically dependent within cluster. We propose a…

统计方法学 · 统计学 2021-08-26 Luis G. Leon-Novelo , Terrance D. Savitsky

Co-clustering targets on grouping the samples (e.g., documents, users) and the features (e.g., words, ratings) simultaneously. It employs the dual relation and the bilateral information between the samples and features. In many realworld…

机器学习 · 计算机科学 2016-11-18 Ping Li , Jiajun Bu , Chun Chen , Zhanying He , Deng Cai