English
Related papers

Related papers: Generalized Rescaled Polya urn and its statistical…

200 papers

We consider clustering based on significance tests for Gaussian Mixture Models (GMMs). Our starting point is the SigClust method developed by Liu et al. (2008), which introduces a test based on the k-means objective (with k = 2) to decide…

Methodology · Statistics 2019-10-08 Purvasha Chakravarti , Sivaraman Balakrishnan , Larry Wasserman

ProfileGLMM is an R package integrating Generalised Linear Mixed Models (GLMMs) as the outcome model for Bayesian profile regression. This statistical framework simultaneously i) explains the variation in the outcome and ii) clusters the…

Methodology · Statistics 2026-04-23 Matteo Amestoy , Mark A. van de Wiel , Wessel N. van Wieringen

The Recurrent Chinese Restaurant Process (RCRP) is a powerful statistical method for modeling evolving clusters in large scale social media data. With the RCRP, one can allow both the number of clusters and the cluster parameters in a model…

Artificial Intelligence · Computer Science 2017-08-22 Wei Wei , Kennth Joseph , Kathleen Carley

Quantifying distributional separation across groups is fundamental in statistical learning and scientific discovery, yet most classical discrepancy measures are tailored to two-group comparisons. We generalize the underlap coefficient…

Methodology · Statistics 2026-02-26 Zhaoxi Zhang , Vanda Inacio , Sara Wade

The clustering of bounded data presents unique challenges in statistical analysis due to the constraints imposed on the data values. This paper introduces a novel method for model-based clustering specifically designed for bounded data.…

Methodology · Statistics 2025-05-16 Luca Scrucca

Distribution shifts between sites can seriously degrade model performance since models are prone to exploiting unstable correlations. Thus, many methods try to find features that are stable across sites and discard unstable features.…

Machine Learning · Computer Science 2024-09-11 Minh Nguyen , Alan Q. Wang , Heejong Kim , Mert R. Sabuncu

Gaussian process (GP) regression is a flexible, nonparametric approach to regression that naturally quantifies uncertainty. In many applications, the number of responses and covariates are both large, and a goal is to select covariates that…

Methodology · Statistics 2022-10-12 Jian Cao , Joseph Guinness , Marc G. Genton , Matthias Katzfuss

Generalized $k$-means can be incorporated with any similarity or dissimilarity measure for clustering. By choosing the dissimilarity measure as the well known likelihood ratio or $F$-statistic, this work proposes a method based on…

Methodology · Statistics 2020-08-11 Tonglin Zhang , Ge Lin

Copulas, generalized estimating equations, and generalized linear mixed models promote the analysis of grouped data where non-normal responses are correlated. Unfortunately, parameter estimation remains challenging in these three…

Methodology · Statistics 2024-10-16 Sarah S. Ji , Benjamin B. Chu , Hua Zhou , Kenneth Lange

The generalized orthogonal Procrustes problem (GOPP) plays a fundamental role in several scientific disciplines including statistics, imaging science and computer vision. Despite its tremendous practical importance, it is generally an…

Information Theory · Computer Science 2024-12-25 Shuyang Ling

This work generalizes graph neural networks (GNNs) beyond those based on the Weisfeiler-Lehman (WL) algorithm, graph Laplacians, and diffusions. Our approach, denoted Relational Pooling (RP), draws from the theory of finite partial…

Machine Learning · Computer Science 2019-05-16 Ryan L. Murphy , Balasubramaniam Srinivasan , Vinayak Rao , Bruno Ribeiro

Belief Propagation (BP) is one of the most popular methods for inference in probabilistic graphical models. BP is guaranteed to return the correct answer for tree structures, but can be incorrect or non-convergent for loopy graphical…

Artificial Intelligence · Computer Science 2012-06-22 Siamak Ravanbakhsh , Chun-Nam Yu , Russell Greiner

Gaussian Process Regression (GPR) is a popular regression method, which unlike most Machine Learning techniques, provides estimates of uncertainty for its predictions. These uncertainty estimates however, are based on the assumption that…

Machine Learning · Computer Science 2024-08-29 Harris Papadopoulos

In consensus clustering, a clustering algorithm is used in combination with a subsampling procedure to detect stable clusters. Previous studies on both simulated and real data suggest that consensus clustering outperforms native algorithms.…

This paper provides a user's guide to the general theory of approximate randomization tests developed in Canay, Romano, and Shaikh (2017) when specialized to linear regressions with clustered data. An important feature of the methodology is…

Econometrics · Economics 2022-03-16 Yong Cai , Ivan A. Canay , Deborah Kim , Azeem M. Shaikh

Our understanding of observed Gravitational Waves (GWs) comes from matching data to known signal models describing General Relativity (GR). These models, expressed in the post-Newtonian formalism, contain the mathematical constant $\pi$.…

General Relativity and Quantum Cosmology · Physics 2020-05-13 Carl-Johan Haster

In order to sample marginalized and/or hard-to-reach populations, respondent-driven sampling (RDS) and similar techniques reach their participants via peer referral. Under a Markov model for RDS, previous research has shown that if the…

Statistics Theory · Mathematics 2022-06-08 Sebastien Roch , Karl Rohe

Every day, we judge the probability of propositions. When we communicate graded confidence (e.g. "I am 90% sure"), we enable others to gauge how much weight to attach to our judgment. Ideally, people should share their judgments to reach…

Quantitative Methods · Quantitative Biology 2025-01-10 Patrick Stinson , Jasper van den Bosch , Trenton Jerde , Nikolaus Kriegeskorte

Count data is becoming more and more ubiquitous in a wide range of applications, with datasets growing both in size and in dimension. In this context, an increasing amount of work is dedicated to the construction of statistical models…

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

Methodology · Statistics 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu