中文
相关论文

相关论文: Large-scale entity resolution via microclustering …

200 篇论文

We present a novel probabilistic clustering model for objects that are represented via pairwise distances and observed at different time points. The proposed method utilizes the information given by adjacent time points to find the…

Based on the self-energy-functional approach proposed recently [M. Potthoff, Eur. Phys. J. B 32, 429 (2003)], we present an extension of the cluster-perturbation theory to systems with spontaneously broken symmetry. Our method applies to…

强关联电子 · 物理学 2007-05-23 C. Dahnken , M. Aichhorn , W. Hanke , E. Arrigoni , M. Potthoff

This paper investigates the asymptotic properties of parameter estimation for the Ewens--Pitman partition with parameters $0<\alpha<1$ and $\theta>-\alpha$. Especially, we show that the maximum likelihood estimator (MLE) of $\alpha$ is…

统计理论 · 数学 2025-05-06 Takuya Koriyama , Takeru Matsuda , Fumiyasu Komaki

Probabilistic clustering models (or equivalently, mixture models) are basic building blocks in countless statistical models and involve latent random variables over discrete spaces. For these models, posterior inference methods can be…

机器学习 · 统计学 2020-06-24 Ari Pakman , Yueqi Wang , Catalin Mitelut , JinHyung Lee , Liam Paninski

Modern data-driven and distributed learning frameworks deal with diverse massive data generated by clients spread across heterogeneous environments. Indeed, data heterogeneity is a major bottleneck in scaling up many distributed learning…

机器学习 · 计算机科学 2023-08-23 Amirhossein Reisizadeh , Khashayar Gatmiry , Asuman Ozdaglar

A novel family of twelve mixture models with random covariates, nested in the linear $t$ cluster-weighted model (CWM), is introduced for model-based clustering. The linear $t$ CWM was recently presented as a robust alternative to the better…

统计计算 · 统计学 2015-03-10 Salvatore Ingrassia , Simona C. Minotti , Antonio Punzo

Subsampling techniques can reduce the computational costs of processing big data. Practical subsampling plans typically involve initial uniform sampling and refined sampling. With a subsample, big data inferences are generally built on the…

统计方法学 · 统计学 2022-09-13 Yan Fan , Yang Liu , Yukun Liu , Jing Qin

We propose a novel methodology for feature screening in clustering massive datasets, in which both the number of features and the number of observations can potentially be very large. Taking advantage of a fusion penalization based convex…

统计方法学 · 统计学 2017-10-05 Trambak Banerjee , Gourab Mukherjee , Peter Radchenko

We study the sparse high-dimensional Gaussian mixture model when the number of clusters is allowed to grow with the sample size. A minimax lower bound for parameter estimation is established, and we show that a constrained maximum…

统计理论 · 数学 2024-02-26 Dapeng Yao , Fangzheng Xie , Yanxun Xu

Any limiting point process for the time normalized exceedances of high levels by a stationary sequence is necessarily compound Poisson under appropriate long range dependence conditions. Typically exceedances appear in clusters. The…

应用统计 · 统计学 2009-03-03 Christian Y. Robert

One iteration of standard $k$-means (i.e., Lloyd's algorithm) or standard EM for Gaussian mixture models (GMMs) scales linearly with the number of clusters $C$, data points $N$, and data dimensionality $D$. In this study, we explore whether…

机器学习 · 统计学 2018-04-18 Dennis Forster , Jörg Lücke

This study introduces a general semiparametric clusterwise index distribution model to analyze how latent clusters affect the covariate-response relationships. By employing sufficient dimension reduction to account for the effects of…

统计方法学 · 统计学 2025-09-30 Jen-Chieh Teng , Chin-Tsang Chiang

Despite the remarkable success of Large Language Models (LLMs) in text understanding and generation, their potential for text clustering tasks remains underexplored. We observed that powerful closed-source LLMs provide good quality…

We present a novel framework for concomitant dimension reduction and clustering. This framework is based on a novel class of Bayesian clustering factor models. These models assume a factor model structure where the vectors of common factors…

统计方法学 · 统计学 2025-05-09 Hwasoo Shin , Marco A. R. Ferreira , Allison N. Tegge

Research on cluster analysis for categorical data continues to develop, with new clustering algorithms being proposed. However, in this context, the determination of the number of clusters is rarely addressed. In this paper, we propose a…

统计方法学 · 统计学 2014-09-29 Cláudia Silvestre , Margarida G. M. S. Cardoso , Mário A. T. Figueiredo

Variable clustering is important for explanatory analysis. However, only few dedicated methods for variable clustering with the Gaussian graphical model have been proposed. Even more severe, small insignificant partial correlations due to…

应用统计 · 统计学 2018-06-18 Daniel Andrade , Akiko Takeda , Kenji Fukumizu

In cluster analysis, it can be useful to interpret the partition built from the data in the light of external categorical variables which were not directly involved to cluster the data. An approach is proposed in the model-based clustering…

Clustering is a fundamental task in data mining and machine learning, particularly for analyzing large-scale data. In this paper, we introduce Clust-Splitter, an efficient algorithm based on nonsmooth optimization, designed to solve the…

机器学习 · 计算机科学 2026-03-19 Jenni Lampainen , Kaisa Joki , Napsu Karmitsa , Marko M. Mäkelä

In this article we study a problem within Dempster-Shafer theory where 2**n - 1 pieces of evidence are clustered by a neural structure into n clusters. The clustering is done by minimizing a metaconflict function. Previously we developed a…

人工智能 · 计算机科学 2007-05-23 Johan Schubert

We present a structural clustering algorithm for large-scale datasets of small labeled graphs, utilizing a frequent subgraph sampling strategy. A set of representatives provides an intuitive description of each cluster, supports the…

数据库 · 计算机科学 2016-10-03 Till Schäfer , Petra Mutzel