中文
相关论文

相关论文: Large-scale entity resolution via microclustering …

200 篇论文

In this paper, we consider the task of clustering a set of individual time series while modeling each cluster, that is, model-based time series clustering. The task requires a parametric model with sufficient flexibility to describe the…

机器学习 · 计算机科学 2023-02-23 Ryohei Umatani , Takashi Imai , Kaoru Kawamoto , Shutaro Kunimasa

Ewens-Pitman model has been successfully applied to various fields including Bayesian statistics. There are four important estimators $K_{n},M_{l,n}$,$K_{m}^{(n)},M_{l,m}^{(n)}$. In particular, $M_{1,n}, M_{1,m}^{(n)}$ are related to…

概率论 · 数学 2018-11-20 Youzhou Zhou

We use statistical mechanics to study model-based Bayesian data clustering. In this approach, each partition of the data into clusters is regarded as a microscopic system state, the negative data log-likelihood gives the energy of each…

无序系统与神经网络 · 物理学 2019-11-19 Alexander Mozeika , Anthony CC Coolen

Composite endpoints are increasingly used in clinical trials to capture treatment effects across multiple or hierarchically ordered outcomes. Although inference procedures based on win statistics, such as the win ratio, win odds, and net…

统计方法学 · 统计学 2025-10-28 Xi Fang , Zhiqiang Cao , Fan Li

A simple explicit construction is provided of a partition-valued fragmentation process whose distribution on partitions of $[n]=\{1,...,n\}$ at time $\theta \ge 0$ is governed by the Ewens sampling formula with parameter $\theta$. These…

概率论 · 数学 2007-05-23 Alexander Gnedin , Jim Pitman

The increasing amount of data on the Web, in particular of Linked Data, has led to a diverse landscape of datasets, which make entity retrieval a challenging task. Explicit cross-dataset links, for instance to indicate co-references or…

信息检索 · 计算机科学 2017-03-31 Besnik Fetahu , Ujwal Gadiraju , Stefan Dietze

We study the large sample behavior of a convex clustering framework, which minimizes the sample within cluster sum of squares under an~$\ell_1$ fusion constraint on the cluster centroids. This recently proposed approach has been gaining in…

统计方法学 · 统计学 2016-12-30 Peter Radchenko , Gourab Mukherjee

Networks often exhibit structure at disparate scales. We propose a method for identifying community structure at different scales based on multiresolution modularity and consensus clustering. Our contribution consists of two parts. First,…

社会与信息网络 · 计算机科学 2018-02-01 Lucas G. S. Jeub , Olaf Sporns , Santo Fortunato

We introduce a new method for performing clustering with the aim of fitting clusters with different scatters and weights. It is designed by allowing to handle a proportion $\alpha$ of contaminating data to guarantee the robustness of the…

We introduce a new Partition of Unity Method for the numerical homogenization of elliptic partial differential equations with arbitrarily rough coefficients. We do not restrict to a particular ansatz space or the existence of a finite…

数值分析 · 数学 2016-05-04 Daniel Peterseim , Patrick Henning , Philipp Morgenstern

As the data size in Machine Learning fields grows exponentially, it is inevitable to accelerate the computation by utilizing the ever-growing large number of available cores provided by high-performance computing hardware. However, existing…

机器学习 · 计算机科学 2021-04-23 Kun Li , Liang Yuan , Yunquan Zhang , Gongwei Chen

Clustering is one of the most widely used procedures in the analysis of microarray data, for example with the goal of discovering cancer subtypes based on observed heterogeneity of genetic marks between different tissues. It is well-known…

统计方法学 · 统计学 2009-04-21 Heng Lian

This work proposes stochastic partial differential equations (SPDEs) as a practical tool to replicate clustering effects of more detailed particle-based dynamics. Inspired by membrane-mediated receptor dynamics on cell surfaces, we…

In this paper we present a family of algorithms that can simultaneously align and cluster sets of multidimensional curves measured on a discrete time grid. Our approach is based on a generative mixture model that allows non-linear time…

应用统计 · 统计学 2012-12-12 Darya Chudova , Scott Gaffney , Padhraic Smyth

With the recent growth in data availability and complexity, and the associated outburst of elaborate modelling approaches, model selection tools have become a lifeline, providing objective criteria to deal with this increasingly challenging…

统计方法学 · 统计学 2020-10-08 Alessandro Casa , Luca Scrucca , Giovanna Menardi

Bayesian mixture models are widely used for clustering of high-dimensional data with appropriate uncertainty quantification. However, as the dimension of the observations increases, posterior inference often tends to favor too many or too…

统计方法学 · 统计学 2022-11-22 Noirrit Kiran Chandra , Antonio Canale , David B. Dunson

We introduce a dimension reduction method for visualizing the clustering structure obtained from a finite mixture of Gaussian densities. Information on the dimension reduction subspace is obtained from the variation on group means and,…

统计方法学 · 统计学 2015-08-10 Luca Scrucca

The expectation-maximization (EM) algorithm is an iterative method for finding maximum likelihood estimates when data are incomplete or are treated as being incomplete. The EM algorithm and its variants are commonly used for parameter…

统计计算 · 统计学 2013-06-26 Ryan P. Browne , Sanjeena Subedi , Paul McNicholas

We introduce a general semiparametric clusterwise elliptical distribution to assess how latent cluster structure shapes continuous outcomes. Using a subjectwise representation, we first estimate cluster-specific mean vectors and a…

统计方法学 · 统计学 2026-04-10 Jen-Chieh Teng , Sheng-Hsin Fan , Chin-Tsang Chiang , Ming-Yueh Huang , Alvin Lim

Algorithms based on spectral graph cut objectives such as normalized cuts, ratio cuts and ratio association have become popular in recent years because they are widely applicable and simple to implement via standard eigenvector…

计算机视觉与模式识别 · 计算机科学 2014-11-27 Xiangyang Zhou , Jiaxin Zhang , Brian Kulis