中文
相关论文

相关论文: Comparing high dimensional partitions, with the Co…

200 篇论文

Dimensionality reduction (DR) techniques are often characterized by whether they preserve global, high-level structures in the data or local, neighborhood structures. This distinction matters in visualization: global methods can obscure…

机器学习 · 计算机科学 2026-05-04 Kaviru Gunaratne , Stephen Kobourov , Jacob Miller

Adjusted similarity measures, such as Cohen's kappa for inter-rater reliability and the adjusted Rand index used to compare clustering algorithms, are a vital tool for comparing discrete labellings. These measures are intended to have the…

统计方法学 · 统计学 2026-01-16 William L. Lippitt , Edward J. Bedrick , Nichole E. Carlson

In this paper, we present a novel method for co-clustering, an unsupervised learning approach that aims at discovering homogeneous groups of data instances and features by grouping them simultaneously. The proposed method uses the entropy…

机器学习 · 统计学 2017-05-22 Charlotte Laclau , Ievgen Redko , Basarab Matei , Younès Bennani , Vincent Brault

In this paper, we provide an approach to clustering relational matrices whose entries correspond to either similarities or dissimilarities between objects. Our approach is based on the value of information, a parameterized,…

人工智能 · 计算机科学 2017-10-31 Isaac J. Sledge , Jose C. Principe

Rand (1971) proposed what has since become a well-known index for comparing two partitions obtained on the same set of units. The index takes a value on the interval between 0 and 1, where a higher value indicates more similar partitions.…

统计方法学 · 统计学 2018-05-22 Marjan Cugmas , Anuška Ferligoj

Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying…

机器学习 · 统计学 2008-03-26 Benhuai Xie , Wei Pan , Xiaotong Shen

In modern randomized experiments, large-scale data collection increasingly yields rich baseline covariates and auxiliary information from multiple sources. Such information offers opportunities for more precise treatment effect estimation,…

统计方法学 · 统计学 2026-03-10 Wei Ma , Zeqi Wu , Zheng Zhang

Subspace clustering refers to the problem of segmenting high dimensional data drawn from a union of subspaces into the respective subspaces. In some applications, partial side-information to indicate "must-link" or "cannot-link" in…

计算机视觉与模式识别 · 计算机科学 2018-05-23 Chun-Guang Li , Junjian Zhang , Jun Guo

Matrix valued data has become increasingly prevalent in many applications. Most of the existing clustering methods for this type of data are tailored to the mean model and do not account for the dependence structure of the features, which…

机器学习 · 统计学 2023-12-07 Inbeom Lee , Siyi Deng , Yang Ning

The hybrid clustering-classification neural network is proposed. This network allows increasing a quality of information processing under the condition of overlapping classes due to the rational choice of a learning rate parameter and…

机器学习 · 计算机科学 2016-10-26 Yevgeniy Bodyanskiy , Olena Vynokurova , Volodymyr Savvo , Tatiana Tverdokhlib , Pavlo Mulesa

Clustering categorical data is an integral part of data mining and has attracted much attention recently. In this paper, we present k-ANMI, a new efficient algorithm for clustering categorical data. The k-ANMI algorithm works in a way that…

人工智能 · 计算机科学 2007-05-23 Zengyou He , Xiaofei Xu , Shengchun Deng

We propose a MAP Bayesian approach to perform and evaluate a co-clustering of mixed-type data tables. The proposed model infers an optimal segmentation of all variables then performs a co-clustering by minimizing a Bayesian model selection…

机器学习 · 统计学 2019-02-07 Aichetou Bouchareb , Marc Boullé , Fabrice Rossi , Fabrice Clérot

Modern high-dimensional methods often adopt the "bet on sparsity" principle, while in supervised multivariate learning statisticians may face "dense" problems with a large number of nonzero coefficients. This paper proposes a novel…

机器学习 · 统计学 2022-02-10 Yiyuan She , Jiahui Shen , Chao Zhang

Clustering, like covariate selection for classification, is an important step to compress and interpret the data. However, clustering of covariates is often performed independently of the classification step, which can lead to undesirable…

统计计算 · 统计学 2020-04-08 Daniel Andrade , Kenji Fukumizu , Yuzuru Okajima

Cluster analysis is one of the essential tasks in data mining and knowledge discovery. Each type of data poses unique challenges in achieving relatively efficient partitioning of the data into homogeneous groups. While the algorithms for…

机器学习 · 计算机科学 2018-12-11 Ruben A. Gevorgyan , Yenok B. Hakobyan

Predicting missing links in incomplete complex networks efficiently and accurately is still a challenging problem. The recently proposed CAR (Cannistrai-Alanis-Ravai) index shows the power of local link/triangle information in improving…

社会与信息网络 · 计算机科学 2016-03-23 Zhihao Wu , Youfang Lin , Jing Wang , Steve Gregory

An appropriate distance metric is crucial for categorical data clustering, as the distance between categorical data cannot be directly calculated. However, the distances between attribute values usually vary in different clusters induced by…

机器学习 · 计算机科学 2026-03-09 Taixi Chen , Yiu-ming Cheung , Yiqun Zhang

This study introduces a general semiparametric clusterwise index distribution model to analyze how latent clusters affect the covariate-response relationships. By employing sufficient dimension reduction to account for the effects of…

统计方法学 · 统计学 2025-09-30 Jen-Chieh Teng , Chin-Tsang Chiang

A clustered adaptive intervention (cAI) is a pre-specified sequence of decision rules that guides practitioners on how best - and based on which measures - to tailor cluster-level intervention to improve outcomes at the level of individuals…

统计方法学 · 统计学 2025-05-05 Yao Song , Kelly Speth , Amy Kilbourne , Andrew Quanbeck , Daniel Almirall , Lu Wang

Bi-clustering refers to the task of finding sub-matrices (indexed by a group of columns and a group of rows) within a matrix of data such that the elements of each sub-matrix (data and features) are related in a particular way, for…

机器学习 · 计算机科学 2021-11-15 Kaijie Xu