中文
相关论文

相关论文: High-Dimensional Data Clustering

200 篇论文

We consider the problem of clustering noisy high-dimensional data points into a union of low-dimensional subspaces and a set of outliers. The number of subspaces, their dimensions, and their orientations are unknown. A probabilistic…

信息论 · 计算机科学 2013-07-19 Reinhard Heckel , Helmut Bölcskei

This paper studies the large-scale subspace clustering (LSSC) problem with million data points. Many popular subspace clustering methods cannot directly handle the LSSC problem although they have been considered as state-of-the-art methods…

机器学习 · 计算机科学 2020-04-10 Jun Li , Hongfu Liu , Zhiqiang Tao , Handong Zhao , Yun Fu

Multi-view clustering has been widely used in recent years in comparison to single-view clustering, for clear reasons, as it offers more insights into the data, which has brought with it some challenges, such as how to combine these views…

机器学习 · 计算机科学 2025-11-25 Alaeddine Zahir , Khalide Jbilou , Ahmed Ratnani

A mixture of joint generalized hyperbolic distributions (MJGHD) is introduced for asymmetric clustering for high-dimensional data. The MJGHD approach takes into account the cluster-specific subspace, thereby limiting the number of…

统计方法学 · 统计学 2018-11-02 Yang Tang , Ryan P. Browne , Paul D. McNicholas

Clustering aims to group unlabeled objects based on similarity inherent among them into clusters. It is important for many tasks such as anomaly detection, database sharding, record linkage, and others. Some clustering methods are taken as…

数据库 · 计算机科学 2024-12-02 Binbin Gu , Saeed Kargar , Faisal Nawab

Hierarchical clustering is one of the most powerful solutions to the problem of clustering, on the grounds that it performs a multi scale organization of the data. In recent years, research on hierarchical clustering methods has attracted…

机器学习 · 计算机科学 2019-08-02 Antonia Korba

Clustering and estimating cluster means are core problems in statistics and machine learning, with k-means and Expectation Maximization (EM) being two widely used algorithms. In this work, we provide a theoretical explanation for the…

机器学习 · 统计学 2025-06-19 David Silva-Sánchez , Roy R. Lederman

Developing an understanding of high-dimensional data can be facilitated by visualizing that data using dimensionality reduction. However, the low-dimensional embeddings are often difficult to interpret. To facilitate the exploration and…

机器学习 · 计算机科学 2025-04-16 Fuyin Lai , Edith Heiter , Guillaume Bied , Jefrey Lijffijt

For numerous reasons there raises a need for dimension reduction that preserves certain characteristics of data. In this work we focus on data coming from a mixture of Gaussian distributions and we propose a method that preserves…

统计理论 · 数学 2014-07-30 Ewa Nowakowska , Jacek Koronacki , Stan Lipovetsky

In this paper we propose a Deep Autoencoder MIxture Clustering (DAMIC) algorithm based on a mixture of deep autoencoders where each cluster is represented by an autoencoder. A clustering network transforms the data into another space and…

机器学习 · 计算机科学 2019-03-28 Shlomo E. Chazan , Sharon Gannot , Jacob Goldberger

Hyperdimensional computing (HDC) is an emerging computational framework that takes inspiration from attributes of neuronal circuits such as hyperdimensionality, fully distributed holographic representation, and (pseudo)randomness. When…

新兴技术 · 计算机科学 2020-04-10 Geethan Karunaratne , Manuel Le Gallo , Giovanni Cherubini , Luca Benini , Abbas Rahimi , Abu Sebastian

Data analysis and data mining are concerned with unsupervised pattern finding and structure determination in data sets. "Structure" can be understood as symmetry and a range of symmetries are expressed by hierarchy. Such symmetries directly…

机器学习 · 统计学 2015-03-17 Fionn Murtagh , Pedro Contreras

Hyperspectral image (HSI) clustering is a challenging task due to the high complexity of HSI data. Subspace clustering has been proven to be powerful for exploiting the intrinsic relationship between data points. Despite the impressive…

计算机视觉与模式识别 · 计算机科学 2020-09-01 Yaoming Cai , Zijia Zhang , Zhihua Cai , Xiaobo Liu , Xinwei Jiang , Qin Yan

We study two practically important cases of model based clustering using Gaussian Mixture Models: (1) when there is misspecification and (2) on high dimensional data, in the light of recent advances in Gradient Descent (GD) based…

机器学习 · 统计学 2020-07-28 Siva Rajesh Kasa , Vaibhav Rajan

Most of the research on clustering ensemble focuses on designing practical consistency learning algorithms.To solve the problems that the quality of base clusters varies and the low-quality base clusters have an impact on the performance of…

机器学习 · 计算机科学 2024-11-04 Jianwen Gan , Yan Chen , Peng Zhou , Liang Du

Subspace clustering (SC) is a promising clustering technology to identify clusters based on their associations with subspaces in high dimensional spaces. SC can be classified into hard subspace clustering (HSC) and soft subspace clustering…

机器学习 · 计算机科学 2016-04-11 Zhaohong Deng , Kup-Sze Choi , Yizhang Jiang , Jun Wang , Shitong Wang

Clustering large, mixed data is a central problem in data mining. Many approaches adopt the idea of k-means, and hence are sensitive to initialisation, detect only spherical clusters, and require a priori the unknown number of clusters. We…

机器学习 · 统计学 2020-11-13 Joshua Tobin , Mimi Zhang

Distribution learning focuses on learning the probability density function from a set of data samples. In contrast, clustering aims to group similar objects together in an unsupervised manner. Usually, these two tasks are considered…

机器学习 · 计算机科学 2023-08-31 Guanfang Dong , Chenqiu Zhao , Anup Basu

Deep clustering (DC), a fusion of deep representation learning and clustering, has recently demonstrated positive results in data science, particularly text processing and computer vision. However, joint optimization of feature learning and…

数据库 · 计算机科学 2024-05-29 Hafiz Tayyab Rauf , Andre Freitas , Norman W. Paton

Clustering methods seek to partition data such that elements are more similar to elements in the same cluster than to elements in different clusters. The main challenge in this task is the lack of a unified definition of a cluster,…

统计理论 · 数学 2022-07-06 Franz Besold , Vladimir Spokoiny