中文
相关论文

相关论文: Fuzzy clustering of distribution-valued data using…

200 篇论文

Biclustering is an effective technique in data mining and pattern recognition. Biclustering algorithms based on traditional clustering face two fundamental limitations when processing high-dimensional data: (1) The distance concentration…

机器学习 · 计算机科学 2025-05-01 Yan Huang , Da-Qing Zhang

The most widely used internal measure for clustering evaluation is the silhouette coefficient, whose naive computation requires a quadratic number of distance calculations, which is clearly unfeasible for massive datasets. Surprisingly,…

数据结构与算法 · 计算机科学 2021-01-21 Federico Altieri , Andrea Pietracaprina , Geppino Pucci , Fabio Vandin

Scale invariance (fractality) is a prominent feature of the large-scale behavior of many stochastic systems. In this work, we construct an algorithm for the statistical identification of the Hurst distribution (in particular, the scaling…

统计方法学 · 统计学 2025-01-31 Patrice Abry , Gustavo Didier , Oliver Orejola , Herwig Wendt

We propose a method for estimating coefficients in multivariate regression when there is a clustering structure to the response variables. The proposed method includes a fusion penalty, to shrink the difference in fitted values from…

机器学习 · 统计学 2018-03-28 Bradley S. Price , Ben Sherwood

We introduce a principled way of computing the Wasserstein distance between two distributions in a federated manner. Namely, we show how to estimate the Wasserstein distance between two samples stored and kept on different devices/clients…

机器学习 · 计算机科学 2023-10-04 Alain Rakotomamonjy , Kimia Nadjahi , Liva Ralaivola

Clustering is an effective technique in data mining to generate groups that are the matter of interest. Among various clustering approaches, the family of k-means algorithms and min-cut algorithms gain most popularity due to their…

机器学习 · 计算机科学 2014-11-25 Xiaojun Chang , Feiping Nie , Zhigang Ma , Yi Yang

We introduce a modified model of random walk, and then develop two novel clustering algorithms based on it. In the algorithms, each data point in a dataset is considered as a particle which can move at random in space according to the…

机器学习 · 计算机科学 2008-10-31 Qiang Li , Yan He , Jing-ping Jiang

A new interpoint distance-based measure is proposed to identify the optimal number of clusters present in a data set. Designed in nonparametric approach, it is independent of the distribution of given data. Interpoint distances between the…

机器学习 · 计算机科学 2022-10-18 Soumita Modak

Learning conditional densities and identifying factors that influence the entire distribution are vital tasks in data-driven applications. Conventional approaches work mostly with summary statistics, and are hence inadequate for a…

统计方法学 · 统计学 2022-09-13 Chengliang Tang , Nathan Lenssen , Ying Wei , Tian Zheng

Recent work on overfitting Bayesian mixtures of distributions offers a powerful framework for clustering multivariate data using a latent Gaussian model which resembles the factor analysis model. The flexibility provided by overfitting…

统计方法学 · 统计学 2019-08-29 Panagiotis Papastamoulis

The optimal number of clusters is one of the main concerns when applying cluster analysis. Several cluster validity indexes have been introduced to address this problem. However, in some situations, there is more than one option that can be…

机器学习 · 统计学 2025-12-24 Nathakhun Wiroonsri , Onthada Preedasawakul

In this article, a new method, called FWP, is proposed for clustering longitudinal curves. In the proposed method, clusters of mean functions are identified through a weighted concave pairwise fusion method. The EM algorithm and the…

统计方法学 · 统计学 2023-06-14 Xin Wang

We consider a generalization of the discrete-time Self Healing Umbrella Sampling method, which is an adaptive importance technique useful to sample multimodal target distributions. The importance function is based on the weights (namely the…

概率论 · 数学 2017-09-04 Gersende Fort , Benjamin Jourdain , Tony Lelièvre , Gabriel Stoltz

The purpose of this paper is to study the algorithm FCM and some of its famous innovations, analyse and discover the method of applying hedge algebra theory that uses algebra to represent linguistic-valued variables, to FCM. Then, this…

机器学习 · 计算机科学 2017-11-06 Hung Thai Le , Khang Ding Tran , Hung Van Le

The literature on clustering for continuous data is rich and wide; differently, that one developed for categorical data is still limited. In some cases, the problem is made more difficult by the presence of noise variables/dimensions that…

统计方法学 · 统计学 2015-04-14 Monia Ranalli , Roberto Rocci

In this work we analyse a number of variants of the Wasserstein distance which allow to focus the classification on the prescribed parts (fragments) of classified 2D curves. These variants are based on the use of a number of discrete…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Agnieszka Kaliszewska , Monika Syga

This paper describes a method for clustering data that are spread out over large regions and which dimensions are on different scales of measurement. Such an algorithm was developed to implement a robotics application consisting in sorting…

机器学习 · 计算机科学 2017-03-23 Joris Guérin , Olivier Gibaru , Stéphane Thiery , Eric Nyiri

Data distillation and coresets have emerged as popular approaches to generate a smaller representative set of samples for downstream learning tasks to handle large-scale datasets. At the same time, machine learning is being increasingly…

Objective: The main objective of this paper is to construct a distributed clustering algorithm based upon spatial data correlation among sensor nodes and perform data accuracy for each distributed cluster at their respective cluster head…

网络与互联网体系结构 · 计算机科学 2011-08-15 Jyotirmoy Karjee , H. S Jamadagni

As the multi-view data grows in the real world, multi-view clus-tering has become a prominent technique in data mining, pattern recognition, and machine learning. How to exploit the relation-ship between different views effectively using…

机器学习 · 计算机科学 2019-08-14 Zhaohong Deng , Chen Cui , Peng Xu , Ling Liang , Haoran Chen , Te Zhang , Shitong Wang