English

Scalable Co-Clustering for Large-Scale Data through Dynamic Partitioning and Hierarchical Merging

Distributed, Parallel, and Cluster Computing 2025-03-20 v3 Machine Learning

Abstract

Co-clustering simultaneously clusters rows and columns, revealing more fine-grained groups. However, existing co-clustering methods suffer from poor scalability and cannot handle large-scale data. This paper presents a novel and scalable co-clustering method designed to uncover intricate patterns in high-dimensional, large-scale datasets. Specifically, we first propose a large matrix partitioning algorithm that partitions a large matrix into smaller submatrices, enabling parallel co-clustering. This method employs a probabilistic model to optimize the configuration of submatrices, balancing the computational efficiency and depth of analysis. Additionally, we propose a hierarchical co-cluster merging algorithm that efficiently identifies and merges co-clusters from these submatrices, enhancing the robustness and reliability of the process. Extensive evaluations validate the effectiveness and efficiency of our method. Experimental results demonstrate a significant reduction in computation time, with an approximate 83% decrease for dense matrices and up to 30% for sparse matrices.

Keywords

Cite

@article{arxiv.2410.18113,
  title  = {Scalable Co-Clustering for Large-Scale Data through Dynamic Partitioning and Hierarchical Merging},
  author = {Zihan Wu and Zhaoke Huang and Hong Yan},
  journal= {arXiv preprint arXiv:2410.18113},
  year   = {2025}
}

Comments

8 pages, 2 figures

R2 v1 2026-06-28T19:33:15.415Z