English

On bicluster aggregation and its benefits for enumerative solutions

Machine Learning 2015-06-04 v1

Abstract

Biclustering involves the simultaneous clustering of objects and their attributes, thus defining local two-way clustering models. Recently, efficient algorithms were conceived to enumerate all biclusters in real-valued datasets. In this case, the solution composes a complete set of maximal and non-redundant biclusters. However, the ability to enumerate biclusters revealed a challenging scenario: in noisy datasets, each true bicluster may become highly fragmented and with a high degree of overlapping. It prevents a direct analysis of the obtained results. To revert the fragmentation, we propose here two approaches for properly aggregating the whole set of enumerated biclusters: one based on single linkage and the other directly exploring the rate of overlapping. Both proposals were compared with each other and with the actual state-of-the-art in several experiments, and they not only significantly reduced the number of biclusters but also consistently increased the quality of the solution.

Keywords

Cite

@article{arxiv.1506.01077,
  title  = {On bicluster aggregation and its benefits for enumerative solutions},
  author = {Saullo Haniell Galvão de Oliveira and Rosana Veroneze and Fernando José Von Zuben},
  journal= {arXiv preprint arXiv:1506.01077},
  year   = {2015}
}

Comments

15 pages, will be published by Springer Verlag in the LNAI Series in the book Advances in Data Mining

R2 v1 2026-06-22T09:46:12.309Z