中文

Language-Assisted Image Clustering Guided by Discriminative Relational Signals and Adaptive Semantic Centers

计算机视觉与模式识别 2026-03-27 v2

摘要

Language-Assisted Image Clustering (LAIC) augments the input images with additional texts with the help of vision-language models (VLMs) to promote clustering performance. Despite recent progress, existing LAIC methods often overlook two issues: (i) textual features constructed for each image are highly similar, leading to weak inter-class discriminability; (ii) the clustering step is restricted to pre-built image-text alignments, limiting the potential for better utilization of the text modality. To address these issues, we propose a new LAIC framework with two complementary components. First, we exploit cross-modal relations to produce more discriminative self-supervision signals for clustering, as it compatible with most VLMs training mechanisms. Second, we learn category-wise continuous semantic centers via prompt learning to produce the final clustering assignments. Extensive experiments on eight benchmark datasets demonstrate that our method achieves an average improvement of 2.6% over state-of-the-art methods, and the learned semantic centers exhibit strong interpretability. Code is available in the supplementary material.

关键词

引用

@article{arxiv.2603.24278,
  title  = {TopoMesh: High-Fidelity Mesh Autoencoding via Topological Unification},
  author = {Guan Luo and Xiu Li and Rui Chen and Xuanyu Yi and Jing Lin and Chia-Hao Chen and Jiahang Liu and Song-Hai Zhang and Jianfeng Zhang},
  journal= {arXiv preprint arXiv:2603.24278},
  year   = {2026}
}

备注

Accepted to CVPR 2026. Project page: https://logan0601.github.io/projects/topomesh/index.html