English

Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model

Computer Vision and Pattern Recognition 2024-12-30 v1 Artificial Intelligence Machine Learning

Abstract

Open-vocabulary panoptic segmentation remains a challenging problem. One of the biggest difficulties lies in training models to generalize to an unlimited number of classes using limited categorized training data. Recent popular methods involve large-scale vision-language pre-trained foundation models, such as CLIP. In this paper, we propose OMTSeg for open-vocabulary segmentation using another large-scale vision-language pre-trained model called BEiT-3 and leveraging the cross-modal attention between visual and linguistic features in BEiT-3 to achieve better performance. Experiments result demonstrates that OMTSeg performs favorably against state-of-the-art models.

Keywords

Cite

@article{arxiv.2412.18917,
  title  = {Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model},
  author = {Yi-Chia Chen and Wei-Hua Li and Chu-Song Chen},
  journal= {arXiv preprint arXiv:2412.18917},
  year   = {2024}
}

Comments

ICIP 2024

R2 v1 2026-06-28T20:48:46.623Z