English

Vision Transformers Are Good Mask Auto-Labelers

Computer Vision and Pattern Recognition 2023-01-11 v1 Machine Learning Multimedia

Abstract

We propose Mask Auto-Labeler (MAL), a high-quality Transformer-based mask auto-labeling framework for instance segmentation using only box annotations. MAL takes box-cropped images as inputs and conditionally generates their mask pseudo-labels.We show that Vision Transformers are good mask auto-labelers. Our method significantly reduces the gap between auto-labeling and human annotation regarding mask quality. Instance segmentation models trained using the MAL-generated masks can nearly match the performance of their fully-supervised counterparts, retaining up to 97.4\% performance of fully supervised models. The best model achieves 44.1\% mAP on COCO instance segmentation (test-dev 2017), outperforming state-of-the-art box-supervised methods by significant margins. Qualitative results indicate that masks produced by MAL are, in some cases, even better than human annotations.

Keywords

Cite

@article{arxiv.2301.03992,
  title  = {Vision Transformers Are Good Mask Auto-Labelers},
  author = {Shiyi Lan and Xitong Yang and Zhiding Yu and Zuxuan Wu and Jose M. Alvarez and Anima Anandkumar},
  journal= {arXiv preprint arXiv:2301.03992},
  year   = {2023}
}
R2 v1 2026-06-28T08:08:33.994Z