English

R-MAE: Regions Meet Masked Autoencoders

Computer Vision and Pattern Recognition 2024-01-08 v2

Abstract

In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we design an architecture which efficiently addresses the one-to-many mapping between images and regions, while being highly effective especially with high-quality regions. When integrated with MAE, our approach (R-MAE) demonstrates consistent improvements across various pre-training datasets and downstream detection and segmentation benchmarks, with negligible computational overheads. Beyond the quantitative evaluation, our analysis indicates the models pre-trained with masked region autoencoding unlock the potential for interactive segmentation. The code is provided at https://github.com/facebookresearch/r-mae.

Keywords

Cite

@article{arxiv.2306.05411,
  title  = {R-MAE: Regions Meet Masked Autoencoders},
  author = {Duy-Kien Nguyen and Vaibhav Aggarwal and Yanghao Li and Martin R. Oswald and Alexander Kirillov and Cees G. M. Snoek and Xinlei Chen},
  journal= {arXiv preprint arXiv:2306.05411},
  year   = {2024}
}
R2 v1 2026-06-28T11:00:20.100Z