English

UNINEXT-Cutie: The 1st Solution for LSVOS Challenge RVOS Track

Computer Vision and Pattern Recognition 2024-08-27 v2

Abstract

Referring video object segmentation (RVOS) relies on natural language expressions to segment target objects in video. In this year, LSVOS Challenge RVOS Track replaced the origin YouTube-RVOS benchmark with MeViS. MeViS focuses on referring the target object in a video through its motion descriptions instead of static attributes, posing a greater challenge to RVOS task. In this work, we integrate strengths of that leading RVOS and VOS models to build up a simple and effective pipeline for RVOS. Firstly, We finetune the state-of-the-art RVOS model to obtain mask sequences that are correlated with language descriptions. Secondly, based on a reliable and high-quality key frames, we leverage VOS model to enhance the quality and temporal consistency of the mask results. Finally, we further improve the performance of the RVOS model using semi-supervised learning. Our solution achieved 62.57 J&F on the MeViS test set and ranked 1st place for 6th LSVOS Challenge RVOS Track.

Keywords

Cite

@article{arxiv.2408.10129,
  title  = {UNINEXT-Cutie: The 1st Solution for LSVOS Challenge RVOS Track},
  author = {Hao Fang and Feiyu Pan and Xiankai Lu and Wei Zhang and Runmin Cong},
  journal= {arXiv preprint arXiv:2408.10129},
  year   = {2024}
}