English

CNN-based RGB-D Salient Object Detection: Learn, Select and Fuse

Computer Vision and Pattern Recognition 2019-09-23 v1

Abstract

The goal of this work is to present a systematic solution for RGB-D salient object detection, which addresses the following three aspects with a unified framework: modal-specific representation learning, complementary cue selection and cross-modal complement fusion. To learn discriminative modal-specific features, we propose a hierarchical cross-modal distillation scheme, in which the well-learned source modality provides supervisory signals to facilitate the learning process for the new modality. To better extract the complementary cues, we formulate a residual function to incorporate complements from the paired modality adaptively. Furthermore, a top-down fusion structure is constructed for sufficient cross-modal interactions and cross-level transmissions. The experimental results demonstrate the effectiveness of the proposed cross-modal distillation scheme in zero-shot saliency detection and pre-training on a new modality, as well as the advantages in selecting and fusing cross-modal/cross-level complements.

Keywords

Cite

@article{arxiv.1909.09309,
  title  = {CNN-based RGB-D Salient Object Detection: Learn, Select and Fuse},
  author = {Hao Chen and Youfu Li},
  journal= {arXiv preprint arXiv:1909.09309},
  year   = {2019}
}

Comments

submitted to a journal in 12-October-2018

R2 v1 2026-06-23T11:20:56.838Z