FOSNet:一种端到端可训练的场景识别深度神经网络
计算机视觉与模式识别
2019-07-19 v2
摘要
场景识别是一种图像识别问题,旨在预测拍摄图像所处地点的类别。本文提出一种使用卷积神经网络(CNN)的新场景识别方法。该方法基于给定图像中物体信息与场景信息的融合,其 CNN 框架命名为 FOS(物体与场景融合)Net。此外,提出一种称为场景一致性损失(SCL)的新损失来训练 FOSNet 并提升场景识别性能。所提 SCL 基于场景的独特特性,即“场景性”在整幅图像上扩散且场景类别不改变。所提 FOSNet 在三个最流行的场景识别数据集上进行了实验,并在其中两个集合上取得最优性能:Places 2 上 60.14%,MIT indoor 67 上 90.37%。在 SUN 397 上取得第二高的性能 77.28%。
引用
@article{arxiv.1907.07570,
title = {FOSNet: An End-to-End Trainable Deep Neural Network for Scene Recognition},
author = {Hongje Seong and Junhyuk Hyun and Euntai Kim},
journal= {arXiv preprint arXiv:1907.07570},
year = {2019}
}
备注
2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works