用于精细视频理解的长期特征库
计算机视觉与模式识别
2019-04-19 v2
摘要
为了理解世界,我们人类不断需要将当下与过去联系起来,并将事件置于上下文中。在本文中,我们使现有的视频模型能够做到这一点。我们提出一个长期特征库——提取自整个视频跨度上的辅助信息——以增强最先进的视频模型,否则这些模型只能观看2-5秒的短片段。我们的实验表明,用长期特征库增强3D卷积网络在三个具有挑战性的视频数据集上取得了最先进的结果:AVA、EPIC-Kitchens和Charades。
引用
@article{arxiv.1812.05038,
title = {Long-Term Feature Banks for Detailed Video Understanding},
author = {Chao-Yuan Wu and Christoph Feichtenhofer and Haoqi Fan and Kaiming He and Philipp Krähenbühl and Ross Girshick},
journal= {arXiv preprint arXiv:1812.05038},
year = {2019}
}
备注
Code and models are available at https://github.com/facebookresearch/video-long-term-feature-banks