中文

基于卷积编码器 - 解码器神经网络的一步式时间依赖未来视频帧预测

计算机视觉与模式识别 2017-07-25 v2

摘要

自动驾驶汽车、无人机和其他机器人 inherently 需要对其环境的行为方式有一个概念,并能够预见近期的变化。在本工作中,我们专注于根据视频的当前帧来预测未来的外观。现有工作要么侧重于将未来外观预测为视频的下一帧,要么侧重于从单个视频帧开始预测未来运动(如光流或运动轨迹)。这项工作拓展了 CNN(卷积神经网络)的能力,使其能够预测任意给定未来时刻的外观 anticipation,而不必局限于视频的下一帧。我们将预测的未来外观条件化于一个连续时间变量,这使得我们能够直接从输入视频帧预测给定时间距离的未来帧。我们表明,CNN 可以学习随时间变化的典型外观变化的内在表示,并成功地在近未来的特定时间差生成逼真的预测。

关键词

引用

@article{arxiv.1702.04125,
  title  = {One-Step Time-Dependent Future Video Frame Prediction with a Convolutional Encoder-Decoder Neural Network},
  author = {Vedran Vukotić and Silvia-Laura Pintea and Christian Raymond and Guillaume Gravier and Jan Van Gemert},
  journal= {arXiv preprint arXiv:1702.04125},
  year   = {2017}
}

备注

11 pages, 1 figures, published in the International Conference of Image Analysis and Processing (ICIAP) 2017 and in the Netherlands Conference on Computer Vision (NCCV) 2016