视频中密集描述事件:SYSU参加ActivityNet Challenge 2020
计算机视觉与模式识别
2020-08-13 v2
摘要
本技术报告简要介绍了我们提交至ActivityNet Challenge 2020密集视频描述任务的情况。我们的方法遵循两阶段流程:首先,提取一组时间事件提案;然后我们提出一个多事件描述模型以捕获事件级时间关系并有效融合多模态信息。我们的方法在测试集上取得了9.28的METEOR分数。
引用
@article{arxiv.2006.11693,
title = {Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020},
author = {Teng Wang and Huicheng Zheng and Mingjing Yu},
journal= {arXiv preprint arXiv:2006.11693},
year = {2020}
}
备注
Second-place solution to TASK 2 (Dense video captioning) in ActivityNet Challenge 2020. Code is available at https://github.com/ttengwang/dense-video-captioning-pytorch