FitCLIP:改进大规模预训练图像-文本模型用于零样本视频理解任务
计算机视觉与模式识别
2022-10-07 v2
摘要
大规模预训练图像-文本模型在少量任务中展现了令人难以置信的零样本性能,包括视频任务如动作识别和文本到视频检索。然而,这些模型尚未适应视频,主要是因为它们没有考虑时间维度,而且视频帧与典型图像不同(例如,包含运动模糊,清晰度较低)。在本文中,我们提出了一种微调策略,以改进这些大规模预训练图像-文本模型用于零样本视频理解任务。我们表明,通过仔细调整这些模型,我们在两个零样本动作识别任务和三个零样本文本到视频检索任务上获得了显著的改进。代码可在https://github.com/bryant1410/fitclip获取。
引用
@article{arxiv.2203.13371,
title = {FitCLIP: Refining Large-Scale Pretrained Image-Text Models for Zero-Shot Video Understanding Tasks},
author = {Santiago Castro and Fabian Caba Heilbron},
journal= {arXiv preprint arXiv:2203.13371},
year = {2022}
}
备注
Accepted at BMVC 2022. It includes the supplementary material. The margins and page size were modified to fit the arXiv ID stamp on the left side