感知测试:面向多模态视频模型的诊断性基准
计算机视觉与模式识别
2023-11-01 v2 人工智能
机器学习
摘要
我们提出一种新颖的多模态视频基准——感知测试(Perception Test),用于评估预训练多模态模型(如 Flamingo、SeViLA 或 GPT-4)的感知与推理能力。相较于聚焦于计算任务(如分类、检测或跟踪)的现有基准,感知测试聚焦于跨视频、音频和文本模态的技能(记忆、抽象、物理、语义)与推理类型(描述性、解释性、预测性、反事实性),以提供全面且高效的评估工具。该基准在零样本/少样本或有限微调设定下探查预训练模型的迁移能力。为此,感知测试引入了 11.6k 个真实世界视频,平均时长 23 秒,旨在展现感知上有趣的情景,由全球约 100 名参与者拍摄。这些视频被密集标注了六类标签(多选与接地视频问答、物体与点轨迹、时序动作与声音片段),支持语言与非语言评估。除带有保留测试集的挑战服务器外,该基准的微调与验证划分已公开可用(CC-BY 许可)。人类基线结果对比最先进视频问答模型显示出显著性能差距(91.4% 对 46.2%),表明多模态视频理解仍有较大提升空间。数据集、基线代码与挑战服务器可在 https://github.com/deepmind/perception_test 获取。
引用
@article{arxiv.2305.13786,
title = {Perception Test: A Diagnostic Benchmark for Multimodal Video Models},
author = {Viorica Pătrăucean and Lucas Smaira and Ankush Gupta and Adrià Recasens Continente and Larisa Markeeva and Dylan Banarse and Skanda Koppula and Joseph Heyward and Mateusz Malinowski and Yi Yang and Carl Doersch and Tatiana Matejovicova and Yury Sulsky and Antoine Miech and Alex Frechette and Hanna Klimczak and Raphael Koster and Junlin Zhang and Stephanie Winkler and Yusuf Aytar and Simon Osindero and Dima Damen and Andrew Zisserman and João Carreira},
journal= {arXiv preprint arXiv:2305.13786},
year = {2023}
}
备注
37th Conference on Neural Information Processing Systems (NeurIPS 2023) Track on Datasets and Benchmarks