推断计算最优视频视觉语言模型
计算机视觉与模式识别
2025-05-27 v1 计算与语言
摘要
本工作探讨了在视频视觉语言模型中,关于语言模型规模、帧数以及每帧视觉标记数这三个关键扩展因子上的推断计算最优分配问题。虽然先前研究通常致力于优化模型效率或提升性能,而不考虑资源约束,但我们则在固定推断计算预算下识别最优模型配置。我们进行大规模训练扫描和精细的参数化建模,以识别推断计算最优的前沿。实验结果揭示了任务性能如何依赖于扩展因子和微调数据规模,以及数据规模变化如何影响计算最优前沿。这些发现translates to practical tips for selecting these scaling factors.
引用
@article{arxiv.2505.18855,
title = {Inference Compute-Optimal Video Vision Language Models},
author = {Peiqi Wang and ShengYun Peng and Xuewen Zhang and Hanchao Yu and Yibo Yang and Lifu Huang and Fujun Liu and Qifan Wang},
journal= {arXiv preprint arXiv:2505.18855},
year = {2025}
}
备注
Annual Meeting of the Association for Computational Linguistics (ACL), 2025