中文

Grid-LOGAT:用于视频问答的基于网格的局部与全局区域转录

计算机视觉与模式识别 2025-07-29 v3 人工智能

摘要

在本文中,我们提出了一种用于视频问答(VideoQA)的基于网格的局部与全局区域转录(Grid-LoGAT)系统。该系统分两个阶段运行。首先,使用视觉语言模型(VLM)从视频帧中提取文本转录。接着,利用这些转录处理问题,并通过大语言模型(LLM)生成答案。这种设计通过将 VLM 部署在边缘设备上、将 LLM 部署在云端,确保了图像隐私。为提高转录质量,我们提出了基于网格的视觉提示,该方法从每个网格单元中提取复杂的局部细节,并将其与全局信息相整合。评估结果表明,使用开源 VLM(LLaVA-1.6-7B)和 LLM(Llama-3.1-8B)的 Grid-LoGAT,在 NExT-QA 和 STAR-QA 数据集上以 65.9% 和 50.11% 的准确率优于具有相似基线模型的最先进方法。此外,在我们使用 NExT-QA 创建的基于定位的问题上,我们的方法比非网格版本高出 24 个百分点。(本文已被 IEEE ICIP 2025 接收。)

关键词

引用

@article{arxiv.2505.24371,
  title  = {Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering},
  author = {Md Intisar Chowdhury and Kittinun Aukkapinyo and Hiroshi Fujimura and Joo Ann Woo and Wasu Wasusatein and Fadoua Ghourabi},
  journal= {arXiv preprint arXiv:2505.24371},
  year   = {2025}
}

备注

Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works