中文

VDMA:基于动态生成多智能体的视频问答

计算机视觉与模式识别 2024-07-08 v1

摘要

本技术报告详细描述了我们针对EgoSchema Challenge 2024的方法。EgoSchema Challenge旨在识别针对给定视频片段相关问题的最合适回答。在本文中,我们提出了基于动态生成多智能体的视频问答(VDMA)。该方法通过采用具有动态生成专家智能体的多智能体系统,作为现有响应生成系统的补充方法。该方法旨在提供最准确且符合上下文的回答。本报告详述了我们方法的各个阶段、所使用的工具以及实验结果。

关键词

引用

@article{arxiv.2407.03610,
  title  = {VDMA: Video Question Answering with Dynamically Generated Multi-Agents},
  author = {Noriyuki Kugo and Tatsuya Ishibashi and Kosuke Ono and Yuji Sato},
  journal= {arXiv preprint arXiv:2407.03610},
  year   = {2024}
}

备注

4 pages, 2 figures