English

What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations

Computation and Language 2025-05-27 v4 Artificial Intelligence Computer Vision and Pattern Recognition

Abstract

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains. VISTA contains 18,599 recorded AI conference presentations paired with their corresponding paper abstracts. We benchmark the performance of state-of-the-art large models and apply a plan-based framework to better capture the structured nature of abstracts. Both human and automated evaluations confirm that explicit planning enhances summary quality and factual consistency. However, a considerable gap remains between models and human performance, highlighting the challenges of our dataset. This study aims to pave the way for future research on scientific video-to-text summarization.

Keywords

Cite

@article{arxiv.2502.08279,
  title  = {What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations},
  author = {Dongqi Liu and Chenxi Whitehouse and Xi Yu and Louis Mahon and Rohit Saxena and Zheng Zhao and Yifu Qiu and Mirella Lapata and Vera Demberg},
  journal= {arXiv preprint arXiv:2502.08279},
  year   = {2025}
}

Comments

ACL 2025 Main & Long Conference Paper