English

TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning

Computation and Language 2024-10-11 v4 Artificial Intelligence Computer Vision and Pattern Recognition

Abstract

It is challenging for models to understand complex, multimodal content such as television clips, and this is in part because video-language models often rely on single-modality reasoning and lack interpretability. To combat these issues we propose TV-TREES, the first multimodal entailment tree generator. TV-TREES serves as an approach to video understanding that promotes interpretable joint-modality reasoning by searching for trees of entailment relationships between simple text-video evidence and higher-level conclusions that prove question-answer pairs. We also introduce the task of multimodal entailment tree generation to evaluate reasoning quality. Our method's performance on the challenging TVQA benchmark demonstrates interpretable, state-of-the-art zero-shot performance on full clips, illustrating that multimodal entailment tree generation can be a best-of-both-worlds alternative to black-box systems.

Keywords

Cite

@article{arxiv.2402.19467,
  title  = {TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning},
  author = {Kate Sanders and Nathaniel Weir and Benjamin Van Durme},
  journal= {arXiv preprint arXiv:2402.19467},
  year   = {2024}
}

Comments

9 pages, EMNLP 2024