It is challenging for models to understand complex, multimodal content such as television clips, and this is in part because video-language models often rely on single-modality reasoning and lack interpretability. To combat these issues we propose TV-TREES, the first multimodal entailment tree generator. TV-TREES serves as an approach to video understanding that promotes interpretable joint-modality reasoning by searching for trees of entailment relationships between simple text-video evidence and higher-level conclusions that prove question-answer pairs. We also introduce the task of multimodal entailment tree generation to evaluate reasoning quality. Our method's performance on the challenging TVQA benchmark demonstrates interpretable, state-of-the-art zero-shot performance on full clips, illustrating that multimodal entailment tree generation can be a best-of-both-worlds alternative to black-box systems.
@article{arxiv.2402.19467,
title = {TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning},
author = {Kate Sanders and Nathaniel Weir and Benjamin Van Durme},
journal= {arXiv preprint arXiv:2402.19467},
year = {2024}
}