English

Towards Neuro-Symbolic Video Understanding

Computer Vision and Pattern Recognition 2024-12-04 v3 Artificial Intelligence

Abstract

The unprecedented surge in video data production in recent years necessitates efficient tools to extract meaningful frames from videos for downstream tasks. Long-term temporal reasoning is a key desideratum for frame retrieval systems. While state-of-the-art foundation models, like VideoLLaMA and ViCLIP, are proficient in short-term semantic understanding, they surprisingly fail at long-term reasoning across frames. A key reason for this failure is that they intertwine per-frame perception and temporal reasoning into a single deep network. Hence, decoupling but co-designing semantic understanding and temporal reasoning is essential for efficient scene identification. We propose a system that leverages vision-language models for semantic understanding of individual frames but effectively reasons about the long-term evolution of events using state machines and temporal logic (TL) formulae that inherently capture memory. Our TL-based reasoning improves the F1 score of complex event identification by 9-15% compared to benchmarks that use GPT4 for reasoning on state-of-the-art self-driving datasets such as Waymo and NuScenes.

Keywords

Cite

@article{arxiv.2403.11021,
  title  = {Towards Neuro-Symbolic Video Understanding},
  author = {Minkyu Choi and Harsh Goel and Mohammad Omama and Yunhao Yang and Sahil Shah and Sandeep Chinchali},
  journal= {arXiv preprint arXiv:2403.11021},
  year   = {2024}
}

Comments

Accepted by The European Conference on Computer Vision (ECCV) 2024

R2 v1 2026-06-28T15:22:56.542Z