English

F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos

Computer Vision and Pattern Recognition 2025-04-16 v2 Artificial Intelligence

Abstract

Analyzing Fast, Frequent, and Fine-grained (F3^3) events presents a significant challenge in video analytics and multi-modal LLMs. Current methods struggle to identify events that satisfy all the F3^3 criteria with high accuracy due to challenges such as motion blur and subtle visual discrepancies. To advance research in video understanding, we introduce F3^3Set, a benchmark that consists of video datasets for precise F3^3 event detection. Datasets in F3^3Set are characterized by their extensive scale and comprehensive detail, usually encompassing over 1,000 event types with precise timestamps and supporting multi-level granularity. Currently, F3^3Set contains several sports datasets, and this framework may be extended to other applications as well. We evaluated popular temporal action understanding methods on F3^3Set, revealing substantial challenges for existing techniques. Additionally, we propose a new method, F3^3ED, for F3^3 event detections, achieving superior performance. The dataset, model, and benchmark code are available at https://github.com/F3Set/F3Set.

Keywords

Cite

@article{arxiv.2504.08222,
  title  = {F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos},
  author = {Zhaoyu Liu and Kan Jiang and Murong Ma and Zhe Hou and Yun Lin and Jin Song Dong},
  journal= {arXiv preprint arXiv:2504.08222},
  year   = {2025}
}

Comments

ICLR 2025; Website URL: https://lzyandy.github.io/f3set-website/