English
Related papers

Related papers: EventFormer: A Node-graph Hierarchical Attention T…

200 papers

We introduce a new task called Referring Atomic Video Action Recognition (RAVAR), aimed at identifying atomic actions of a particular person based on a textual description and the video data of this person. This task differs from…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Kunyu Peng , Jia Fu , Kailun Yang , Di Wen , Yufan Chen , Ruiping Liu , Junwei Zheng , Jiaming Zhang , M. Saquib Sarfraz , Rainer Stiefelhagen , Alina Roitberg

Efficiently modeling spatial-temporal information in videos is crucial for action recognition. To achieve this goal, state-of-the-art methods typically employ the convolution operator and the dense interaction modules such as non-local…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Yuan Tian , Yichao Yan , Guangtao Zhai , Guodong Guo , Zhiyong Gao

Understanding events in texts is a core objective of natural language understanding, which requires detecting event occurrences, extracting event arguments, and analyzing inter-event relationships. However, due to the annotation challenges…

Computation and Language · Computer Science 2024-06-21 Xiaozhi Wang , Hao Peng , Yong Guan , Kaisheng Zeng , Jianhui Chen , Lei Hou , Xu Han , Yankai Lin , Zhiyuan Liu , Ruobing Xie , Jie Zhou , Juanzi Li

Event-based Action Recognition (EAR) possesses the advantages of high-temporal resolution capturing and privacy preservation compared with traditional action recognition. Current leading EAR solutions typically follow two regimes: project…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Meiqi Cao , Xiangbo Shu , Jiachao Zhang , Rui Yan , Zechao Li , Jinhui Tang

Event cameras offer significant advantages over conventional frame-based counterparts, including high temporal resolution, low latency, and energy efficiency. These characteristics make them suitable for high-speed and high-dynamic range…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Ramna Maqsood , Paulo Nunes , Luís Ducla Soares , Caroline Conti

Human intention prediction is a growing area of research where an activity in a video has to be anticipated by a vision-based system. To this end, the model creates a representation of the past, and subsequently, it produces future…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 Nada Osman , Guglielmo Camporese , Lamberto Ballan

Human Activity Recognition (HAR) has recently witnessed advancements with Transformer-based models. Especially, ActionFormer shows us a new perspectives for HAR in the sense that this approach gives us additional outputs which detect the…

Machine Learning · Computer Science 2025-05-28 Kunpeng Zhao , Asahi Miyazaki , Tsuyoshi Okita

Event cameras are novel sensors that report brightness changes in the form of asynchronous "events" instead of intensity frames. They have significant advantages over conventional cameras: high temporal resolution, high dynamic range, and…

Computer Vision and Pattern Recognition · Computer Science 2019-04-18 Henri Rebecq , René Ranftl , Vladlen Koltun , Davide Scaramuzza

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the importance of features…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Michael S. Ryoo , AJ Piergiovanni , Juhana Kangaspunta , Anelia Angelova

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Mengmeng Wang , Jiazheng Xing , Boyuan Jiang , Jun Chen , Jianbiao Mei , Xingxing Zuo , Guang Dai , Jingdong Wang , Yong Liu

Robust and flexible event representations are important to many core areas in language understanding. Scripts were proposed early on as a way of representing sequences of events for such understanding, and has recently attracted renewed…

Computation and Language · Computer Science 2017-11-22 Noah Weber , Niranjan Balasubramanian , Nathanael Chambers

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Aniket Rege , Arka Sadhu , Yuliang Li , Kejie Li , Ramya Korlakai Vinayak , Yuning Chai , Yong Jae Lee , Hyo Jin Kim

Action recognition and localization in complex, untrimmed videos remain a formidable challenge in computer vision, largely due to the limitations of existing methods in capturing fine-grained actions, long-term temporal dependencies, and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Liyang Peng , Sihan Zhu , Yunjie Guo

Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and efficiently, as standard uniform sampling is expensive and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Martin Q. Ma , Willis Guo , Aditya Agrawal , Ankit Gupta , Paul Pu Liang , Ruslan Salakhutdinov , Louis-Philippe Morency

Procedure Planning in instructional videos entails generating a sequence of action steps based on visual observations of the initial and target states. Despite the rapid progress in this task, there remain several critical challenges to be…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Ali Zare , Yulei Niu , Hammad Ayyubi , Shih-fu Chang

Script is a kind of structured knowledge extracted from texts, which contains a sequence of events. Based on such knowledge, script event prediction aims to predict the subsequent event. To do so, two aspects should be considered for…

Computation and Language · Computer Science 2022-12-19 Long Bai , Saiping Guan , Zixuan Li , Jiafeng Guo , Xiaolong Jin , Xueqi Cheng

Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support. We…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Han Zhang , Wanting Jiang , Tomasz Kornuta , Tian Zheng , Vidya Murali

The aim of this work is to detect and automatically generate high-level explanations of anomalous events in video. Understanding the cause of an anomalous event is crucial as the required response is dependant on its nature and severity.…

Computer Vision and Pattern Recognition · Computer Science 2021-12-13 Stanislaw Szymanowicz , James Charles , Roberto Cipolla

In recent days, streaming technology has greatly promoted the development in the field of livestream. Due to the excessive length of livestream records, it's quite essential to extract highlight segments with the aim of effective…

Multimedia · Computer Science 2022-06-13 Yang Zhao , Xuan Lin , Wenqiang Xu , Maozong Zheng , Zhengyong Liu , Zhou Zhao