中文
相关论文

相关论文: EventFormer: A Node-graph Hierarchical Attention T…

200 篇论文

We introduce a new task called Referring Atomic Video Action Recognition (RAVAR), aimed at identifying atomic actions of a particular person based on a textual description and the video data of this person. This task differs from…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Kunyu Peng , Jia Fu , Kailun Yang , Di Wen , Yufan Chen , Ruiping Liu , Junwei Zheng , Jiaming Zhang , M. Saquib Sarfraz , Rainer Stiefelhagen , Alina Roitberg

Efficiently modeling spatial-temporal information in videos is crucial for action recognition. To achieve this goal, state-of-the-art methods typically employ the convolution operator and the dense interaction modules such as non-local…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Yuan Tian , Yichao Yan , Guangtao Zhai , Guodong Guo , Zhiyong Gao

Understanding events in texts is a core objective of natural language understanding, which requires detecting event occurrences, extracting event arguments, and analyzing inter-event relationships. However, due to the annotation challenges…

计算与语言 · 计算机科学 2024-06-21 Xiaozhi Wang , Hao Peng , Yong Guan , Kaisheng Zeng , Jianhui Chen , Lei Hou , Xu Han , Yankai Lin , Zhiyuan Liu , Ruobing Xie , Jie Zhou , Juanzi Li

Event-based Action Recognition (EAR) possesses the advantages of high-temporal resolution capturing and privacy preservation compared with traditional action recognition. Current leading EAR solutions typically follow two regimes: project…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Meiqi Cao , Xiangbo Shu , Jiachao Zhang , Rui Yan , Zechao Li , Jinhui Tang

Event cameras offer significant advantages over conventional frame-based counterparts, including high temporal resolution, low latency, and energy efficiency. These characteristics make them suitable for high-speed and high-dynamic range…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Ramna Maqsood , Paulo Nunes , Luís Ducla Soares , Caroline Conti

Human intention prediction is a growing area of research where an activity in a video has to be anticipated by a vision-based system. To this end, the model creates a representation of the past, and subsequently, it produces future…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Nada Osman , Guglielmo Camporese , Lamberto Ballan

Human Activity Recognition (HAR) has recently witnessed advancements with Transformer-based models. Especially, ActionFormer shows us a new perspectives for HAR in the sense that this approach gives us additional outputs which detect the…

机器学习 · 计算机科学 2025-05-28 Kunpeng Zhao , Asahi Miyazaki , Tsuyoshi Okita

Event cameras are novel sensors that report brightness changes in the form of asynchronous "events" instead of intensity frames. They have significant advantages over conventional cameras: high temporal resolution, high dynamic range, and…

计算机视觉与模式识别 · 计算机科学 2019-04-18 Henri Rebecq , René Ranftl , Vladlen Koltun , Davide Scaramuzza

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the importance of features…

计算机视觉与模式识别 · 计算机科学 2020-08-19 Michael S. Ryoo , AJ Piergiovanni , Juhana Kangaspunta , Anelia Angelova

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Mengmeng Wang , Jiazheng Xing , Boyuan Jiang , Jun Chen , Jianbiao Mei , Xingxing Zuo , Guang Dai , Jingdong Wang , Yong Liu

Robust and flexible event representations are important to many core areas in language understanding. Scripts were proposed early on as a way of representing sequences of events for such understanding, and has recently attracted renewed…

计算与语言 · 计算机科学 2017-11-22 Noah Weber , Niranjan Balasubramanian , Nathanael Chambers

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous,…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Aniket Rege , Arka Sadhu , Yuliang Li , Kejie Li , Ramya Korlakai Vinayak , Yuning Chai , Yong Jae Lee , Hyo Jin Kim

Action recognition and localization in complex, untrimmed videos remain a formidable challenge in computer vision, largely due to the limitations of existing methods in capturing fine-grained actions, long-term temporal dependencies, and…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Liyang Peng , Sihan Zhu , Yunjie Guo

Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and efficiently, as standard uniform sampling is expensive and…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Martin Q. Ma , Willis Guo , Aditya Agrawal , Ankit Gupta , Paul Pu Liang , Ruslan Salakhutdinov , Louis-Philippe Morency

Procedure Planning in instructional videos entails generating a sequence of action steps based on visual observations of the initial and target states. Despite the rapid progress in this task, there remain several critical challenges to be…

计算机视觉与模式识别 · 计算机科学 2024-09-26 Ali Zare , Yulei Niu , Hammad Ayyubi , Shih-fu Chang

Script is a kind of structured knowledge extracted from texts, which contains a sequence of events. Based on such knowledge, script event prediction aims to predict the subsequent event. To do so, two aspects should be considered for…

计算与语言 · 计算机科学 2022-12-19 Long Bai , Saiping Guan , Zixuan Li , Jiafeng Guo , Xiaolong Jin , Xueqi Cheng

Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support. We…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Han Zhang , Wanting Jiang , Tomasz Kornuta , Tian Zheng , Vidya Murali

The aim of this work is to detect and automatically generate high-level explanations of anomalous events in video. Understanding the cause of an anomalous event is crucial as the required response is dependant on its nature and severity.…

计算机视觉与模式识别 · 计算机科学 2021-12-13 Stanislaw Szymanowicz , James Charles , Roberto Cipolla

In recent days, streaming technology has greatly promoted the development in the field of livestream. Due to the excessive length of livestream records, it's quite essential to extract highlight segments with the aim of effective…

多媒体 · 计算机科学 2022-06-13 Yang Zhao , Xuan Lin , Wenqiang Xu , Maozong Zheng , Zhengyong Liu , Zhou Zhao