English
Related papers

Related papers: EventFormer: A Node-graph Hierarchical Attention T…

200 papers

We introduce the Action Transformer model for recognizing and localizing human actions in video clips. We repurpose a Transformer-style architecture to aggregate features from the spatiotemporal context around the person whose actions we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-20 Rohit Girdhar , João Carreira , Carl Doersch , Andrew Zisserman

Human action recognition in long-term videos, characterized by complex backgrounds and subtle action differences, poses significant challenges for traditional deep learning models due to computational overhead, difficulty in capturing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Kaining Li , Shuwei He , Zihan Xu

Understanding videos is an important research topic for multimodal learning. Leveraging large-scale datasets of web-crawled video-text pairs as weak supervision has become a pre-training paradigm for learning joint representations and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Gengyuan Zhang , Jinhe Bi , Jindong Gu , Yanyu Chen , Volker Tresp

As a branch of time series forecasting, stock movement forecasting is one of the challenging problems for investors and researchers. Since Transformer was introduced to analyze financial data, many researchers have dedicated themselves to…

Statistical Finance · Quantitative Finance 2024-04-12 Chufeng Li , Jianyong Chen

We present TubeFormer-DeepLab, the first attempt to tackle multiple core video segmentation tasks in a unified manner. Different video segmentation tasks (e.g., video semantic/instance/panoptic segmentation) are usually considered as…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Dahun Kim , Jun Xie , Huiyu Wang , Siyuan Qiao , Qihang Yu , Hong-Seok Kim , Hartwig Adam , In So Kweon , Liang-Chieh Chen

High-speed vision sensing is essential for real-time perception in applications such as robotics, autonomous vehicles, and industrial automation. Traditional frame-based vision systems suffer from motion blur, high latency, and redundant…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Riadul Islam , Joey Mulé , Dhandeep Challagundla , Shahmir Rizvi , Sean Carson

Despite advances in the application of MLLMs for various video tasks, video event prediction (VEP) remains relatively underexplored. VEP requires the model to perform fine-grained temporal modeling of videos and establish logical…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Qile Su , Jing Tang , Rui Chen , Lei Sun , Xiangxiang Chu

The paper-to-video task converts a research paper into a structured video abstract, distilling key concepts, methods, and conclusions into an accessible, well-organized format. While state-of-the-art video generation models demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Jingwei Liu , Ling Yang , Hao Luo , Fan Wang , Hongyan Li , Mengdi Wang

Continuously-observed event occurrences, often exhibit self- and mutually-exciting effects, which can be well modeled using temporal point processes. Beyond that, these event dynamics may also change over time, with certain periodic trends.…

Machine Learning · Computer Science 2024-03-11 Sikun Yang , Hongyuan Zha

Moving object segmentation is critical to interpret scene dynamics for robotic navigation systems in challenging environments. Neuromorphic vision sensors are tailored for motion perception due to their asynchronous nature, high temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Yusra Alkendi , Rana Azzam , Sajid Javed , Lakmal Seneviratne , Yahya Zweiri

Capturing complex hierarchical human activities, from atomic actions (e.g., picking up one present, moving to the sofa, unwrapping the present) to contextual events (e.g., celebrating Christmas) is crucial for achieving high-performance…

Computer Vision and Pattern Recognition · Computer Science 2024-09-16 Yanan Wang , Shuichiro Haruta , Donghuo Zeng , Julio Vizcarra , Mori Kurokawa

Predictive Process Monitoring focuses on predicting future states of ongoing process executions, such as forecasting the remaining time. Recent developments in Object-Centric Process Mining have enriched event data with objects and their…

Machine Learning · Computer Science 2024-04-17 Tim K. Smit , Hajo A. Reijers , Xixi Lu

Recently, event cameras have shown large applicability in several computer vision fields especially concerning tasks that require high temporal resolution. In this work, we investigate the usage of such kind of data for emotion recognition…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Lorenzo Berlincioni , Luca Cultrera , Chiara Albisani , Lisa Cresti , Andrea Leonardo , Sara Picchioni , Federico Becattini , Alberto Del Bimbo

Event-based vision sensors, inspired by biological neural systems, asynchronously capture local pixel-level intensity changes as a sparse event stream containing position, polarity, and timestamp information. These neuromorphic sensors…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Tiantian Xie , Pengpai Wang , Rosa H. M. Chan

Event cameras are advantageous for tasks that require vision sensors with low-latency and sparse output responses. However, the development of deep network algorithms using event cameras has been slow because of the lack of large labelled…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Joachim Ott , Zuowen Wang , Shih-Chii Liu

Real-world graph typically evolve via a series of events, modeling dynamic interactions between objects across various domains. For dynamic graph learning, dynamic graph neural networks (DGNNs) have emerged as popular solutions. Recently,…

Machine Learning · Computer Science 2026-05-22 Xingtong Yu , Ruijuan Liang , Renhe Jiang , Dongyuan Li , Yunxiao Zhao , Xinming Zhang , Yuan Fang

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in space and time,…

Given the enormous number of instructional videos available online, learning a diverse array of multi-step task models from videos is an appealing goal. We introduce a new pre-trained video model, VideoTaskformer, focused on representing…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Medhini Narasimhan , Licheng Yu , Sean Bell , Ning Zhang , Trevor Darrell

Analyzing laparoscopic surgery videos presents a complex and multifaceted challenge, with applications including surgical training, intra-operative surgical complication prediction, and post-operative surgical assessment. Identifying…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Sahar Nasirihaghighi , Negin Ghamsarian , Heinrich Husslein , Klaus Schoeffmann

Audio-visual event localization (AVEL) plays a critical role in multimodal scene understanding. While existing datasets for AVEL predominantly comprise landscape-oriented long videos with clean and simple audio context, short videos have…

Multimedia · Computer Science 2025-04-10 Wuyang Liu , Yi Chai , Yongpeng Yan , Yanzhen Ren
‹ Prev 1 4 5 6 7 8 10 Next ›