中文
相关论文

相关论文: EventFormer: A Node-graph Hierarchical Attention T…

200 篇论文

We introduce the Action Transformer model for recognizing and localizing human actions in video clips. We repurpose a Transformer-style architecture to aggregate features from the spatiotemporal context around the person whose actions we…

计算机视觉与模式识别 · 计算机科学 2019-05-20 Rohit Girdhar , João Carreira , Carl Doersch , Andrew Zisserman

Human action recognition in long-term videos, characterized by complex backgrounds and subtle action differences, poses significant challenges for traditional deep learning models due to computational overhead, difficulty in capturing…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Kaining Li , Shuwei He , Zihan Xu

Understanding videos is an important research topic for multimodal learning. Leveraging large-scale datasets of web-crawled video-text pairs as weak supervision has become a pre-training paradigm for learning joint representations and…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Gengyuan Zhang , Jinhe Bi , Jindong Gu , Yanyu Chen , Volker Tresp

As a branch of time series forecasting, stock movement forecasting is one of the challenging problems for investors and researchers. Since Transformer was introduced to analyze financial data, many researchers have dedicated themselves to…

统计金融 · 定量金融 2024-04-12 Chufeng Li , Jianyong Chen

We present TubeFormer-DeepLab, the first attempt to tackle multiple core video segmentation tasks in a unified manner. Different video segmentation tasks (e.g., video semantic/instance/panoptic segmentation) are usually considered as…

计算机视觉与模式识别 · 计算机科学 2023-03-07 Dahun Kim , Jun Xie , Huiyu Wang , Siyuan Qiao , Qihang Yu , Hong-Seok Kim , Hartwig Adam , In So Kweon , Liang-Chieh Chen

High-speed vision sensing is essential for real-time perception in applications such as robotics, autonomous vehicles, and industrial automation. Traditional frame-based vision systems suffer from motion blur, high latency, and redundant…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Riadul Islam , Joey Mulé , Dhandeep Challagundla , Shahmir Rizvi , Sean Carson

Despite advances in the application of MLLMs for various video tasks, video event prediction (VEP) remains relatively underexplored. VEP requires the model to perform fine-grained temporal modeling of videos and establish logical…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Qile Su , Jing Tang , Rui Chen , Lei Sun , Xiangxiang Chu

The paper-to-video task converts a research paper into a structured video abstract, distilling key concepts, methods, and conclusions into an accessible, well-organized format. While state-of-the-art video generation models demonstrate…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Jingwei Liu , Ling Yang , Hao Luo , Fan Wang , Hongyan Li , Mengdi Wang

Continuously-observed event occurrences, often exhibit self- and mutually-exciting effects, which can be well modeled using temporal point processes. Beyond that, these event dynamics may also change over time, with certain periodic trends.…

机器学习 · 计算机科学 2024-03-11 Sikun Yang , Hongyuan Zha

Moving object segmentation is critical to interpret scene dynamics for robotic navigation systems in challenging environments. Neuromorphic vision sensors are tailored for motion perception due to their asynchronous nature, high temporal…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Yusra Alkendi , Rana Azzam , Sajid Javed , Lakmal Seneviratne , Yahya Zweiri

Capturing complex hierarchical human activities, from atomic actions (e.g., picking up one present, moving to the sofa, unwrapping the present) to contextual events (e.g., celebrating Christmas) is crucial for achieving high-performance…

计算机视觉与模式识别 · 计算机科学 2024-09-16 Yanan Wang , Shuichiro Haruta , Donghuo Zeng , Julio Vizcarra , Mori Kurokawa

Predictive Process Monitoring focuses on predicting future states of ongoing process executions, such as forecasting the remaining time. Recent developments in Object-Centric Process Mining have enriched event data with objects and their…

机器学习 · 计算机科学 2024-04-17 Tim K. Smit , Hajo A. Reijers , Xixi Lu

Recently, event cameras have shown large applicability in several computer vision fields especially concerning tasks that require high temporal resolution. In this work, we investigate the usage of such kind of data for emotion recognition…

计算机视觉与模式识别 · 计算机科学 2023-04-14 Lorenzo Berlincioni , Luca Cultrera , Chiara Albisani , Lisa Cresti , Andrea Leonardo , Sara Picchioni , Federico Becattini , Alberto Del Bimbo

Event-based vision sensors, inspired by biological neural systems, asynchronously capture local pixel-level intensity changes as a sparse event stream containing position, polarity, and timestamp information. These neuromorphic sensors…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Tiantian Xie , Pengpai Wang , Rosa H. M. Chan

Event cameras are advantageous for tasks that require vision sensors with low-latency and sparse output responses. However, the development of deep network algorithms using event cameras has been slow because of the lack of large labelled…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Joachim Ott , Zuowen Wang , Shih-Chii Liu

Real-world graph typically evolve via a series of events, modeling dynamic interactions between objects across various domains. For dynamic graph learning, dynamic graph neural networks (DGNNs) have emerged as popular solutions. Recently,…

机器学习 · 计算机科学 2026-05-22 Xingtong Yu , Ruijuan Liang , Renhe Jiang , Dongyuan Li , Yunxiao Zhao , Xinming Zhang , Yuan Fang

This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in space and time,…

Given the enormous number of instructional videos available online, learning a diverse array of multi-step task models from videos is an appealing goal. We introduce a new pre-trained video model, VideoTaskformer, focused on representing…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Medhini Narasimhan , Licheng Yu , Sean Bell , Ning Zhang , Trevor Darrell

Analyzing laparoscopic surgery videos presents a complex and multifaceted challenge, with applications including surgical training, intra-operative surgical complication prediction, and post-operative surgical assessment. Identifying…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Sahar Nasirihaghighi , Negin Ghamsarian , Heinrich Husslein , Klaus Schoeffmann

Audio-visual event localization (AVEL) plays a critical role in multimodal scene understanding. While existing datasets for AVEL predominantly comprise landscape-oriented long videos with clean and simple audio context, short videos have…

多媒体 · 计算机科学 2025-04-10 Wuyang Liu , Yi Chai , Yongpeng Yan , Yanzhen Ren