English
Related papers

Related papers: EventFormer: A Node-graph Hierarchical Attention T…

200 papers

Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Baoyu Liang , Qile Su , Shoutai Zhu , Yuchen Liang , Chao Tong

While language models have become impactful in many real-world applications, video generation remains largely confined to entertainment. Motivated by video's inherent capacity to demonstrate physical-world information that is difficult to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Junhao Cheng , Liang Hou , Xin Tao , Jing Liao

An audio-visual event (AVE) is denoted by the correspondence of the visual and auditory signals in a video segment. Precise localization of the AVEs is very challenging since it demands effective multi-modal feature correspondence to ground…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Tanvir Mahmud , Diana Marculescu

Tracking using bio-inspired event cameras has drawn more and more attention in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The first category needs…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Xiao Wang , Shiao Wang , Chuanming Tang , Lin Zhu , Bo Jiang , Yonghong Tian , Jin Tang

Video Corpus Moment Retrieval (VCMR) is a practical video retrieval task focused on identifying a specific moment within a vast corpus of untrimmed videos using the natural language query. Existing methods for VCMR typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Danyang Hou , Liang Pang , Huawei Shen , Xueqi Cheng

Audio-visual event parsing plays a crucial role in understanding multimodal video content, but existing methods typically rely on offline processing of entire videos with huge model sizes, limiting their real-time applicability. We…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Xiao Yu , Yan Fang , Xiaojie Jin , Yao Zhao , Yunchao Wei

Scripts are structured sequences of events together with the participants, which are extracted from the texts.Script event prediction aims to predict the subsequent event given the historical events in the script. Two kinds of information…

Computation and Language · Computer Science 2021-11-01 Long Bai , Saiping Guan , Jiafeng Guo , Zixuan Li , Xiaolong Jin , Xueqi Cheng

The event-based Vision-Language Model (VLM) recently has made good progress for practical vision tasks. However, most of these works just utilize CLIP for focusing on traditional perception tasks, which obstruct model understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Pengteng Li , Yunfan Lu , Pinghao Song , Wuyang Li , Huizai Yao , Hui Xiong

Recent advances in zero-shot and few-shot classification heavily rely on the success of pre-trained vision-language models (VLMs) such as CLIP. Due to a shortage of large-scale datasets, training such models for event camera data remains…

Computer Vision and Pattern Recognition · Computer Science 2023-11-20 Ziyi Wu , Xudong Liu , Igor Gilitschenski

The neuromorphic event cameras, which capture the optical changes of a scene, have drawn increasing attention due to their high speed and low power consumption. However, the event data are noisy, sparse, and nonuniform in the…

Computer Vision and Pattern Recognition · Computer Science 2021-03-23 Chang Liu , Xiaojuan Qi , Edmund Lam , Ngai Wong

The event camera's low power consumption and ability to capture microsecond brightness changes make it attractive for various computer vision tasks. Existing event representation methods typically convert events into frames, voxel grids, or…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Bin Jiang , Zhihao Li , M. Salman Asif , Xun Cao , Zhan Ma

In this paper, we present ENTER, an interpretable Video Question Answering (VideoQA) system based on event graphs. Event graphs convert videos into graphical representations, where video events form the nodes and event-event relationships…

We propose Anticipative Video Transformer (AVT), an end-to-end attention-based video modeling architecture that attends to the previously observed video in order to anticipate future actions. We train the model jointly to predict the next…

Computer Vision and Pattern Recognition · Computer Science 2021-09-23 Rohit Girdhar , Kristen Grauman

We then introduce a novel hierarchical knowledge distillation strategy that incorporates the similarity matrix, feature representation, and response map-based distillation to guide the learning of the student Transformer network. We also…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Shiao Wang , Xiao Wang , Chao Wang , Liye Jin , Lin Zhu , Bo Jiang , Yonghong Tian , Jin Tang

Facial action units (AUs) play an indispensable role in human emotion analysis. We observe that although AU-based high-level emotion analysis is urgently needed by real-world applications, frame-level AU results provided by previous works…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Yingjie Chen , Jiarui Zhang , Tao Wang , Yun Liang

To improve the performance of long text generation, recent studies have leveraged automatically planned event structures (i.e. storylines) to guide story generation. Such prior works mostly employ end-to-end neural generation models to…

Computation and Language · Computer Science 2022-10-20 Chen Tang , Zhihao Zhang , Tyler Loakman , Chenghua Lin , Frank Guerin

Event cameras are neuromorphic vision sensors that record a scene as sparse and asynchronous event streams. Most event-based methods project events into dense frames and process them using conventional vision models, resulting in high…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Bochen Xie , Yongjian Deng , Zhanpeng Shao , Qingsong Xu , Youfu Li

Temporal action localization aims to predict the boundary and category of each action instance in untrimmed long videos. Most of previous methods based on anchors or proposals neglect the global-local context interaction in entire video…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Yizheng Ouyang , Tianjin Zhang , Weibo Gu , Hongfa Wang

Event cameras provide robust visual signals under fast motion and challenging illumination conditions thanks to their microsecond latency and high dynamic range. However, their unique sensing characteristics and limited labeled data make it…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Jianwen Cao , Jiaxu Xing , Nico Messikommer , Davide Scaramuzza

Next-token prediction serves as the foundational learning task enabling reasoning in LLMs. But what should the learning task be when aiming to equip MLLMs with temporal reasoning capabilities over video inputs? Existing tasks such as video…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Haonan Wang , Hongfu Liu , Xiangyan Liu , Chao Du , Kenji Kawaguchi , Ye Wang , Tianyu Pang
‹ Prev 1 2 3 10 Next ›