English
Related papers

Related papers: Beyond Grounding: Extracting Fine-Grained Event Hi…

200 papers

Most prior art in visual understanding relies solely on analyzing the "what" (e.g., event recognition) and "where" (e.g., event localization), which in some cases, fails to describe correct contextual relationships between events or leads…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Aman Chadha , Gurneet Arora , Navpreet Kaloty

In this work we deal with the problem of high-level event detection in video. Specifically, we study the challenging problems of i) learning to detect video events from solely a textual description of the event, without using any positive…

Machine Learning · Computer Science 2015-11-26 Christos Tzelepis , Damianos Galanopoulos , Vasileios Mezaris , Ioannis Patras

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models remains opaque. We provide empirical clarity through…

Multi-modal large language models have demonstrated impressive performance across various tasks in different modalities. However, existing multi-modal models primarily emphasize capturing global information within each modality while…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Zhaowei Li , Qi Xu , Dong Zhang , Hang Song , Yiqing Cai , Qi Qi , Ran Zhou , Junting Pan , Zefeng Li , Van Tu Vu , Zhida Huang , Tao Wang

Human understanding of narrative is mainly driven by reasoning about causal relations between events and thus recognizing them is a key capability for computational models of language understanding. Computational work in this area has…

Computation and Language · Computer Science 2017-09-01 Zhichao Hu , Elahe Rahimtoroghi , Marilyn A Walker

Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by requiring identification of salient actions and associated…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Balaji Darur , Amanmeet Garg , Makarand Tapaswi

Dense video captioning aims to generate corresponding text descriptions for a series of events in the untrimmed video, which can be divided into two sub-tasks, event detection and event captioning. Unlike previous works that tackle the two…

Computer Vision and Pattern Recognition · Computer Science 2023-07-24 Qi Zhang , Yuqing Song , Qin Jin

Multimedia data is highly expressive and has traditionally been very difficult for a machine to interpret. Middleware systems such as complex event processing (CEP) mine patterns from data streams and send notifications to users in a timely…

Artificial Intelligence · Computer Science 2020-10-01 Piyush Yadav , Edward Curry

Tracking using bio-inspired event cameras has drawn more and more attention in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The first category needs…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Xiao Wang , Shiao Wang , Chuanming Tang , Lin Zhu , Bo Jiang , Yonghong Tian , Jin Tang

Multi-modal Event Reasoning (MMER) endeavors to endow machines with the ability to comprehend intricate event relations across diverse data modalities. MMER is fundamental and underlies a wide broad of applications. Despite extensive…

Artificial Intelligence · Computer Science 2024-04-17 Zhengwei Tao , Zhi Jin , Junqiang Huang , Xiancai Chen , Xiaoying Bai , Haiyan Zhao , Yifan Zhang , Chongyang Tao

Esports has rapidly emerged as a global phenomenon with an ever-expanding audience via platforms, like YouTube. Due to the inherent complexity nature of the game, it is challenging for newcomers to comprehend what the event entails. The…

Computation and Language · Computer Science 2024-06-14 Thye Shan Ng , Feiqi Cao , Soyeon Caren Han

Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevance accurately. To address this, we introduce GETReason…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Shikhhar Siingh , Abhinav Rawat , Chitta Baral , Vivek Gupta

Social platforms have emerged as crucial platforms for distributing information and discussing social events, offering researchers an excellent opportunity to design and implement novel event detection frameworks. Identifying unspecified…

Computation and Language · Computer Science 2025-06-12 Mohammadali Sefidi Esfahani , Mohammad Akbari

Event extraction is of practical utility in natural language processing. In the real world, it is a common phenomenon that multiple events existing in the same sentence, where extracting them are more difficult than extracting a single…

Computation and Language · Computer Science 2022-12-19 Xiao Liu , Zhunchen Luo , Heyan Huang

Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora. Existing large vision-language models (LVLMs) often underperform in this regime because…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Pengyu Yan , Akhil Gorugantu , Mahesh Bhosale , Abdul Wasi , Vishvesh Trivedi , David Doermann

Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A complex event relates to dense crowds, anomalous individuals, or…

Computer Vision and Pattern Recognition · Computer Science 2023-07-14 Weiyao Lin , Huabin Liu , Shizhan Liu , Yuxi Li , Rui Qian , Tao Wang , Ning Xu , Hongkai Xiong , Guo-Jun Qi , Nicu Sebe

Video-text retrieval (VTR) is an attractive yet challenging task for multi-modal understanding, which aims to search for relevant video (text) given a query (video). Existing methods typically employ completely heterogeneous visual-textual…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Haoran Wang , Di Xu , Dongliang He , Fu Li , Zhong Ji , Jungong Han , Errui Ding

Event cameras capture changes in brightness with microsecond precision and remain reliable under motion blur and challenging illumination, offering clear advantages for modeling highly dynamic scenes. Yet, their integration with natural…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Lingdong Kong , Dongyue Lu , Ao Liang , Rong Li , Yuhao Dong , Tianshuai Hu , Lai Xing Ng , Wei Tsang Ooi , Benoit R. Cottereau

Grounding events in videos serves as a fundamental capability in video analysis. While Vision Language Models (VLMs) are increasingly employed for this task, existing approaches predominantly train models to associate events with timestamps…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Fangxu Yu , Ziyao Lu , Liqiang Niu , Fandong Meng , Jie Zhou

Memes are a powerful tool for communication over social media. Their affinity for evolving across politics, history, and sociocultural phenomena makes them an ideal communication vehicle. To comprehend the subtle message conveyed within a…

Computation and Language · Computer Science 2023-05-30 Shivam Sharma , Ramaneswaran S , Udit Arora , Md. Shad Akhtar , Tanmoy Chakraborty