English
Related papers

Related papers: EA-VTR: Event-Aware Video-Text Retrieval

200 papers

Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Zhaoyang Yang , Yurun Jin , Lizhe Qi , Cong Huang , Kai Chen

Video Anomaly Understanding (VAU) is essential for applications such as smart cities, security surveillance, and disaster alert systems, yet remains challenging due to its demand for fine-grained spatio-temporal perception and robust…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Liyun Zhu , Qixiang Chen , Xi Shen , Xiaodong Cun

Multimedia Event Extraction (MEE) aims to identify events and their arguments from documents that contain both text and images. It requires grounding event semantics across different modalities. Progress in MEE is limited by the lack of…

Computation and Language · Computer Science 2026-05-28 Yongkang Jin , Jianwen Luo , Jingjing Wang , Jianmin Yao , Yu Hong

In this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The proposed parsing approach simultaneously detects the temporal…

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges…

Computation and Language · Computer Science 2021-10-14 Ankit P. Shah , Shijie Geng , Peng Gao , Anoop Cherian , Takaaki Hori , Tim K. Marks , Jonathan Le Roux , Chiori Hori

We present CEIA, an effective framework for open-world event-based understanding. Currently training a large event-text model still poses a huge challenge due to the shortage of paired event-text data. In response to this challenge, CEIA…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Wenhao Xu , Wenming Weng , Yueyi Zhang , Zhiwei Xiong

The task of text-video retrieval aims to understand the correspondence between language and vision, has gained increasing attention in recent years. Previous studies either adopt off-the-shelf 2D/3D-CNN and then use average/max pooling to…

Computer Vision and Pattern Recognition · Computer Science 2022-06-02 Ning Han , Jingjing Chen , Chuhao Shi , Yawen Zeng , Guangyi Xiao , Hao Chen

In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Eitan Shaar , Ariel Shaulov , Gal Chechik , Lior Wolf

Video understanding is inherently intention-driven-humans naturally focus on relevant frames based on their goals. Recent advancements in multimodal large language models (MLLMs) have enabled flexible query-driven reasoning; however,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Ziqiang Xu , Qi Dai , Tian Xie , Yifan Yang , Kai Qiu , DongDong Chen , Zuxuan Wu , Chong Luo

The bio-inspired event cameras or dynamic vision sensors are capable of asynchronously capturing per-pixel brightness changes (called event-streams) in high temporal resolution and high dynamic range. However, the non-structural…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Qiang Qu , Yiran Shen , Xiaoming Chen , Yuk Ying Chung , Tongliang Liu

Vision-language retrieval is an important multi-modal learning topic, where the goal is to retrieve the most relevant visual candidate for a given text query. Recently, pre-trained models, e.g., CLIP, show great potential on retrieval…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Haojun Jiang , Jianke Zhang , Rui Huang , Chunjiang Ge , Zanlin Ni , Shiji Song , Gao Huang

Good temporal representations are crucial for video understanding, and the state-of-the-art video recognition framework is based on two-stream networks. In such framework, besides the regular ConvNets responsible for RGB frame inputs, a…

Computer Vision and Pattern Recognition · Computer Science 2018-05-22 Wanjia Liu , Huaijin Chen , Rishab Goel , Yuzhong Huang , Ashok Veeraraghavan , Ankit Patel

Grounding events in videos serves as a fundamental capability in video analysis. While Vision Language Models (VLMs) are increasingly employed for this task, existing approaches predominantly train models to associate events with timestamps…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Fangxu Yu , Ziyao Lu , Liqiang Niu , Fandong Meng , Jie Zhou

Composed video retrieval (CoVR) is a challenging problem in computer vision which has recently highlighted the integration of modification text with visual queries for more sophisticated video search in large databases. Existing works…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Omkar Thawakar , Muzammal Naseer , Rao Muhammad Anwer , Salman Khan , Michael Felsberg , Mubarak Shah , Fahad Shahbaz Khan

Video Moment Retrieval (MR) aims to localize moments within a video based on a given natural language query. Given the prevalent use of platforms like YouTube for information retrieval, the demand for MR techniques is significantly growing.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Seojeong Park , Jiho Choi , Kyungjune Baek , Hyunjung Shim

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of videos. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Thomas Hummel , Shyamgopal Karthik , Mariana-Iuliana Georgescu , Zeynep Akata

Long-form video understanding requires efficient navigation of extensive visual data to pinpoint sparse yet critical information. Current approaches to longform video understanding either suffer from severe computational overhead due to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Te Yang , Xiangyu Zhu , Bo Wang , Quan Chen , Peng Jiang , Zhen Lei

Video-text retrieval is an important yet challenging task in vision-language understanding, which aims to learn a joint embedding space where related video and text instances are close to each other. Most current works simply measure the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-02 Peng Wu , Xiangteng He , Mingqian Tang , Yiliang Lv , Jing Liu

Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mohamed Eltahir , Osamah Sarraj , Mohammed Bremoo , Mohammed Khurd , Abdulrahman Alfrihidi , Taha Alshatiri , Mohammad Almatrafi , Tanveer Hussain

The relations expressed in user queries are vital for cross-modal information retrieval. Relation-focused cross-modal retrieval aims to retrieve information that corresponds to these relations, enabling effective retrieval across different…

Computer Vision and Pattern Recognition · Computer Science 2023-07-31 Yan Gong , Georgina Cosma , Axel Finke
‹ Prev 1 8 9 10 Next ›