English
Related papers

Related papers: Localizing Moments in Video with Natural Language

200 papers

Given an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by the query. To address this task, existing methods can be roughly grouped into two groups: 1)…

Computer Vision and Pattern Recognition · Computer Science 2022-11-02 Shaoning Xiao , Long Chen , Jian Shao , Yueting Zhuang , Jun Xiao

Video corpus moment retrieval~(VCMR) is the task of retrieving a relevant video moment from a large corpus of untrimmed videos via a natural language query. State-of-the-art work for VCMR is based on two-stage method. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Danyang Hou , Liang Pang , Yanyan Lan , Huawei Shen , Xueqi Cheng

Untrimmed videos have interrelated events, dependencies, context, overlapping events, object-object interactions, domain specificity, and other semantics that are worth highlighting while describing a video in natural language. Owing to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Iqra Qasim , Alexander Horsch , Dilip K. Prasad

Moment retrieval in videos is a challenging task that aims to retrieve the most relevant video moment in an untrimmed video given a sentence description. Previous methods tend to perform self-modal learning and cross-modal interaction in a…

Computer Vision and Pattern Recognition · Computer Science 2023-02-21 Xin Sun , Xuan Wang , Jialin Gao , Qiong Liu , Xi Zhou

Video Moment Retrieval (VMR) aims to retrieve relevant moments of an untrimmed video corresponding to the query. While cross-modal interaction approaches have shown progress in filtering out query-irrelevant information in videos, they…

Artificial Intelligence · Computer Science 2024-08-26 Chenghua Gao , Min Li , Jianshuo Liu , Junxing Ren , Lin Chen , Haoyu Liu , Bo Meng , Jitao Fu , Wenwen Su

Video Corpus Moment Retrieval (VCMR) is a practical video retrieval task focused on identifying a specific moment within a vast corpus of untrimmed videos using the natural language query. Existing methods for VCMR typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Danyang Hou , Liang Pang , Huawei Shen , Xueqi Cheng

Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each query. This assumption does not always hold in real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Yiming Ding , Siyu Cao , Luyuan Jiao , Yixuan Li , Zitong Wang , Zhiyong Liu , Lu Zhang

Temporal action localization is an important and challenging task that aims to locate temporal regions in real-world untrimmed videos where actions occur and recognize their classes. It is widely acknowledged that video context is a…

Computer Vision and Pattern Recognition · Computer Science 2021-03-10 Xin Qin , Hanbin Zhao , Guangchen Lin , Hao Zeng , Songcen Xu , Xi Li

Describing visual data into natural language is a very challenging task, at the intersection of computer vision, natural language processing and machine learning. Language goes well beyond the description of physical objects and their…

Computer Vision and Pattern Recognition · Computer Science 2020-05-26 Iulia Duta , Andrei Liviu Nicolicioiu , Simion-Vlad Bogolin , Marius Leordeanu

In this paper we undertake the task of text-based video moment retrieval from a corpus of videos. To train the model, text-moment paired datasets were used to learn the correct correspondences. In typical training methods, ground-truth…

Computer Vision and Pattern Recognition · Computer Science 2021-06-28 Sho Maeoki , Yusuke Mukuta , Tatsuya Harada

The task of retrieving clips within videos based on a given natural language query requires cross-modal reasoning over multiple frames. Prior approaches such as sliding window classifiers are inefficient, while text-clip similarity driven…

Computation and Language · Computer Science 2019-04-08 Soham Ghosh , Anuva Agarwal , Zarana Parekh , Alexander Hauptmann

We present CLIP2Video network to transfer the image-language pre-training model to video-text retrieval in an end-to-end manner. Leading approaches in the domain of video-and-language learning try to distill the spatio-temporal video…

Computer Vision and Pattern Recognition · Computer Science 2021-06-22 Han Fang , Pengfei Xiong , Luhui Xu , Yu Chen

Associating image regions with text queries has been recently explored as a new way to bridge visual and linguistic representations. A few pioneering approaches have been proposed based on recurrent neural language models trained…

Computer Vision and Pattern Recognition · Computer Science 2017-04-18 Yuting Zhang , Luyao Yuan , Yijie Guo , Zhiyuan He , I-An Huang , Honglak Lee

Localizing events in videos based on semantic queries is a pivotal task in video understanding, with the growing significance of user-oriented applications like video search. Yet, current research predominantly relies on natural language…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Gengyuan Zhang , Mang Ling Ada Fok , Jialu Ma , Yan Xia , Daniel Cremers , Philip Torr , Volker Tresp , Jindong Gu

We present a Temporal Context Network (TCN) for precise temporal localization of human activities. Similar to the Faster-RCNN architecture, proposals are placed at equal intervals in a video which span multiple temporal scales. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2017-08-09 Xiyang Dai , Bharat Singh , Guyue Zhang , Larry S. Davis , Yan Qiu Chen

Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Weijia Liu , Jiuxin Cao , Bo Miao , Zhiheng Fu , Xuelin Zhu , Jiawei Ge , Bo Liu , Mehwish Nasim , Ajmal Mian

Multimodal machine translation (MMT), which mainly focuses on enhancing text-only translation with visual features, has attracted considerable attention from both computer vision and natural language processing communities. Most current MMT…

Computation and Language · Computer Science 2020-09-07 Huan Lin , Fandong Meng , Jinsong Su , Yongjing Yin , Zhengyuan Yang , Yubin Ge , Jie Zhou , Jiebo Luo

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Vladimir Iashin , Esa Rahtu

Video moment retrieval (VMR) aims to localize target moments in untrimmed videos pertinent to a given textual query. Existing retrieval systems tend to rely on retrieval bias as a shortcut and thus, fail to sufficiently learn multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Sunjae Yoon , Ji Woo Hong , Eunseop Yoon , Dahyun Kim , Junyeong Kim , Hee Suk Yoon , Chang D. Yoo

Giving machines the ability to imagine possible new objects or scenes from linguistic descriptions and produce their realistic renderings is arguably one of the most challenging problems in computer vision. Recent advances in deep…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Levent Karacan , Tolga Kerimoğlu , İsmail İnan , Tolga Birdal , Erkut Erdem , Aykut Erdem